copernic

Version Control System for Structured Data

Visit Website
February 28, 2023 README.md neats the point

Copernic use a novel approach to store triples in an ordered key-value store. It use FoundationDB database storage engine to deliver a pragmatic versatile ACID-compliant versioned triple store where people can cooperate around the making of knowledge. Copernic only stores changes between versions. It has also a snapshot of the latest version. copernic does not rely on the theory of patches introduced by Darcs but re-use some its vocabulary. Copernic is the future.

https://github.com/amirouche/copernic#readme

Comment

February 28, 2023 Easy does not sell

I know people have based other internal products on Copernic. The algorithm is so easy that it does not require maintenance. The only thing that may require more documentation is FoundationDB production use. I do not host Copernic for myself.

Anyway, my point is that nobody requested any support on it. Easy does not sell.

PS: I got one request for a talk, still no money.

Comment

February 28, 2021 Copernic in production

I re-used the algorithm, and re-implemented the datastructures on top of FoundationDB 6.3, taking advantage of asynchronous caps of CPython 3.7+ for a client.

That was part of a bigger task, involing a time sensitive extract-transform-load data pipeline.

The new code was 350 times faster using a single core on my dev setup, 30 times less energy hungry.

Comment

June 25, 2020 120 times fewer storage requirements

Since I started I know how crazy the idea to replace wikidata platform with a modular monolith is. I think that is the future. That is why I am working on it.

When I started last year the requirements in terms of SSD disk space were the following relative to the original size of the data:

  • 120 times, for storing and querying the data (SRFI-168 vnstore factor when n=5)
  • 2.55 times overhead for a single replication database (FoundationDB)
  • 2 times for single machine resilience (RAID 1)

You end up with a factor of 612. That is big already. My demo application needs to store the whole wikidata that is around 4TB. Eventually, using the original design, it required 2 448 TB that is around 2.5 petabytes.

Today, I end up a suite of experiments to reduce the required disk space to 20 TB :-)

Comment

February 20, 2020 Released v0.1.0

I changed the name from datae to copernic.space. I reworked the product almost completly. Instead of Scheme LISP, I rely on Python and Django. Instead of both WiredTiger and FoundationDB, I will focus on FoundationDB.

In practice, instead of "git for structured data", I will focus on a "scalable wikidata with change-request mechanic". It is still versioned. However there is no more general Directed-Acyclic-Graph history. In other words, the history is a single branch with stashed changes.

I posted http://copernic.space/ as Show HN and several related subreddit. The server was hit 5000 times, 1500 unique visitors. I managed to have 2 points on HN and -1 on reddit. lobste.rs was more generous with 3 points.

I think the current index page is not good enough, far from it. But it goes to the point, after editing:

copernic aims to make practical cooperation around the creation, publication, storage, re-use and maintenance of knowledge bases, and structured data that are bigger than memory.

Since I am confident with the backend and database code, I will focus on the ui/ux and onboarding.

The main take-aways, I have to share:

  • Do not try to bring too much new (big) things at the same time, I my case selling the idea of "versioned structured data at scale", is already big enough. The purpose of the product should be focused.

  • There is a lot of legacy and debt on in this field of work (data warehouses), hence lot of momentum in existing products. A new product, that is not backward compatible, will causes lot of frictions. In practice, moving several terabytes of data from one system to another, takes several months of hardware work. Do not expect people to think about moving to your solution, just because it is the good thing.

Comment

June 19, 2019 I decided to work full-time on datae

I have quit my job and decided to work full-time on datae. Not sure, how this will work as I have an opinionated idea about how my business will work. I want to work in the open. I want to make a living as consulting shop. And I want to use Scheme programming language to build that product.

Comment

About

There is a need to foster cooperation around data.