
SirixDB
Efficient snapshotting through a novel versioning algorithm
Hi,
I'm just finishing a native JSON layer for Sirix.io, a temporal database system, which implements a novel versioning algorithm and brings versioning to a sub-file level. It is heavily inspired by ZFS (the operating system). I began work on the project already during my studies at the University of Konstanz in 2006 (well, I guess now it's almost completely rewritten due to constant refactoring), whereas Marc Kramis started the project back then and Sebastian Graf added and refactored a lot of stuff, too during his Ph.D. I'm now more than ever confident that retaining the history of data for trend analysis, simple undo/redo operations, data audits and so on is crucial, as the nature of flash drives as for instance SSDs also encourages to write in batches while it's not simply possible to overwrite data in-place and random reads are a lot faster in stark contrast to mechanical disks.
I've extended an XQuery processor (both to work with XML and JSON) to provide functions and temporal axis for time travel queries, implemented a diff-algorithm to import the differences (to detect and apply a hopefully minimal or near minimal edit script) between several revisions of (currently) XML-documents, implemented an ID-based diff-algorithm to compare revisions once stored in Sirix (the node identifiers are always stable and never reorganized). Furthermore the differences for our XML storage can now be emitted in the form of XQuery Update Scripts. Once we add JSON update expressions in our XQuery binding, we can also emit a script for the differences encountered in JSON resources).
I've also implemented a RESTful, non-blocking API for both the XML and JSON storage as well as user defined, typed index-structures in the form of versioned AVL-trees.
As I'm currently finishing the JSON stuff, I'd be happy to receive any kind of feedback :-) I also want to publish a new version (0.9) in a couple of days.
You can find some documentation on the recently built website. However I'm still working on it :-)
Kind regards
Johannes
#product-feedback-request
About
We believe that a temporal database has to be both concise (novel versioning algorithm called sliding window), easy to use (allow sophisticated time travel queries) as well as efficient.

Comment