4 ms·
could you describe major differences and overlaps between your solution and Datomics, MarkLogic temporal documents (https://docs.marklogic.com/guide/temporal/in
by platform 8y ago
could you describe major differences and overlaps between your solution and Datomics, MarkLogic temporal documents (https://docs.marklogic.com/guide/temporal/intro https://docs.marklogic.com/guide/temporal/intro).
Some compare with any other database 'time travel' feature would be helpful as well.
Thank you
- lichtenberger 8y agoThanks so much for asking. I'm not sure how they are actually storing their bitemporal documents, but I think it must be some kind of a huge B(+)-tree maybe. I think the cool thing is that we have revision root pages and are able to reconstruct each revision in merely the same time. We currently have a tree, very similar to how ZFS stores the objects on-disk. It's more or less a form of a hash array based trie. The number of levels of the indirect pages currently is static, but I'll change this and only create a new level once it's really needed as in ZFS. During every change to the resources in Sirix never in-place changes occur, instead it does a copy-on-write of the involved page and a pointer to the former version is created. This is necessary for our versioning algorithms as it has to fetch at most N former versions of the page, depending on if the page has been modified (usually N should be between 2 and 5 or something like that). The index structures currently are AVL-trees (thus also versioned), simply stored as data records in the leafes, but it would be best to plug in a B-tree in the future. On a higher level I have added many operations, which are usually not found in XML databases, for instance as our internal tree structure is a kind of persistent DOM firstChild/rightSibling/leftSibling/parent encoding I added move operations to move subtree. A path summary stores all paths in the resource and is kept updated at all times as are optional indexes on paths, elements, attributes or content-and-structure. The interesting thing is with the help of extensions to Brackit(.org), which also has some basic JSON navigation primitives already built in I was able to add temporal axis. I added several XPath axis to navigate in time, for instance first::, last::, all-time::, next::, previous::, future::, past::*... we also allow diffing, opening a resource in a specific revision...
- platform 8y agointeresting. last:: -- I assume, means 'most-recent'. maintaining pre-calculated 'most recent' is very useful (as long as I can ask 'what was most recent, say, yesterday at 1pm'). Most of the 'hand-made' append-only data schema design suffer from not being able to return 'joined' most-recent datums, quickly. Because a 'usual' implementation requires doing join with select max on business or system (or both) time stamp My understanding that both DBs I mentioned previously do copy of write for temporal data, but I am not sure if they do any document level-diffs.
- lichtenberger 8y agoYes, the revisions can be reconstructed in merely the same time, for instance it doesn't matter if you open the most recent revision or any past revision regarding system/transaction time. We also do not have to store the transaction time more than once (in the revision - root page). And if you do a fine granular modification, we usually do not copy and rewrite the whole page with currently 512 records (it depends on the versioning algorithm you use).
- lichtenberger 8y agoI think they might store a bitemporal that is two dimensional index. However, I think this is not really efficient to reconstruct a specific revision in comparison to our approach, where we just have to follow pointers from the revision root pages.
- lichtenberger 8y agoPh.D. thesis which might also help from Marc (from which the idea and everything originates) http://www.uni-konstanz.de/mmsp/pubsys/publishedFiles/Kramis2014.pdf http://www.uni-konstanz.de/mmsp/pubsys/publishedFiles/Kramis... From Sebastian Graf: https://kops.uni-konstanz.de/bitstream/handle/123456789/27250/Graf_272505.pdf?sequence=1&isAllowed=y https://kops.uni-konstanz.de/bitstream/handle/123456789/2725... The XQuery compiler from Sebastian Baechle: http://wwwlgis.informatik.uni-kl.de/cms/fileadmin/publications/2013/Dissertation-Baechle.pdf http://wwwlgis.informatik.uni-kl.de/cms/fileadmin/publicatio...
- lichtenberger 8y agoDuring our copy-on-write we also do not necessarily store the whole record-page (leaf page with data/records), but depending on the versioning algorithm only changed records since some past state. The sliding snapshot algorithm for instance (with a sliding window) is very clever in balancing read-/write-performance.