4 ms·
TerminusDB (co founder here) was partially inspired by the git scraping approach to the revision control for data problem. We built a database that gives you al
by LukeEF 6y ago
TerminusDB (co founder here) was partially inspired by the git scraping approach to the revision control for data problem. We built a database that gives you all of the functionality of git, but in a database so you can query and with a commit graph etc. Has git semantics for clone, fork, rebase, merge and the other major functions. We store & share deltas and use succinct data structures that allow us to pass around in-memory DBs that you can then query in place. We're relatively new, but happy to see versioning data wherever it might land. Open source forever so really trying to be git for data. Technical white paper on our structure:
https://github.com/terminusdb/terminusdb-server/blob/dev/docs/whitepaper/terminusdb.pdf https://github.com/terminusdb/terminusdb-server/blob/dev/doc...
- LukeEF 6y agoOne comment on using git, isn't there a pull:push bottle neck that means it can't service any workload that has any write rate greater than (say) one write every 30 seconds. Write is subject to interpretation of course given commit != push (and i am unashamedly suggesting that a DB is better)
- Quenty 6y agoVery cool that it’s based off of RDF triples. I’m really curious about your product. What does performance look like? The TerminusDB website says it’s in memory? Can you have a database larger than your RAM?
- LukeEF 6y agoRDF triples turn out to be a crucial part of the architecture as they make describing deltas really straightforward. It is just these triples were added and these triples were taken away. Performance is good - you get a degradation in query time as you build more appended layers, but you can squash these to a single plane to speed up. Often we have a query branch where the layers are optimized for query and another branch with all the commit history in place. We are working on something we call Delta roll ups at the moment - these are like squashes that keep the history. Hopefully you'll soon be able to automate the roll ups to keep query performance at a specified level (something like the vacuum cleaner in postgres). It is in-memory, so you are limited to what's in RAM for querying, but it persists to disk, and we are betting that memory is going to get bigger and cheaper over the next while.
- mark_l_watson 6y ago+1 just wanted to say thanks for pointing out the use of RDF, I would have missed that. I have been into the semantic web since year zero, and after years of doing deep learning work I just started a Knowledge Graph job two weeks ago. Anyway thanks, I am going to dig into TerminusDB as soon as I get back from my morning hike. EDIT: wow, TerminousDB is written in Swi-Prolog.
- LukeEF 6y agoYes - the server is in SWIPL and the distributed store is in Rust. Great combo we think. We tried to take some of the best ideas from the semantic web and make them as practical as possible. Great to hear that people are getting knowledge graph jobs out there!