6 ms·
Graviton Database: ZFS for key-value stores
- ysleepy 6y agoNice! I implemented pretty much the same trade off set in an authenticated storage system. single writer, radix merkle tree, persistent storage, hashed keys, proofs. I guess it is a local maxima within that trade off space. I like how the time travelling/history is always touted as a feature (which it is), but it really just means the garbage collector/pruning part of the transaction engine is missing. Postgres and other mvcc systems could all be doing this, but they don't. The hard part of the feature is being able to turn it off. I'll probably have a look around later, the diffing looks interesting, not sure yet if it's done using the merkle tree (likely) or some commit walking algorithm.
- mulander 6y ago> I like how the time travelling/history is always touted as a feature (which it is), but it really just means the garbage collector/pruning part of the transaction engine is missing. Postgres and other mvcc systems could all be doing this, but they don't. Postgres actually did tout it as a feature in "THE IMPLEMENTATION OF POSTGRES" by Michael Stonebraker, Lawrence A. Rowe and Michael Hirohama[1] search for "time travel" in the PDF. I added the relevant quotes below for easier access ;) This was back when PostgreSQL had the postquel language (before SQL was added) there was special syntax to access data at specific points in time: > The second benefit of a no-overwrite storage manager is the possibility of time travel. As noted earlier, a user can ask a historical query and POSTGRES will automatically return information from the record valid at the correct time. Quoting the paper again: > For example to find the salary of Sam at time T one would query: retrieve (EMP.salary) using EMP [T] where EMP.name = "Sam" > POSTGRES will automatically find the version of Sam’s record valid at the correct time and get the appropriate salary. [1] - https://dsf.berkeley.edu/papers/ERL-M90-34.pdf https://dsf.berkeley.edu/papers/ERL-M90-34.pdf
- rmetzler 6y agoIs this still possible with Postgres?
- mulander 6y agoYes and no or to be precise - to a certain degree but not through an exposed language feature. PostgreSQL still does copy-on-write so the old versions of the row exist and are present in storage. However now there is an autovacuum process going over the records regularly marking those no longer seen by any transactions as re-usable so eventually the old records would get overwritten. You can get at the older versions of the rows directly on disk or perhaps it would be possible to get the db to return such older versions of the rows. It seems that by default even trying to get at them with `ctid` is not possible so that may require hacking PostgreSQL itself or using some extension which seem to actually exist[1]. [1] - https://github.com/omniti-labs/pgtreats/tree/master/contrib/pg_dirtyread https://github.com/omniti-labs/pgtreats/tree/master/contrib/...
- BryanG2 6y agoSomeone paste timing results of diffing for very large data sets.
- Rochus 6y agoWhat is the use case? Why is it important that "All keys, values are backed by blake 256 bit checksum"?
- jopari 6y agoIt seems to be intended as the backend database for the Dero blockchain smart contract platform: https://medium.com/deroproject/graviton-zfs-for-key-value-stores-4e48a4831a6a https://medium.com/deroproject/graviton-zfs-for-key-value-st... The post claims: "The features included in Graviton provide the missing functionality that prevented Stargate RC1 from reaching deployment on our mainnet." I'm not sure, but I guess that this checksumming is relevant for storing the Merkle trees encoding the blockchain. I don't know why the previous choice of database wasn't suitable.
- naivedevops 6y agoZFS stores the checksums of files to prevent bit rotting. Since they are comparing their database to ZFS, I guess it stores the checksums for the same reason. If bit rotting occurs, you don't need to discard the entire database, just the affected entry. If the entry was already there for some time, you might even be able to restore it from a backup.
- GordonS 6y agoIsn't a 256-bit Blake hash a little OTT, versus a simple CRC, or even a faster, smaller hash like MurmurHash or Jenkins-one-at-a-time?
- jlokier 6y agoIt's a cryptographic hash, so it will detect tampering with the data, which a simple CRC, MurmerHash or Jenkins would not.
- GordonS 6y agoStill, I'd like an option to use a faster, more efficient CRC or hash - bit rot is usually the main threat, rather than tampering. Not to mention that if a user can tamper with the data they can probably just create a new hash at the same time. Using a cryptographic hash as a souped-up CRC seems rather odd, given how many more CPU cycles and RAM it will use, but I don't know the reasoning behind the decision; there must be one.
- nickcw 6y agoWhat I'd really like is a multiprocess safe embeddable database written in pure Go. So a database which is safe to read and write from separate processes. Unfortunately I don't think this one is multiprocess safe.
- carlosf 6y agoNot sure if easy to embed, but Consul does offer strong consistency if explicitly configured.
- sneak 6y agoI too feel the “pure Go” pull, but is your use case so precarious or latency-sensitive that you can’t simply use SQLite? That’s what I do in these situations.
- jopari 6y agoI checked with the devs and they say writing is only safe from a single process, but reading is multiprocess safe. So I think this confirms your thought.
- ramoz 6y agoBadger supports concurrent ACID. https://github.com/dgraph-io/badger https://github.com/dgraph-io/badger
- lalalarororo 6y agoDERO the real BITCOIN
- deleted 6y ago[deleted]
- TomTinks 6y agoThis is definitely something to look into. so far dero looks like a pretty solid project with out of the box thinking.
- AtlasBarfed 6y ago... doesn't cassandra do a lot of this?
- ramoz 6y agoCassandra is not, traditionally/practically, an embedded db.
- byteshock 6y agoIf latency and performance is an issue there are also solutions like RocksDB or LevelDB
- jopari 6y agoThere's a brief comparison with RocksDB and LevelDB in the README file, which concludes: "If you require a high random write throughput or you need to use spinning disks then LevelDB could be a good choice unless there are requirements of versioning, authenticated proofs or other features of Graviton database."
- byteshock 6y agoThis was a reply to another comment in the thread that suggested a user use sqlite. I commented using the Octal ios app. Not sure why it didn’t post it correctly....
- moralestapia 6y agoI love the idea but I think you (author) need a lot of time/support polishing this. You need a team probably. Also, >Superfast proof generation time of around 1000 proofs per second per core. Does this limit in any way things like read/write perfomance or usability in general?
- jopari 6y agoI asked the devs, and apparently "proof generation won't affect general read/write performance". But I guess this is the kind of thing where benchmarks might be useful, and there don't seem to be any public yet.
- ysleepy 6y agoThe proof is just the nodes in the path of the merkle tree to the root. (or all sibling nodes of the path to the root). So proof "generation" is just fetching the nodes and sending them to the client.
- jjirsa 6y ago> You need a team probably The cardinal rule of database development: http://www.dbms2.com/2013/03/18/dbms-development-marklogic-hadoop/ http://www.dbms2.com/2013/03/18/dbms-development-marklogic-h...
- moralestapia 6y agoYup, I do not mean to discourage the authors. I truly like the project and I have a few things in mind that could make use of it already. (Heck, one of them is pretty much Graviton + a front end). But I cannot just jump into it as there's some real money involved and no one wants to experiment with that. I see a bright future for Graviton, once it becomes tested and stable in production environments.
- jopari 6y agoI'm sure the devs would be interested in hearing about your use cases, should you open an issue on the repo, or got in touch via https://dero.io/#contacts-section https://dero.io/#contacts-section (I'm not on the team, just interested in the project.)
- ramoz 6y agoComparison to Badger? Badger is also go-native and, for me, has been exceptional at scale and for read-heavy workloads on SSD. Ref: https://github.com/dgraph-io/badger https://github.com/dgraph-io/badger
- jopari 6y agoI think the key differentiating feature of Graviton is the tree of authenticated proofs of data consistency. (AFAICT this is particularly important for scalably updating and verifying a large blockchain history.)
- ramoz 6y agoAh, figured as much but am not as familiar with that use case. Thanks!
- jjirsa 6y agoDefine “at scale”
- ramoz 6y agosounds like you want to define that for me
- derefr 6y agoDoes anyone know of an embedded key-value store that does do versioning/snapshots, but doesn’t bother with cryptographic integrity (and so gets better OLAP performance than a Merkle-tree-based implementation)? My use-case is a system that serves as an OLAP data warehouse of representations of how another system’s state looked at various points in history. You’d open a handle against the store, passing in a snapshot version; and then do OLAP queries against that snapshot. Things that make this a hard problem: The dataset is too large to just store the versions as independent copies; so it really needs some level of data-sharing between the snapshots. But it also needs to be fast for reads, especially whole-bucket reads—it’s an OLAP data warehouse. Merkle-tree-based designs really suck for doing indexed table scans. But, things that can be traded off: there’d only need to be one (trusted) writer, who would just be batch-inserting new snapshots generated by reducing over a CQRS/ES event stream. It’d be that (out-of-band) event stream that’d be the canonical, integrity-verified, etc. representation for all this data. These CQRS state-aggregate snapshots would just be a cache. If the whole thing got corrupted, I could just throw it all away and regenerate it from the CQRS/ES event stream; or, hopefully, “rewind” the database back to the last-known-good commit (i.e. purge all snapshots above that one) and then regenerate only the rest from the event stream. I’m not personally aware of anything that targets exactly this use case. I’m working on something for it myself right now. Two avenues I’m looking into: • something that acts like a hybrid between LMDB and btrfs (i.e. a B-tree with copy-on-write ref-counted pages shared between snapshots, where those snapshots appear as B-tree nodes themselves) • “keyframe” snapshots as regular independent B-trees, maybe relying on L2ARC-like block-level dedup between them; “interstitial” snapshots as on-disk HAMT ‘overlays’ of the last keyframe B-tree, that share nodes with other on-disk HAMTs, but only within their “generation” (i.e. up to the next keyframe), such that they can all be rewritten/compacted/finalized once the next keyframe arrives, or maybe even converted into “B-frames” that have forward-references to data embedded in the next keyframe.
- ysleepy 6y agoWell, it all depends on what operations you need. exact key lookup/write? -> Easy, just use any KV store and append the version to the key, then do a ceil/floor lookup in the KV with the key+target version. Supporting efficient range scans is hard though. EDIT: yeah, OLAP will be hard.
- bdcravens 6y agoYou can run a Graviton database. You can also run a database on a Graviton: https://aws.amazon.com/about-aws/whats-new/2020/07/announcing-preview-for-amazon-rds-m6g-and-r6g-instance-types/ https://aws.amazon.com/about-aws/whats-new/2020/07/announcin... For best results, run Graviton on a Graviton: https://aws.amazon.com/ec2/graviton/ https://aws.amazon.com/ec2/graviton/
- Mister_Snuggles 6y agoNaming things is difficult.
- ethbr0 6y agoNot according to git! We can just call this: ad58cd9088995cfb528187b11c275dad60ce2ec5 And the chip: 59b54f61dd17c27744e884542e35b34172e2cc79 So easy!
- aftbit 6y ago>Graviton is currently alpha software. More like the "BTRFS for key-value stores" ;) Kidding aside, I dislike when new unproven software claims the name of industry standards like this. When I saw the headline, I was hoping this somehow actually leveraged ZFS's storage layer, but actually it is just a new database that thinks Copy-on-Write is cool.
- innagadadavida 6y agoTitle is very click baity, this is just another kv store, completely unrelated to ZFS.
- mayama 6y agoEven their README is click baity then. I quickly glanced at their repo and thought it somehow is related to ZFS, before reading comments here.
- BryanG2 6y agoEvery new unproven software is alpha during launch. Yet to see any software perfect since launch. Graviton provides ZFS leverages and features beyond ZFS. :)