8 ms·
An embedded database written in Rust
- tormeh 8y agoNo clustering :( I feel like good multi-master asynchronous and synchronous clustering is truly the frontier in DBs.
- keithnz 8y agowhat do you think of FoundationDb?
- krenoten 8y agoI am a total devotee to their approach to building simulable systems, although I seek to push it even farther and integrate lineage driven fault injection from an early stage. I see a lot of cool things in what they have done, technically. Sled is free from day one.
- jinqueeny 8y agoWhat about TiKV (https://github.com/pingcap/tikv https://github.com/pingcap/tikv), a distributed transactional Key-Value store?
- krenoten 8y agoUse it for large scale HTAP! It's great for its flexible use cases at high scales :)
- Groxx 8y agoclustering and embedded seem like almost opposite ends of the spectrum
- spacenick88 8y agoThen again embedded databases are a great building block for things like etcd. It also might even make sense to have something like this in process because that would remove quite a few failure modes coming from the client connection and simplify deployment
- fleetfox 8y agoWhat would be the use-case for multi-master clustering in embedded?
- fowl2 8y ago"real" serverless - mesh databases! ;p
- krenoten 8y agoThis is honestly a use case I'm experimenting with using a mix of CRDTs and OT. Our systems are becoming more and more location agnostic and I don't feel that our current data infrastructure is adequate to serve the workloads we're going to be facing as compute migrates to the edge.
- jacquesm 8y agoYou can't cluster an embedded database, this comment makes not sense. Compare with SQLite, not with Orcale or Postgres.
- krenoten 8y agoCurrently it's even more basic. The current usable parts are a pagecache following the llama approach, some great testing utility libraries, and an index (that you can use as a kv) that follows the bwtree approach. Later it will have structured access support, but it needs some more db components to get there. It is a construction kit as well as a kv.
- Hello71 8y agoSo... SQLite FoundationDB?
- spacenick88 8y agoNot really, when the embedded database is allowed to run its own threads and network connections embedding something like etcd makes perfect sense. It removes the failure modes of the client connection and simplifies deployment.
- krenoten 8y agoIt's modular, and there is a paxos implementation, but it has been built totally in simulation so far and I haven't plugged it into an io layer yet. But this is trivial. That said, sled will always be a bwtree index, and the other modular crates will stand on their own.
- tzahola 8y agoYou don't say?? No clustering in an embedded database?
- stringer 8y agoHow does it compare relatively to dgraph's Badger library?
- dmitrygr 8y agoHow well does that handle the storage disappearing halfway through a write? How well does it handle power being cut halfway through an update of some sort? How well does it handle some of the written blocks actually making it to the disc and others not? How about if the ones that made it were not the first or the last it issued to be written? (For predictable answers to these, and many other complex questions, when dealing with data you care about, use sqlite)
- krenoten 8y agoALICE showed that's not always true with sqlite. Sled is being built with an extreme bias toward reliability over features, but as the readme says, it has some time to go before reaching maturity. The tests are quite good at finding new issues and deterministically replaying them, so you can help bake it in by mining bugs using the default test suite and help it get there.
- toolslive 8y agoMost database designers assume that a power failure will only affect writes that are pending. Alas, for SSDs and NVMEs that's not always true. A power failure can cause all kinds of corruption. Long story short: even append-only strategies will not save you. https://www.usenix.org/system/files/conference/fast13/fast13-final80.pdf https://www.usenix.org/system/files/conference/fast13/fast13...
- spacenick88 8y agoThat paper sounds like problems that can not be worked around in software and need hardware fixes?
- toolslive 8y agoYou could work around it (erasure coding, fe), if you had insights on the failure mode specifics, but it's vendor specific and vendors are not exactly forthcoming. So the only thing you can do is add checksumming schemes that allow you to detect you have been affected. The paper is from 2013, so the situation might have improved meanwhile (I wouldn't put any money on it)
- dingo_bat 8y agoDoes rust compile reliably to embedded targets yet? Last time I checked there were a lot of problems with armv5.
- smilekzs 8y agoquoting catwell: > I think they mean "embedded" as in "embedded in a program", not "targeting embedded hardware".
- steveklabnik 8y agoI'm not sure about ARMv5, but v7 and v8-A are Tier 1 for Firefox, so we get a pretty decent smoke test out of them. The thing about embedded is that it's quite diverse, so speaking about it in broad terms is tough. It goes from "lol nope" to "barely works" to "pretty decent" to "great", depending on which thing you're talking about. We have a whole working group this year working on embedded.
- Nullabillity 8y agoI've been using embedded Rust for my thesis project (ARMv6-M), and it's been a pleasure.
- nixpulvis 8y agoIt's been a focus of the team, and been getting better. With efforts being made to move some external tools into the mainline toolchains. Last I checked it was possible, but still a bit cumbersome. P.S. Really looking forward to writing Rust on AVRs.
- dajonker 8y agoI see that MVCC is a planned feature. Why would you need MVCC for an embedded database? It seems like unnecessary overhead that conflicts with the performance goals.
- krenoten 8y agoIt's not needed for the single-key atomic record store, which is the sled bwtree index that is the current highest level module. MVCC is implemented in most popular embedded DBs because it is an effective way to manage mixed workloads that seek to read snapshots of the entire database at a single point in time as well as not blocking writes as this happens. This functionality is desirable for transactions that support mixed workloads. That's why I'm building it for a higher level module. This is a collection of modules that let you choose the abstraction and associated complexity that you want. There's also a modular pagecache that is totally decoupled and reusable for your own database experiments.
- eternalban 8y agoMVCC is a viable general purpose approach to deal with concurrency. It is the mother of all DAGs.
- catwell 8y agoI think they mean "embedded" as in "embedded in a program", not "targeting embedded hardware". In particular the README states they intend to use it in https://github.com/spacejam/rasputin https://github.com/spacejam/rasputin. This is a use case similar to LMDB, for instance, which has MVCC.
- EugeneOZ 8y agoIn Rust abstractions are free.
- fooker 8y agoIf that was so, there wouldn't have been a need for the unsafe mode.
- jsnell 8y agoOther people have had trouble wringing competitive performance out of Bw-Trees despite heroic optimization efforts [0]. Why is this implementation going to beat other index structures with just a bit of tuning? [0] https://news.ycombinator.com/item?id=17041616 https://news.ycombinator.com/item?id=17041616
- krenoten 8y agoIt might not. But the critiques of bw trees in terms of performance that I've seen have not had compelling data in terms of things that matter outside of academia or benchmarking shootouts, like write or space amplification. The bw tree is a cheap thing to abandon after I implement a persistent ART and measure it though. The bwtree is only like 1k of rust on top of the modular pagecache, which is the real heart of the system.
- lifepillar 8y agoI am not familiar with bw-trees, but are you aware of this paper from this year’s SIGMOD? https://db.cs.cmu.edu/papers/2018/mod342-wangA.pdf https://db.cs.cmu.edu/papers/2018/mod342-wangA.pdf Moving to ARTs (possibly combined with B+ trees) might be a smart choice after all. Edit: sorry, I missed that grand-parent has cited the same paper already.
- arthursilva 8y agoThat paper is pretty good but it's comparing bw-tree with much simpler in memory data structures. I think bw-tress might work specially well for fast-disk storage.
- krenoten 8y agoThis was my interpretation as well. I'm going to compare a disk-backed bwtree with a disk-backed ART, both backed by the same pagecache, and maybe end up with an ART that scatters partial pages on disk, bwtree style. But I need to measure apples to apples on the metrics that matter for storage first. The pagecache is where most of the complexity is in my implementation, and it makes building different kinds of persistent structures on top of it pretty easy. docs.rs/pagecache
- samuell 8y agoNice with more developments in embedded databases! Would be interesting with a comparison with mentat [1]. Do you plan to support any query langauge such as datalog, like mentat? [1] https://github.com/mozilla/mentat https://github.com/mozilla/mentat
- krenoten 8y agoYeah, I'm curious about using sled as a more ssd friendly storage engine for mentat. I'm just starting to experiment with datalog implementations, but I think by having harmony between the storage engine, query language, and hardware properties we can make a really compelling stateful systems. If this is something that interests you, I'd love to work with more people on this.
- jstewartmobile 8y agoI'll take Richard Hipp's C code over just about anyone's Rust code any day of the week. The man is a national treasure!
- judofyr 8y agoI'll also take any project that has thousands (if not millions?) of man-hours together with an extremely solid test suite. That doesn't mean other projects aren't worth exploring or won't be useful.
- tzahola 8y agoBut C is very unsafe! /s
- jstewartmobile 8y agosorry you got strikefrced, but love your other comments
- IshKebab 8y agoIt says it is alpha about 5 times in the readme. What gave you the impression that this was already more robust than SQLite?
- jeffdavis 8y ago"don't wake up operators. bring reliability techniques from academia into real-world practice." What does he mean here? Don't wake up people with operational problems? Or does "wake up" refer to a scheduling strategy? Either way, what are these techniques?
- erwan 8y agoI took that as a reference to pager duty.
- krenoten 8y agoIt means pay more attention to reliability than pop infrastructure and internet companies (who can offset poor reliability with human attention or intentionally deprioritize it to sell more support contacts) tend to put into these things. Specifically, exhaustive concurrency testing of lock-free algorithm interleavings via ptrace driven scheduling, model-based testing in combination with fault injection, ALICE-style file correctness testing, and for the various distributed modules that sit on top, network simulation combined with lineage driven fault injection. This is all very much a work in progress, and I'd love to work with more people on it!