8 ms·
TurboKV: Insanely fast Rust key-value store
- netappblackbox 1mo ago[flagged]
- nine_k 1mo agoI suppose the insane speed is due to this: > TurboKV's persisted Bloom-filter format uses hardware AES. Also, built-in LZ4 compression. I would expect SIMD to be used for scans.
- rgbimbochamp 1mo agoThose help but the main write speed gain is the WAL, that uses preallocated mmap segments to avoid a write(2) per durable mutation while preserving crash recovery. AES hashing mainly helps Bloom filter point lookups and LZ4 mainly helps SSTable I/O. Scans benefit indirectly, but don’t yet use a custom SIMD merge loop.
- dangoodmanUT 1mo agoIirc that’s how badger handles the WAL as well
- bestouff 1mo agoI said elsewhere this doesn't survive a power loss.
- taneq 1mo agoWhile it’s important to make this explicit, at what point do we just assume a high-reliability UPS is table stakes? Of course, if you need SIL2 type reliability then you need to assume any given hardware component can spontaneously combust and become a total loss, at which point the data loss caused by a power cut is a rounding error.
- whilenot-dev 1mo agoWhat's got this to do with a UPS? Not having a UPS is an external threat on the reliability of the power grid. Doing a hard shutdown or tripping over power cords seem much likelier local scenarios than any spontaneous combustion of hardware components.
- toast0 1mo ago> While it’s important to make this explicit, at what point do we just assume a high-reliability UPS is table stakes? Several years after they become commercially available? My experience with small UPSes is they tend to cook the batteries and you don't find out until they switch the load and the battery doesn't hold up. Large facility UPSes tend to do better, but automatic transfer switches have a tendancy to fail ocassionally. If you're hosted in many locations, it's not unusual to have a couple ATS failures per decade. All that said, unexpected power loss is certainly one reason that writes may be lost, but OSes crash too. Disk firmware can also crash, but if thst bricks the disk, writes in progress don't really matter. Sometimes cabling fails. Or you get a uncorrectable ECC error (which will typically cause an OS panic... unless you're running a very fancy OS, but if it's in dirty disk backed page, even a fancy OS wouldn't save you) Plenty of applications don't need or want to pay the cost for full commit to disk, but calling something durable when it's not committed to disk is inaccurate. And that's before we get into the whole thing where the OS and the disk like to return success when things haven't quite finished.
- imtringued 1mo agoNot sure how you managed to do it but you got it completely backwards. If someone demands that the database should use fsync and only respond with success once the write finished, it is not some arbitrarily high reliability demand that needs to be implemented using reliable hardware. In fact, the entire point of implementing the power loss protection in software is so that you don't need perfectly reliable hardware. The power loss event turns into a downtime event which is often completely acceptable. The requirement to have infallible hardware only emerged because the software refused to do its job. Infallible hardware is not a requirement decided by the user, it's a requirement decided by the developer of TurboKV to intentionally restrict his software to exclusively operate in a reliable hardware environment. The fact that the user specified durability of the KV store during power loss does not make the user obsessed over hardware reliability, the software shifted the burden onto the hardware and forced the user to deal with this mess. I don't know how exactly TurboKV works so let's talk about a hypothetical software instead. Let's say the software cannot survive a power loss event and just corrupts the database. If the user wants to operate the software, he is forced by the software to operate it in an infallible environment where power loss can never occur. Based on how the software was designed, power loss is a catastrophic event. The SIL2 type reliability you're talking about only makes sense in contexts with catastrophic events. So how it went is that the user made a reasonable demand with bounded reliability: "please survive power loss with durable writes" and the author says, sure just run the software on a SIL2 type reliability hardware environment. It's not the user who blew up the hardware requirements.
- rgbimbochamp 1mo agoparanoid() does survive power loss.
- vlovich123 1mo agoAnd has performance I believe worse than rocksdb
- rgbimbochamp 1mo agoRocksDB’s default is WriteOptions.sync = false. It can lose recent acknowledged writes after power loss, much like TurboKV Durable. Most folks use that unknowingly.
- deleted 1mo ago[deleted]
- haberman 1mo agoI assume this is for hashing. I've seen several hashing algorithms turn to hardware AES instructions before, but I haven't seen any evidence that this technique outperforms state-of-the-art hashes like RapidHash (https://github.com/Nicoshev/rapidhash https://github.com/Nicoshev/rapidhash) in either quality or speed.
- Sesse__ 1mo agoFor any complex system, there's never one single trick or design choice that makes it fast. It's always a large amount of engineering (or exaggerations, of course).
- itemize123 1mo agothat's optimizing a pretty fast already portion of code. unlikely it's the difference maker.
- dangoodmanUT 1mo ago> DbOptions::durable() > Appended to the WAL without a per-write sync So… it’s not durable? Durable doesn’t mean “survives a process restart”, it means “durably saved to persistent storage”. For example, this “durable” mode wouldn’t survive power loss.
- stingraycharles 1mo agoYeah this should be benchmarked against other systems that have flush() disabled. mmap is nice but it doesn’t support durable semantics in the way that we usually mean with databases. if a write is acknowledged it should not be forgotten, which is not what this is.
- rgbimbochamp 1mo agoYou're right, that mode provides process crash recovery, not power-loss durability. The benchmark compares it against fjall’s equivalent buffered-WAL mode.
- a2ff6eeb0 1mo agoIf that's your design constraint, couldn't you speed it up by getting rid of the WAL?
- oneshadab 1mo agoYou'd lose durability against process crashes. If your system has a reasonable tolerance for power failure (multi-az multi-cloud), this can provide much better throughput
- t098i3 1mo agoIndeed, a common enough pattern for etcd is to run it backed by a RAMdisk and have multi-az availability + periodic backups + tolerance at a business level to be OK losing some recent data.
- 1mo ago
- boguscoder 1mo agoEmbedded could also mean no_std, which this is absolutely not. Still cool though
- LoganDark 1mo agoYep, the correct term here is "embeddable", not embedded.
- AtlasBarfed 1mo agoAphyr or gtfo
- paulsutter 1mo agoOops you built a database!
- yeasin-arafat 1mo ago[flagged]
- karen4830 1mo ago[flagged]
- medv 1mo agoEvery programmer eventually creates own db: https://github.com/antonmedv/medb https://github.com/antonmedv/medb
- otabdeveloper4 1mo agoA file-backed hashtable isn't really a DB.
- riffraff 1mo agoWell, dbm (database manager) is basically that and has been called that for almost 50 years
- otabdeveloper4 1mo agodbm isn't a database either, and people who call it that are simply wrong.
- shakna 1mo agoSo... Databases weren't invented, until after the name was in common use?
- insanitybit 1mo agoIf someone asked me what a DB is, I'd probably start the conversation with the exact description "a file-backed map, like a hashmap".
- mikedelago 1mo agoWhy not?
- otabdeveloper4 1mo ago"Database" implies there's some sort of data aggregation or query primitive under the hood.
- rollulus 1mo agoNow that “blazing fast in Rust” has become a meme, is “insanely” the next thing?
- carlos-menezes 1mo agoInsanely/ridiculously.
- ivolimmen 1mo agoLudacris
- Tepix 1mo agoIs an atomic get+delete operation planned?
- vlovich123 1mo agoIt’s usually very difficult in a KV db to have an efficient operation that returns the item deleted. You’d need a transaction API to do it reliably. The challenge is concurrent writes are impossible to serialize against without transactions. This DB doesn’t have a transaction API.
- mattrighetti 1mo agoI like the fact that the first commits were about the logo, important things first :D
- Surac 1mo agoFast compared to what?
- DarmokTanagra 1mo agoOff topic, but why is tokio still independent of the rust async runtime? It seems pretty ubiquitous yet not a part of the core rust libs.
- insanitybit 1mo agoThe same reason as ever. Not everyone wants to use the same runtime.
- DarmokTanagra 1mo agoSure, but now if my rust application isn't using tokio I have to include it as a new dependency because the author of this lib decided to use it as his async runtime? I'm not trying to be pedantic, but this split over async runtimes was what originally turned me off of rust years ago and it still seems to be an issue.
- timschmidt 1mo agoI'm using async on esp32 with embassy, for instance, where there's simply no room and no need for a runtime as complex as tokio.
- tinco 1mo agoIt's not pedantry, you're trying to find something in Rust that's simply not there. If this sort of thing turns you off on Rust you should be looking at a programming language with other priorities.
- derriz 1mo agoI’ve had the same reaction. We’ve seen this play out in other language eco-systems (Java - JAXP, javax.validation, JPA, etc all ended up with a de facto single implementation) and the idea of pluggable implementations sounds appealing but rarely pays off. The price that Rust paid for this abstraction - which in fact ended up not being useful as tokio is the only reasonable choice - was high in terms of requiring ugly (to my eyes) changes to its type system.
- dist-epoch 1mo agoHow does it compare to RocksDB?
- hnea3ekp5i 1mo agoUnderrated take
- KAdot 1mo agoThe benchmark is setup to test 80 MiB dataset on a machine with 32 GiB RAM, which doesn't represent a typical database workload. How does the key-value store perform on larger than RAM datasets? MMAP is fast when the dataset fits in memory, but it can slow to a crawl when it doesn't, especially if the workload is mostly random point lookups.
- hakesson 1mo ago[dead]
- pjmlp 1mo agoMost of these "insanely fast" projects kind of feel like people rediscovering compiled languages after the scripting languages dark ages. Insanely fast was making 8 bit games possible at all.
- xtracto 1mo agoMore or less. I think the use of CPU specific instructions can make an compiled program different than "the rest". Although nowadays some compilers are clever enough to do better than manual optimization.
- pjmlp 1mo agoMichael Abrash wrote about such optimizations on the late 90's regarding Pentium versus its predecessors, and then everyone started using Python, Ruby, whatever for full stack applications, beyond plain OS scripting tasks.
- 5ersi 1mo ago[dead]
- mellosouls 1mo agoEditorialized? Where does the repo claim to be "Insanely fast"? Has there been a change to the Readme since submission? If not, please just link the repo and its own title, no need for hype.