9 ms·
Ask HN: Why are there no open source NVMe-native key value stores in 2023?
Hi HN,
NVMe disks, when addressed natively in userland, offer massive performance improvements compared to other forms of persistent storage. However, in spite of the existence of projects like SPDK and SplinterDB, there don't seem to be any open source, non-embedded key value stores or DBs out in the wild yet.
Why do you think that is? Are there possibly other projects out there that I'm not familiar with?
- jrmysterylord 3y ago[dead]
- diggan 3y agoI don't remember exactly why I have any of them saved, but these are some experimental data stores that seems to be fitting what you're looking for somewhat: - https://github.com/DataManagementLab/ScaleStore https://github.com/DataManagementLab/ScaleStore - "A Fast and Cost-Efficient Storage Engine using DRAM, NVMe, and RDMA" - https://github.com/unum-cloud/udisk https://github.com/unum-cloud/udisk (https://github.com/unum-cloud/ustore https://github.com/unum-cloud/ustore) - "The fastest ACID-transactional persisted Key-Value store designed for NVMe block-devices with GPU-acceleration and SPDK to bypass the Linux kernel." - https://github.com/capsuleman/ssd-nvme-database https://github.com/capsuleman/ssd-nvme-database - "Columnar database on SSD NVMe"
- ashvardanian 3y agoHey, thanks for the mention! UDisk, however, hasn't been open-sourced yet. Still considering it :)
- geek_at 3y agoyou could also configure Redis to transact everything to disk and choose nvme as the target
- jitl 3y agoThat would save via file system, not bypass the kernel to access the NVMe drive directly from user space. NVMe drives themself have a bunch of features that make them amenable to K/V storage directly. Good overview: https://www.mydistributed.systems/2020/07/towards-building-high-performance-scale.html?m=1 https://www.mydistributed.systems/2020/07/towards-building-h...
- PaulHoule 3y agoSee https://www.snia.org/educational-library/key-value-standardized-2020 https://www.snia.org/educational-library/key-value-standardi... for some description of the special command set to get an nvme drive to natively work as a key-value store. Also https://www.snia.org/sites/default/files/ESF/Key-Value-Storage-Standard-Final.pdf https://www.snia.org/sites/default/files/ESF/Key-Value-Stora...
- mycall 3y agoHow is that implemented? Btree, hashtable?
- anticensor 3y agoImplementation-defined. The API resembles a map access, though.
- aftbit 3y agoInteresting, never heard of this before! Do you have any other resources to share? How can I play with this today?
- laurencerowe 3y agoHow do you tell which NVMe drive models support the KV API? Is this something that you can experiment with on a consumer drive or do you need specific enterprise ssd models? Samsung's uNVMe evaluation guide (from 2019) device support section just states: Guide Version: uNVMe2.0 SDK Evaluation Guide ver 1.2 Supported Product(s): NVMe SSD (Block/KV) Interface(s): NVMe 1.2 https://github.com/OpenMPDK/uNVMe/blob/master/doc/uNVMe2.0_SDK_Evaluation_Guide_v1.2.pdf https://github.com/OpenMPDK/uNVMe/blob/master/doc/uNVMe2.0_S... I can't find detailed spec sheets detailing which NVMe command sets are supported even for their enterprise drives.
- laurencerowe 3y agoI reached out to Samsung support to ask. After being sent from one department to another and receiving some very clearly incorrect advice from their sales support they eventually sent me to an online form for the memory department. Am still waiting for a response a week later.
- jamesblonde 3y agoRonDB is open-source and supports on-disk data on NVMe disks. http://mikaelronstrom.blogspot.com/2022/04/variable-sized-disk-rows-in-rondb.html http://mikaelronstrom.blogspot.com/2022/04/variable-sized-di...
- bestouff 3y agoNaive question: are there really expected gains to address natively an NVMe disk wrt using a regular key-value database on a filesystem ?
- creshal 3y agoLatency ought to be much better, since you're skipping several abstraction layers in the kernel. But that's about it. And the latency is still worse than in-memory solutions. Between that and the non-trivial effort needed to make this work in any sort of cloud setup (be it self-hosted k8s or AWS), it's a hard sell. If I really need latency above all, AWS gives me instances with 24TB RAM, and if I don't… why not just use existing kv-stores and accept the couple of ns extra latency?
- klodolph 3y agoAgreed. The classic reason is when you have latency needs, but your data set is large enough that RAM is cost-prohibitive, and random-access enough that disk won’t work. The cost savings from switching to NVMe have to justify the higher NRE cost, and simultaneously, you have to be sensitive to latency.
- creshal 3y agoIndividual NVMe drives are also rather small – the biggest I can find is 30TB, which is still more than what AWS offers me as RAM, but not much. Once you start adding custom algorithms to spread your data over multiple "raw" NVMe drives to get more capacity, the latency gap between your custom solution and existing, well-optimized file system stacks starts to erode. Might as well stick to existing kv stores on ZFS or something, rather than roll your own project that might be able to beat it, maybe.
- adgjlsfhk1 3y agoWhile you can get 24TB ram, there is a pretty big cost difference. 2 TB of ram costs roughly $10000 compared to $130 for NVME storage (or $230 for 12 TB of a good hard drive). Sure the NVME is ~3.5x more expensive, but the latency will be dramatically lower and the throughput will be dramatically higher. Sure you can build a 24 TB ram system, but at that point the cost of the server will be entirely the ram. The reason for NVME based storage at this point is that at only ~3.5x the cost of a hard drive, you can switch all your storage over and as long as you don't need tons of storage (i.e. less than 100TB), the SSDs will be a minority of the cost of the system.
- delfinom 3y agohttps://github.com/OpenMPDK/KVRocks https://github.com/OpenMPDK/KVRocks Given however, that most of the world has shifted to VMs, I don't think KV storage is accessible for that reason alone because the disks are often split out to multiple users. So the overall demand for this would be low.
- londons_explore 3y agoNVME's allow namespaces to be made - effectively letting multiple users all share an NVME device without interfering with each other.
- moondev 3y agoA note for those unaware, consumer grade NVME devices (basically all m.2 formfactor drives from my experience) only support a single namespace. If you want to explore creating multiple namespaces you will need an enterprise grade u.2 drive. Some u.2 drives even support thin provisioning, like how a hypervisor treats a sparse disk file but for physical hardware.
- zupa-hu 3y agoIs there any performance gain over writing append-only data to a file? I mean, using a merkle tree or something like that to make sense of the underlying data.
- dboreham 3y agoWriting to append-only files is a terrible idea if you want to query quickly. (yes it's fashionable, but it's still terrible for random read performance)
- zupa-hu 3y agoCare to elaborate? How is reading from an append-only file backed by a memory indexed DB slower compared to either 1) a mutated file, or 2) either append-only or mutated raw NVMe disk storage? I mean, what's the trick NVMe can do to be meaningfully faster?
- LAC-Tech 3y agoYour views are intriguing and I wish to subscribe your news letter. But seriously, I've been thinking about an append-only files + memory indexed DB for the past couple of weeks - any prior art or links or papers or anything, lay it on me.
- zupa-hu 3y agoI've been using it in production for 8 years in Boomla. It's closed source though. I haven't found any prior art myself, so just went from first principles. Take a look at the data structure of Git for inspiration. (Merkle tree) Write speed wasn't my primary motivation though. I wanted a data storage solution that is hard to fuck up. Hard to beat append only in this regard. Plus everything is stored in merkle trees like in Git, so there is the added benefit of data integrity checks. Yes, bit rot is real, and I love to have a mechanism in place to detect and fix those.
- brightball 3y agoSolidCache and SolidQueue from Rails will be doing that when released. Otherwise though…you have the file system. Is that not enough?
- andruby 3y agoIs that discussion/implementation of nvme available somewhere in public? https://github.com/rails/solid_cache https://github.com/rails/solid_cache didn't include anything about NVME that I could find.
- andrenotgiant 3y agoI think the original question came up after the recent Rails keynote where they mention that, with NVMe speeds, disk is cheaper and almost as fast as memory, so Redis is not as vital. https://youtu.be/iqXjGiQ_D-A?t=2836 https://youtu.be/iqXjGiQ_D-A?t=2836 So Solid Cache and Solid Queue just use the database (MySQL), which uses NVMe. So now, in addition to: "You don't need a queue, just use Postgres/MySQL", we have "You don't need a cache, just use Postgres/MySQL"
- andruby 3y agoRight, that is cool, but unrelated to the OP of using NVME directly and bypassing the filesystem. Or does MySQL have a storage driver that talks directly on the NVME level? (I haven't used MySQL in more than a decade, mostly PostgerSQL now)
- telegpt 3y ago[dead]
- formerly_proven 3y agoThere's actually an NVMe command set which allows you to use the FTL directly as a K/V store. (This is limited to 16-byte keys [1] however, so it is not that useful and probably not implemented anywhere, my guess is Samsung looked at this for some hyperscaler, whipped up a prototype in their customer-specific firmware and the benefits were lesser than expected so it's dead now) [1] These slides claim up to 32 bytes, which would be a practically useful length: https://www.snia.org/sites/default/files/ESF/Key-Value-Storage-Standard-Final.pdf https://www.snia.org/sites/default/files/ESF/Key-Value-Stora... but the current revision of the standard only permits two 64-bit words as the key ("The maximum KV key size is 16 bytes"): https://nvmexpress.org/wp-content/uploads/NVM-Express-Key-Value-Command-Set-Specification-1.0c-2022.10.03-Ratified-1.pdf https://nvmexpress.org/wp-content/uploads/NVM-Express-Key-Va...
- londons_explore 3y agoPresumably there is some way to use the hash of the actual key as the key, and then store both key and value as data? 16 bytes is long enough that collisions will be super rare, and while you obviously need to write code to support that case, it should have no performance impact.
- deleted 3y ago[deleted]
- formerly_proven 3y agoA 32-byte key would allow using NVMe KV directly for content-addressed storage; many of those systems use 256-bit / 32-byte cryptographic hashes as keys. Notable exception would be git with 20-byte keys.
- londons_explore 3y agoYou can still do this... Just use the first 16 bytes as the key, and the 2nd 16 bytes as the start of the data.
- londons_explore 3y ago
- znpy 3y agoI often attended a presentation by some presales engineer from Aerospike and IIRC they're doing some nvme-in-userspace stuff.
- CubsFan1060 3y agoInteresting article here: https://grafana.com/blog/2023/08/23/how-we-scaled-grafana-cloud-logs-memcached-cluster-to-50tb-and-improved-reliability/ https://grafana.com/blog/2023/08/23/how-we-scaled-grafana-cl... Utilizing: https://memcached.org/blog/nvm-caching/,https://github.com/memcached/memcached/wiki/Extstore https://memcached.org/blog/nvm-caching/,https://github.com/m... TLDR; Grafana Cloud needed tons of Caching, and it was expensive. So they used extstore in memcache to hold most of it on NVMe disks. This massively reduced their costs.
- thskman 3y ago[dead]
- jiggawatts 3y agoNote that some cloud VM types expose entire NVMe drives as-is directly the guest operating system without hypervisors or other abstractions in the way. The Azure Lv3/Lsv3/Lav3/Lasv3 series all provide this capability, for example. Ref: https://learn.microsoft.com/en-us/azure/virtual-machines/lasv3-series https://learn.microsoft.com/en-us/azure/virtual-machines/las...
- rwmj 3y agoIs there not any danger of tenants rewriting the firmware on these drives, and surprising (or compromising) future tenants? AIUI this is the central reason why even "baremetal" cloud instances still have a minimal hypervisor between the tenant and the hardware.
- nixgeek 3y agoI’m not sure what makes you think an “minimal hypervisor” exists — Oracle Cloud Infrastructure doesn’t have a hypervisor of any sort between you and its .metal instance types. Don’t think Amazon EC2 does either.
- rwmj 3y agoAmazon have their own partitioning hypervisor for this purpose. It sits below any hypervisor that might be visible to the tenant.
- wmf 3y agoThe top clouds (AWS/Azure/Google) have custom firmware to solve this problem. Second-tier clouds probably don't so customers can reflash firmware.
- otterley 3y agoIf your second sentence is true -- and I hope it isn't! -- that would be a gaping security hole.
- javierhonduco 3y agoThere’s Kvrocks. It uses the Redis protocol and it’s built on RocksDB https://github.com/apache/kvrocks https://github.com/apache/kvrocks
- eatonphil 3y agoDoes RocksDB speak NVMe directly? > High-performance storage engines. There are a number of storage engines and key-value stores optimized for flash. RocksDB [36] is based on an LSM-Tree that is optimized for low write amplification (at the cost of higher read amplification). RocksDB was designed for flash storage, but at the time of SATA SSDs, and therefore cannot saturate large NVMe arrays. From this slightly tangent mention, I am guessing not. https://web.archive.org/web/20230624195551/https://www.vldb.org/pvldb/vol16/p2090-haas.pdf https://web.archive.org/web/20230624195551/https://www.vldb....
- ilyt 3y agoIt becomes complex when you want to support multiple NVMes Even more complex when you want to have any kind of redundancy, as you'd essentially need to build-in some kind of RAID-like into your database. Also few terabytes in RAID10 NVMes + PostgreSQL and something covers about 99% of companies needs for speed. So you're left with 1% needing that kind of speeds
- deleted 3y ago[deleted]
- gavinray 3y agoWhy do you mean by non-embedded? You might also be interested in xNVMe and the RocksDB/Ceph KV drivers: https://github.com/OpenMPDK/xNVMe https://github.com/OpenMPDK/xNVMe https://github.com/OpenMPDK/KVSSD https://github.com/OpenMPDK/KVSSD https://github.com/OpenMPDK/KVRocks https://github.com/OpenMPDK/KVRocks
- nphase 3y agoSuper helpful, thanks. What I mean is something akin to a single-node daemon with network capabilities. Something as basic as a memcached or redis type of interface to start.
- gavinray 3y agoI think there's actually a standard defined for networked KV API over NVMe, written by the SNIA (as others have mentioned) Though I'm not super knowledgeable about it. I think Redfish/Swordfish are maybe meant for this sort of thing: https://www.snia.org/forums/smi/swordfish https://www.snia.org/forums/smi/swordfish There's a video on NVMe and NVMe-oF management for instance: https://www.youtube.com/watch?v=56VoD_1iGIs&list=PLH_ag5Km-YUYEHj-8YEhmlA6z7Bml_3Kq&index=72 https://www.youtube.com/watch?v=56VoD_1iGIs&list=PLH_ag5Km-Y...
- threeseed 3y agoCrail [1] which is a distributed K/V store on top of NVMEoF. [1] https://craillabs.github.io https://craillabs.github.io
- otterley 3y agoBecause you haven't written it yet!
- rubiquity 3y agoI think it's mostly because while the internal parallelism of NVMe is fantastic our logical use of them is still largely sequential.
- caeril 3y ago> non-embedded key value stores or DBs out in the wild yet I like how you reference the performance benefits of NVMe direct addressing, but then immediately lament that you can't access these benefits across a SEVEN LAYER STACK OF ABSTRACTIONS. You can either lament the dearth of userland direct-addressable performant software, OR lament the dearth of convenient network APIs that thrash your cache lines and dramatically increase your access latency. You don't get to do both simultaneously. Embedded is a feature for performance-aware software, not a bug.
- espoal 3y agoI'm building one: https://github.com/yottaStore/yottaStore https://github.com/yottaStore/yottaStore
- altairprime 3y ago“Lazyweb, find me an NVMe key-value store” is how we phrased requests like this twenty years ago. Who could afford to develop and maintain such a niche thing, in today’s economy, without either a universal basic income or a “non-free” license to guarantee revenue?
- Already__Taken 3y agoA seaweedFS volume store sounds like a good candidate to split some of the performance volumes across the nvme queues. You're supposed to give it a whole disk to use anyway.
- nerpderp82 3y agoAerospike does direct NVME access. https://github.com/aerospike/aerospike-server/blob/master/cf/src/hardware.c#L83 https://github.com/aerospike/aerospike-server/blob/master/cf... There are other occurrences in the codebase, but that is the most prominent one.
- nerpderp82 3y agoEatonphil posted a link to this paper https://web.archive.org/web/20230624195551/https://www.vldb.org/pvldb/vol16/p2090-haas.pdf https://web.archive.org/web/20230624195551/https://www.vldb.... a couple hours after this post (zero comments [0]) > NVMe SSDs based on flash are cheap and offer high throughput. Combining several of these devices into a single server enables 10 million I/O operations per second or more. Our experiments show that existing out-of-memory database systems and storage engines achieve only a fraction of this performance. In this work, we demonstrate that it is possible to close the performance gap between hardware and software through an I/O optimized storage engine design. In a heavy out-of-memory setting, where the dataset is 10 times larger than main memory, our system can achieve more than 1 million TPC-C transactions per second. [0] https://news.ycombinator.com/item?id=37899886 https://news.ycombinator.com/item?id=37899886
- infamouscow 3y agoI work on a database that is a KV-store if you squint enough and we're taking advantage of NVMe. One thing they don't tell you about NVMe is you'll end up bottlenecked on CPU and memory bandwidth if you do it right. The problem is after eliminating all of the speed bumps in your IO pathway, you have a vertical performance mountain face to climb. People are just starting to run into these problems, so it's hard to say what the future holds. It's all very exciting.