5 ms·
You call it "high performance" and provide no benchmarks?
by bufferoverflow 4y ago
You call it "high performance" and provide no benchmarks?
- dilyevsky 4y agoPaper has it https://www.usenix.org/system/files/atc20-conway.pdf https://www.usenix.org/system/files/atc20-conway.pdf but yeah if you check out list of limitations looks more like a research proj at this stage. Pretty interesting architecture overall though
- bufferoverflow 4y agoThe numbers look very good actually. I don't care if it's a research project. If it doesn't crash, doesn't corrupt data, and delivers performance, it's useful. I'd want to see performance against Redis and KeyDB.
- dilyevsky 4y agoWell you should read the limitations… I think they are actually cheating by not calling fsync at all which makes writes not durable. This is different in rocks/pebble and friends. > I'd want to see performance against Redis and KeyDB. I think this is apples to oranges comparison as neither of these provide durability by default and if you enable it redis had terrible performance last I checked + redis needs to fit a whole dataset in memory
- ajhconway 4y agoHi, research lead for SplinterDB here. SplinterDB does make all writes durable and in fact has its own user-level cache which generally performs writes directly to disk (using O_DIRECT for example). Like RocksDB's default behavior (no fsyncs on the log), it does not immediately sync writes to its log when they happen. It waits to sync in batches, so that writes may not be immediately durable, but logging is more efficient. This is a slightly stronger default durability guarantee, and we intend to make this configurable.
- dilyevsky 4y agoI missed the use of direct io and the comment about fsync threw me off, thanks. Very impressive then!
- jlokier 4y agoO_DIRECT doesn't provide power-cut durability on storage devices with write cache. Recently written and acknowledged data can still be lost on a power cut. You still need fsync, fdatasync or equivalent after an O_DIRECT write, to tell the storage device to commit its write cache to the non-volatile layer. (And last time I looked, I think some filesystems even incorrectly failed to flush the device write cache on fsync after O_DIRECT writes because of no dirty page states.)
- dilyevsky 4y agoThere’s a ton devices on the market that would lie to you too saying caches are flushed while they aint. If you really want that data to be there better use “server grade” hw with power loss protection
- otterley 4y agoI’m a little confused. If you don’t ensure data is committed to storage (log or otherwise) before acking the write request, how can you call it durable? If it’s not truly 100% durable by default, it’s best not to suggest that it is. Experience says people will use the default settings and then become very cross if they lose data. It undermines trust and is harmful to reputation.
- ajhconway 4y agoWith many workloads, there's a tradeoff between the granularity of durability and the overall performance. If a workload has many small writes (some of our product workloads do), then syncing each write can cause write amplification and massively affect overall throughput and latency. Suppose I do a 100B write, this causes a 4KiB page write to sync, which is 40x write amp. Suddenly a 2GiB/sec SSD can effectively only write 50MiB/sec. Similarly, the per-write latency goes from <5us to 10us (with the fastest Optane SSDs) or 150us (with flash SSDs). So storage systems tend to offer a range of durability guarantees. Some systems have a special sync operation for applications to ensure that all writes are durable. RocksDB offers a fairly weak guarantee by default too, writing to the write-ahead-log (WAL), but not performing fsyncs (https://github.com/facebook/rocksdb/wiki/WAL-Performance https://github.com/facebook/rocksdb/wiki/WAL-Performance). They make a similar write amplification argument too (https://github.com/facebook/rocksdb/wiki/WAL-Performance#write-amplification https://github.com/facebook/rocksdb/wiki/WAL-Performance#wri...).
- tyingq 4y agoAh, that's helpful, and explains why it exists: "Three novel ideas contribute to the high performance of SplinterDB: the STB-tree, a new compaction policy that exposes more concurrency, and a concurrent memtable and user-level cache that removes scalability bottlenecks. All three components are designed to enable the CPU to drive high IOPS without wasting cycles." "At the heart of SplinterDB is the STB-tree, a novel data structure that combines ideas from log-structured merge tree and B-trees. The STB-tree adapts the idea of size-tiering (also known as fragmentation) from key-value stores such as Cassandra and PebblesDB and applies them to B-trees to reduce write amplification by reducing the number of times a data item is re-written during compaction."
- stingraycharles 4y agoYeah I would appreciate a benchmark against its main alternative, rocksdb. I know benchmark are typically manufactured and not too representative for real world load, but at least a ballpark figure would be nice to know what we’re talking about here. Their main website is at https://splinterdb.org/ https://splinterdb.org/ by the way, for those interested. Also no benchmarks there. :)
- ridruejo 4y agoThe paper referenced in the other comment includes a benchmark against RocksDB https://news.ycombinator.com/item?id=31515765 https://news.ycombinator.com/item?id=31515765
- deleted 4y ago[deleted]