4 ms·
I don't get it. Is this just a faster Cassandra? What's their competitive niche?
by rosslazer 10y ago
I don't get it. Is this just a faster Cassandra? What's their competitive niche?
- nopzor 10y agoA faster, well engineered, drop in replacement for Cassandra is a pretty competitive niche in of itself, don't you think?
- rosslazer 10y agoI guess but I don't really know. Lots of people have Cassandra in production. I suppose there is a certain segment of that market that needs a super high performance version, but is it really that big?
- manigandham 10y agoCassandra is supposed to be a high-performance database, except it has some fundamental issues that keep it from being what it can be. ScyllaDB fixes those fundamental issues so anyone using cassandra can benefit from using scylla instead.
- penberg 10y agoFor many users, that "super high performance" really just means lower request processing latency, which is a very desirable feature to have in various market segments. Look at the various benchmarks available (or run the benchmarks yourself) and you will see that Scylla has a very consistent, low latency that's a direct result of its non-blocking, shared-nothing architecture and implementation in a non-manage language, which gives us more control.
- bkeroack 10y agoAn order of magnitude faster Cassandra with no GC pauses (because it's C++). It's a literal drop in replacement. Even nodetool is compatible.
- nemothekid 10y agoTheir marketing points are really understood and relatable to people who know what a pain operating Cassandra can be. The two biggest pain points IMO are read-repair and compaction - both of which must be run periodically and consume a huge amount of resources. Read Repair is especially a pain because (1) it must be run periodically or you risk losing data (2) it takes forever to complete in some deploys (ex - you must run read-repair in a certain (user-tunable) timeframe, the default of which is every 10 days - I have tables that take 7 days to complete a read-repair, meaning I have repairs pretty much running 24/7) and (3) there are no/few operational tools to manage read repair. The low-tech way is to write a cron job on every node - and even then there is no way to measure progress or detect if a job failed/completed without grepping logs - it's so bad that Spotify wrote a open source tool to manage it. The solution has been to just buy more nodes (if you don't want long repairs, store less than 1TB of data per node) and faster disks. Read Repair maintenance is probably the only thing I hate about Cassandra - and seeing benchmarks that Scylla does these operations on the order of minutes rather than hours is attractive enough for most people (I don't think most deploys are even coming close to the benchmarked txn/s in real-world workloads, for both databases). Both compaction and repair tend to be CPU intensive (both work by essentially reading a ton of data), so I'd imagine the move to C++ and the core-per-thread design is more efficient. In short, the operational efficiency is far more attractive even if you aren't pushing a trillion writes/sec. I've been thinking about testing Scylla for a while, but unfortunately they don't support the features we support, and while our Cassandra deployment is a rather comparatively large cost, there are enough things on my plate right now where trading my current set of evils for other unknown ones isn't very attractive. See this post by Discord App - https://blog.discordapp.com/how-discord-stores-billions-of-messages-7fa6ec7ee4c7#.8ybap8c6o https://blog.discordapp.com/how-discord-stores-billions-of-m... - where they are mentioning moving to Scylla from Cassandra for similar reasons. Performance is fine, but repair efficiency is more of the driving factor. I'd also add that Cassandra advertises itself as a relatively high performance database for distributed workloads. If something like a faster Cassandra doesn't entice you, chances are you'd be better served by something like Postgres anyways.
- Nate75Sanders 10y agoWhat you keep calling "read repair" is actually just "repair" or the longer phrase "anti-entropy repair". "Read repair" is something different.