5 ms·
The "reduced latencies" really reduces down to performance, where tps was a first order proxy for performance. The ease of maintenance, similarly, is sold as e
by jjirsa 5y ago
The "reduced latencies" really reduces down to performance, where tps was a first order proxy for performance.
The ease of maintenance, similarly, is sold as easier due to reduced node count, which is perhaps an extension of performance but probably misunderstands (or ignores) that most people running large cassandra clusters have tooling that parallelizes most maintenance anyway, so the reduction in effort is sorta not that important in real life (if anything, having more machines gives you better blast radius behavior, consolidation onto fewer exposes you to larger percentages of loss/failure when there's inevitably a problem with the fewer, larger machines).
The real comparison, though, is missing in that link, because the real comparison is not performance. It's license. Nobody is running AGPL in prod unless they have zero IP worth protecting, so it's ultimately comparing OSS to proprietary.
(And similar disclosure: cassandra committer)
- thekozmo 5y agoHmm, AGPL is proprietary? This isn't aligned with the OSI.
- jjirsa 5y agoYour options for scylla are AGPL or pay scylla. Nobody risks AGPL in prod so people running it are probably paying scylla.
- throwdbaaway 5y ago> The "reduced latencies" really reduces down to performance, where tps was a first order proxy for performance. Thank you so much for this comment. I have come to a similar conclusion recently -- as long as the P100 latency is acceptable (e.g. all requests are served within 1 seconds), the only thing that matters is the TPS. Back in 2016, I stumbled upon "How NOT to Measure Latency" by Gil Tene, and it really opened my eyes, especially the part about coordinated omission from benchmark tools. Of course, after a few days, I also learnt that Gil was behind the Azul Zing JVM with the pauseless C4 garbage collector, and thus it makes sense for him to emphasize on measuring latency correctly. At around the same time, GCP also boasted about having "consistent single-digit millisecond latency" for its BigTable offering, with Cassandra being the obvious target to attack. I was sold. Then came Scylla, with a focus on maintaining low tail latency while running on fewer larger machines. I tested one of the early version with cassandra-stress, and the result was worse than cassandra. But I continued to follow Scylla blog posts with great interest. Recently, I saw https://www.p99conf.io/ https://www.p99conf.io/, and something "clicked" when I read that it was sponsored by Scylla. Suddenly the hype of P99 seems to be wearing off. The video from Gil above is still correct, but I think it is applicable more to exceptional cases, e.g. - the JVM enters full GC and do no meaningful work for seconds/minutes - the InnoDB engine stalls for seconds/minutes due to a million different reasons For normal cases, I haven't found the P95/P99/P99.9 metrics to be that useful. Instead, something like PostgreSQL/MySQL slow log threshold, where anything that exceeds a soft P100 target is logged, seems to be more useful. Back to the article. If we only concern about TPS, under the "real-life" workload with Gaussian distribution, Scylla beats Cassandra by 2X. So that's it, 2X better performance on one hand, Apache vs. AGPL on the other. (Of course that's an oversimplification. Scylla's shard-per-core architecture also allows it to avoid some silly single-threaded bottlenecks in Cassandra, but nobody wants to talk about those.)
- PeterCorless 5y agoIf you are a fan of Gil's work, and the entire phenomenon of coordinated omission, this is right up your alley. https://www.scylladb.com/2021/04/22/on-coordinated-omission/ https://www.scylladb.com/2021/04/22/on-coordinated-omission/ Also, yes, P99 CONF (https://p99conf.io https://p99conf.io) is sponsored by ScyllaDB, but we're very glad to have speakers from other NoSQL vendors like Couchbase and Redis, as well as folks from across the industry — streaming systems like Kafka (Confluent), Pulsar (Splunk), Redpanda (Vectorized). Plus storage systems like Ceph, Crimson (both Redhat) and Lightbits LightOS. For the P99 CONF I recently conducted a quick poll of what people consider "acceptable" P99s. For some people it's <100 µseconds. For others, its <1 ms. And for some it goes all the way up to 1 second. But "acceptable" is use-case specific. An in-memory database or cache will have a very different expectation than someone writing data to SSD or even today, HDD. 33.3% of respondents expected <1 ms. 37.5% were okay with 1-<10 ms 20.8% were okay with 10-<100 ms Only 8.3% were okay with 100ms - 1 second (The <100 µsec was a "write-in" comment. But I am sure if I had included it, and we had a broader sample poll, you'd see it as a prevalent and vocal minority.) https://twitter.com/P99CONF/status/1440197863875629057?s=20 https://twitter.com/P99CONF/status/1440197863875629057?s=20
- throwdbaaway 5y agoThe result of your poll is exactly my fear -- p99 seems to be at the peak of the hype cycle. I am totally a fan of Gil's work, but I think he over-dramatized the impact of tail latency, with the rather extreme example of "If a typical user session involves 5 pages loads, averaging 40 resources per page" from his talk. Regarding coordinated omission, I am not totally convinced that an open-model system is what happens in production, and I am actually fine with a closed-model system when running benchmark, as long as p100 is acceptable. My goal is to achieve the maximum TPS without stalling, and I don't really care that much whether p99 is 10ms or 100ms.
- deleted 5y ago[deleted]
- jjirsa 5y agoI don’t think anyone avoiding talking about single threaded bottlenecks. The main one that matters is on sstable streaming during bootstrap, and the zero copy streaming was directly to address that. There’s not really any single threaded pieces in cassandra that matter beyond that? If anything the opposite is true - thread pools everywhere leading to context switching and data copying for no reason. On a deeper note, it probably says something about the user base that people who don’t run it in prod think the perf is really bad and yet there are very few people submitting perf improvement patches. Maybe most of the power users aren’t that worried about perf because they either know how to tune a JVM or they happen to size their clusters based on bytes on disk?