3 ms·
In part it's also simply a demonstration of different priorities. Scylla's USP is performance, so a lot of elbow grease is spent there. The Apache Cassandra com
by _benedict 5y ago
In part it's also simply a demonstration of different priorities. Scylla's USP is performance, so a lot of elbow grease is spent there. The Apache Cassandra community is focused primarily on operating at scale, as that's its USP.
Performance is adequate for Cassandra, so the community has (for several years) primarily focused elsewhere. It will be a priority again in future, but in the meantime with many huge scale users out there the community has focused on guaranteeing correctness and stability at scale. For example, the Harry[1] toolkit for validating huge databases, and an adversarial cluster simulator[2] for exposing distributed and other complex bugs. Also a huge amount of behind-the-scenes work that isn't so easy to call out.
The community is now focusing on expanding the utility of the database for these use cases. For example the recently proposed enhancement to bring state-of-the-art general purpose transactions[3] to Apache Cassandra.
[1] https://github.com/apache/cassandra-harry https://github.com/apache/cassandra-harry
[2] https://cwiki.apache.org/confluence/display/CASSANDRA/CEP-10%3A+Cluster+and+Code+Simulations https://cwiki.apache.org/confluence/display/CASSANDRA/CEP-10...
[3] https://cwiki.apache.org/confluence/download/attachments/188744725/Accord.pdf https://cwiki.apache.org/confluence/download/attachments/188...
[edit] disclaimer: I’m an Apache Cassandra contributor involved with some of the above work.
- PeterCorless 5y agoThe only clarification I want to make here is that Cassandra is focused on horizontal scalability, which Scylla can match. However, Cassandra isn't (now, or yet) focused on vertical scalability, such as using it in current instances that have dozens of vCPUs. This was shown in the "4 vs. 40" test of Scylla vs. Cassandra, where 4 boxes of a vertically scaled Scylla (72 vCPU each; 288 vCPUs total) were able to perform the same or better as 40 boxes of Cassandra (with 16 vCPU; 640 vCPU total). Definitely newer JVMs are improving things such as latency, and there are now some that are NUMA-aware, but in a JVM you are literally straight-jacketed from seeing the raw hardware you are running on. And that will impact to greater or lesser degrees your ability to take advantage of it.
- _benedict 5y agoThis is not what I was referring to, no. When operating a database at huge scale surprising things happen, because everything that can happen will happen. So operators are interested in ensuring the database behaves well in these extreme circumstances. This isn’t specifically about horizontal scalability, though that is a necessary component. To your point about vertical scalability, no doubt Scylla performs better here. However the details of your mentioned comparison are perhaps misleading, as Cassandra can happily exploit more than 16vCPUs before its performance materially plateaus. While it’s true that the JVM imposes some restrictions, they do not translate to a difference in performance on the order of that claimed in this post. JVMs have also been NUMA-aware for some time. The main explanatory factor is relative investment and focus.