5 ms·
Does any one on here have some real world experience with scylla? We currently make heavy use of dynamo and are interested in something cheaper/faster. The mar
by throwawaythekey 4y ago
Does any one on here have some real world experience with scylla?
We currently make heavy use of dynamo and are interested in something cheaper/faster. The marketing material is pretty compelling but I'm unsure of how hard scylla is to operate at scale.
- zeckalpha 4y agoThey have some high a availability functionality in the open source version but not commercial. If the experts won’t run it in that mode, why do you think you’ll do a better job?
- manigandham 4y agoWhat are you talking about? High-availability is a core part of the Cassandra data model and node architecture. Scylla is designed to be highly available, and the enterprise version only adds features to make it easier to operate.
- PeterCorless 4y agoJust to note that ScyllaDB Open Source sometimes is some months to a year ahead of the Enterprise release in terms of features. Though I am not sure what specific "high availability" feature poster above is referring to. Otherwise, as you note, ScyllaDB Enterprise will be a superset of our Open Source offering. Also to minimize lag between Open Source an feature availability, we just announced new Feature releases for Enterprise which will be coming later this year and into the future. https://www.scylladb.com/2022/06/08/new-scylladb-enterprise-support-and-versioning/ https://www.scylladb.com/2022/06/08/new-scylladb-enterprise-...
- manigandham 4y agoUsed it on my past adtech startup handling billions of http requests per day (backend having multiple more calls). Solid product, it's since reached parity with Cassandra and added even more features. Great team, very helpful. It's far easier to run than Cassandra because of its performance so you need less nodes overall (which is the single biggest issue in C* ops). The enterprise edition also has better compaction strategies and an automated system to schedule it all. What scale are you looking at?
- throwawaythekey 4y agoIt sounds like you might be a future version of me! We are an adtech startup in the hundreds of millions per day right now, expecting billions per day by the end of the year. Management would be happy if I said I have a plan for tens of billions. The main difference is that we would probably lean towards the cloud version instead of self hosted (do we still need to think about clusters then??)
- manigandham 4y agoIf you use the cloud version then there are no ops to worry about.
- samhw 4y ago> far easier to run than Cassandra I know precisely nothing about Scylla, but somehow I still agree with this. Cassandra is far and away the most horrendous software I've ever had to work with (Kafka coming in a close second). This was true even when we hired a major core contributor to help manage it; even he wasn't 100% comfortable with it. The word 'anticompaction' is still enough to give me nightmares. I very much welcome – not that it's Cassandra's only problem – this trend of rewriting '00s/early-'10s Java software in simpler languages. (And, languages aside, I think Cassandra - certainly inter alia - simply sits at a level of complexity which is beyond the comprehension of any one human being. And software which is beyond the comprehension of any one human being inevitably begets bugs, even as an operator rather than a developer.)
- _benedict 4y agoI'm intrigued: The Apache Cassandra developer community is quite small and well known to each other, and the only major contributor I know to have left the community to run a cluster did so almost a decade ago? Cassandra had a lot of rough edges until very recently, and was only really suitable for sophisticated users, as most contributors would have attested. The project has matured a lot over the past couple of years in particular, as the broader community stepped up in the face of DataStax's (temporary) withdrawal. The investment by some large scale users has transformed it, and the next couple of years will do so even more. None of the issues the project faced were really related to the language chosen, in my opinion, rather than the maturity of the project and how it grew at the time.
- AtlasBarfed 4y agoNot sure why you would use it over Cassandra, which scylladb last I checked was a C++ rewrite that had impressive initial feature releases but has since lagged in feature support badly. I would migrate to cassandra first, and then prototype your workloads on scylladb. You'll want loadtest execution ability regardless so you can tune for your workloads. I'd guess ScyllaDB and Cassandra are comparable for headaches. Our clusters never go down, but there's always SOMETHING to look at. But cassandra has more community and commercial support.
- PeterCorless 4y ago1. Performance https://www.scylladb.com/2021/08/24/apache-cassandra-4-0-vs-scylla-4-4-comparing-performance/ https://www.scylladb.com/2021/08/24/apache-cassandra-4-0-vs-... 2. Features https://www.scylladb.com/2021/11/17/cassandra-and-scylladb-similarities-and-differences/ https://www.scylladb.com/2021/11/17/cassandra-and-scylladb-s... ScyllaDB used to chase Cassandra's feature set. Now, in many ways [MV, compaction strategies, workload prioritization, LWT, CDC] we can either do things Cassandra cannot do, or we're a better implementation of Cassandra's same or similar functionality. Cassandra c. 2018-2020 was a pretty moribund project, not shipping a major new release until last year [2021]. I am pretty excited to see the new work done on 4.0 and now 4.1. During the same interval ScyllaDB went from 2.0 to now 5.0 which is coming out Very Soon Now. The CEP process Cassandra maintains shows some very interesting directions they are heading in. And we showed off some of our roadmap at Scylla Summit 2022 describing where we're taking our product next. Whichever path users decide, the good news is that year-over-year users have increasingly better databases to choose from. [Disclosure: I work at ScyllaDB.]
- _benedict 4y agoRe: your comparison, LWTs in Cassandra in 4.1 (about to release) offer global 1RT reads, and 2RT for writes. Your comparison suggests you take 3RTs? I suspect you may be moving to offering 1RT for those in the local region, and 2RT for all operations in another region? You make a point of mentioning your MVs (and presumably your global secondary indexes) may get out of sync with the base table, which is the reason the Cassandra community decided to mark them experimental. The Cassandra community has become very conservative since it began being driven by users of the software. I should not draw any conclusions around the relative merits of features from this conservatism. The Cassandra community paused feature development for several years, due to this very conservatism, in order to focus on delivering safety and reliability at scale. Now that's done, you can expect a lot more visible activity.