8 ms·
Going Head-To-Head: Scylla vs. Amazon DynamoDB
- aswartz 8y agoScylla is awesome!
- thekozmo 8y agoWe'd love to hear your feedback on it. Dor (article writer)
- Artemis2 8y ago> The global reservation is divided to multiple partitions, each no more than 10TB in size. I think this should be 10GB partitions. By the way, maybe it’s worth mentioning adaptive capacity? https://aws.amazon.com/blogs/database/how-amazon-dynamodb-adaptive-capacity-accommodates-uneven-data-access-patterns-or-why-what-you-know-about-dynamodb-might-be-outdated/ https://aws.amazon.com/blogs/database/how-amazon-dynamodb-ad...
- thekozmo 8y agoYou're right, thanks
- gtaylor 8y agoWith the caveat that I'd love to see Scylla succeed, I'm a bit bummed to see the conditions around this bake-off so slanted in Scylla's favor from the start. Seeing a multi-AZ configuration compared to a single-AZ config is an immediate disqualifier, without getting into some of the other comments on item size disparity and YCSB quirks. This was a really poor response from Scylla on these fronts: https://news.ycombinator.com/item?id=18677511 https://news.ycombinator.com/item?id=18677511
- thekozmo 8y agoI went and run a ping between 2 VMs, one on us-east1c and the other on a different AZ - us-east-1b As you can see, the ping latency is 0.5ms to 1ms on rare occasions. Of course that ping has a cost even within a single AZ so the above numbers reflect much more than the AZ round trip time. You're welcome to add 1ms to our latency numbers, it's 3x-4x better in the 99th case so there's enough room... 64 bytes from 35.173.249.24: icmp_seq=13 ttl=63 time=0.792 ms 64 bytes from 35.173.249.24: icmp_seq=14 ttl=63 time=0.584 ms 64 bytes from 35.173.249.24: icmp_seq=15 ttl=63 time=0.462 ms 64 bytes from 35.173.249.24: icmp_seq=16 ttl=63 time=0.848 ms 64 bytes from 35.173.249.24: icmp_seq=17 ttl=63 time=0.465 ms 64 bytes from 35.173.249.24: icmp_seq=18 ttl=63 time=0.451 ms 64 bytes from 35.173.249.24: icmp_seq=19 ttl=63 time=1.00 ms 64 bytes from 35.173.249.24: icmp_seq=20 ttl=63 time=0.854 ms 64 bytes from 35.173.249.24: icmp_seq=21 ttl=63 time=0.446 ms 64 bytes from 35.173.249.24: icmp_seq=22 ttl=63 time=0.460 ms 64 bytes from 35.173.249.24: icmp_seq=23 ttl=63 time=0.509 ms
- awinder 8y agoThere’s still 2 unresolved problems (at least) here: 1. running a ping test is a fine start but wouldn’t a more realistic test be to rerun the benchmark? 2. in a multi-az setup you’re going to have to pay for inter-az transfer costs between ec2 nodes. data transfer between ec2 and dynamodb is free. Also on the pricing front reserved capacity really drops your costs with dynamodb if you know you’re going to be with it for some term of time. Not that that needed to be included in the benchmark but... aws pricing is complicated, to say the least.
- thekozmo 8y agoTrue. In our brand new service offering we took it all into account. Networking is costly but doesn't change the game that much. Scylla is 4x-6x cheaper, depending whether it's on-demand (vs Dynamo on-demand) or reserved for a year (vs dynamo reservation) https://www.scylladb.com/product/scylla-cloud/#pricing https://www.scylladb.com/product/scylla-cloud/#pricing
- omiSSD 8y agoThe cost savings are mind blowing!
- planckscnst 8y agoI haven't read the whole piece yet, but I saw some things that already stood out as strange. > Please keep reading to see how diligent we were in creating a fair test case I hope this is true, but my cursory reading found several unfair spots that seem at odds with this statement. > 3-node cluster on single DC | RF=3 DynamoDB works across AZs and will continue to be available even when a data center goes down. That cross-AZ operation has benefits as well as latency costs that are not present in the ScyllaDB setup. > We hit errors on ~50% of the YCSB threads causing them to die when using ≥50% of write provisioned capacity That is surprising. It is quite common to use your full provisioned capacity without problems. My guess is that there is something not ideal about the YCSB DynamoDB library or its configuration. I'm not familiar with YCSB: does it give you stack traces that indicate why the threads failed? > Sadly for DynamoDB, each item weighted 1.1kb – YCSB default schema, thus each write originated in two accesses This is what made me come write this comment. You specifically knew this was a pessimistic case which could be easily addressed to allow DynamoDB to operate at a lower cost. Is arbitrarily settling for the default 1.1kB item size on an artificial benchmark fair? Good engineering teams use their tools the way that gives them the most benefit. Calling out that you may have to work to ensure your use case doesn't have pessimistic characteristics would clearly be fair, but I'm not convinced that just picking arbitrary benchmark settings is.
- thekozmo 8y agoFair questions, let me (Dor) answer: 1. Scylla works within different AZs too We are topology aware and can have as good and even better HA than DynamoDB. Our design is based on C* 2. We were surprised with getting small utilization too. As I wrote in the article, I think Dynamo had a hard time reaching to 1TB that quick. The population started fine and deeper in the the population it failed. Only a decrease in the rate solved it. 3. 1.1kb The is the default table setup by YCSB. It isn't optimal for Dynamo but that's life.. think about if you store a blob of 1kb - together with the key it will be more than 1kb. We were upfront about it and you can make your own calculation. Scylla will still be better. The only use case where Dynamo is better in price is when you store lots of data but with a tiny IOPS reservation. We'll address this case over time too. Cheers!
- CyanLite2 8y agoTLDR: When not accounting for DevOps and SRE costs, managing your own VMs are a fraction of the cost of PaaS services.
- jugg1es 8y agoThis all looks great but one of the primary advantages of DynamoDB is that you don't have to spend resources on maintaining or troubleshooting ec2 instances. When you have a large deployment, not having to worry about HA on your data store is hard to beat.
- sheeshkebab 8y agoYou still have to spend a lot of labor/resources maintaining and managing dynamodb. Not managing ec2s but, at scale, it’s not exactly a fire and forget type thing and requires a lot of SRE and devops attention. Also, it sounds like Scilla has a managed solution too, that costs a fraction of dynamodb, at scale. Dynamodb to me is the new generation IMS database from mainframe cobol times - fully proprietary, locked down, runs only on proprietary hardware, serviced by one vendor. It’s only a matter of time before companies using it will be stuck with multimillion $$ yearly bills, if they aren’t already.
- Dunedan 8y ago> You still have to spend a lot of labor/resources maintaining and managing dynamodb. Do you? At least with DynamoDB On-Demand there shouldn't be much maintenance left to do.
- ddorian43 8y agoYou will or you'll pay (a lot) for it.
- jugg1es 8y agoI'm not sure what you mean by having to spend money maintaining DynamoDB. What maintenance are you talking about? You definitely have to spend money designing your application and figuring out how to keep it cost effective, but once you have that set up along with the auto-scale policies, it is pretty much fire and forget.
- aynsof 8y agoAm I the only one who's a little disappointed they didn't call it CharybDB?
- biggestdummy 8y agoFYI. https://www.scylladb.com/2016/02/16/fault-injection-filesystem-software-testing/ https://www.scylladb.com/2016/02/16/fault-injection-filesyst...
- PeterCorless 8y agoWe have a Charybdefs; a fault-injection filesystem for testing. https://github.com/scylladb/charybdefs https://github.com/scylladb/charybdefs
- reilly3000 8y agoA theme of the comments so far has been around the fact that the benchmark is created in a manner that feels biased. QUICK STRAW POLL: What is the right way for that information to be gathered and presented? a. There is no bias. Companies can and should present research that may favor them, as long as they cite good sources. b. An industry analyst, paid to research multiple companies and options and present that information to their respective (most likely paying) customers. c. A journalist, presenting research in public, funded by advertising for unrelated interests. d. Reporting by peers/actual users of the system on how the products compare, and possibly about how it aligns with their technical and business goals. e. Trust no one. Conduct your own research, and publish it if you see fit to so do, or keep it to yourself. I see potential "disinformation vulnerability" with each approach. How ought we get the best info on how to align our organizations with technologies?
- RobLach 8y agoThe best benchmarks I’ve witnessed come out of academia.
- thekozmo 8y agoThe benchmark was done as fair as possible but I'm the vendor, yes, don't take my word (despite good background in OSS and earlier achievements). Listen to our users: - Here's a webinar by one of our customers explaining how they migrated from Mongo+Hive, received better consistency and simplicity while saving 5X: https://www.youtube.com/watch?v=1hXKd_rNyuE https://www.youtube.com/watch?v=1hXKd_rNyuE - See how kiwi.com got a CRAZY gain with Scylla vs Cassandra: https://youtu.be/Bqh09LG_QDE?t=833 https://youtu.be/Bqh09LG_QDE?t=833 - See Grab (South east Asia Uber) about Dynamo vs Scylladb: https://www.scylladb.com/users/case-study-grab-hails-scylla-for-performance-and-ease-of-use/ https://www.scylladb.com/users/case-study-grab-hails-scylla-... Go ahead and give it a try and see for yourself