7 ms·
Revisiting 1M Writes per second
- jbellis 12y agoThis is running Cassandra 1.2.x which is over 18 months old. Here's some performance notes on the latest (2.1rc4): http://www.datastax.com/dev/blog/cassandra-2-1-now-over-50-faster http://www.datastax.com/dev/blog/cassandra-2-1-now-over-50-f...
- arielweisberg 12y agoWould be great to know the exact data set size (not size in the database). I can't get an order of magnitude sense of what I am looking at without that. I know I can divine it from the parameters to stress, but I have no idea if the row keys generated by different clients overlap and I don't know the default number of columns nor their size. I think it's also important in this kind of benchmark to describe the distribution of access especially for a read intensive benchmark. Without that you really don't know what your are looking at. I am a fan of scrambled Zipfian.
- jdf 12y agoIt seems like this info should be front in center in the test. 1M/s 100kb writes is much more impressive than 1M/s 16 byte writes. That said, there's a previous benchmark linked to at the top of the post: http://techblog.netflix.com/2011/11/benchmarking-cassandra-scalability-on.html http://techblog.netflix.com/2011/11/benchmarking-cassandra-s... The client is writing 10 columns per row key, row key randomly chosen from 27 million ids, each column has a key and 10 bytes of data. The total on disk size for each write including all overhead is about 400 bytes. There are 3 replicas, so figure that in as well.
- MoOmer 12y agoI love using Cassandra; it's been a dream for analytics. Thank you Netflix et al. for not only moving the project forward, but providing research and commentary like this and others available at http://wiki.apache.org/cassandra/ArticlesAndPresentations http://wiki.apache.org/cassandra/ArticlesAndPresentations
- rdtsc 12y agoMy other favorite "scalability" study is from WhatsApp: http://www.youtube.com/watch?v=c12cYAUTXXs http://www.youtube.com/watch?v=c12cYAUTXXs That's 'Billion' with a 'B': Scaling to the Next Level at WhatsApp (that walk title was create before the acquisition and was mean to imply message count, after the acquisition it got a secondary meaning). The one thing that is fascinating about it, is how small their team was compared to the volume and complexity of the operation.
- turnip1979 12y agoI wonder where the term sidecar originated from and what is the precise definition? Is this something invented by the netflix OSS people or does it predate that?
- jedberg 12y agoIt's from here: https://www.google.com/search?q=motorcycle+sidecar&source=lnms&tbm=isch https://www.google.com/search?q=motorcycle+sidecar&source=ln...
- turnip1979 12y agoI was hoping for a more precise software concept. Maybe it is loosey goosey.
- jedberg 12y agoThe definition I use is a piece of software that runs independently to encapsulate infrastructure libraries. In other words it gives you a way to access the libraries without having to import them into your code.
- dnBldGVy 12y agoSo the test was run on a 285 node cluster with 60 clients. It would be nice to know how they arrived at those numbers. Was there some sort of formula used to calculate how large each group should be? How much trial and error was involved.
- arrryarr 12y agoHow many companies can possibly afford the cost and management pain of running a 285 node database? Few would have 285 servers of any type? So if this is what it takes to get 1M writes on Cassandra that is some poor ROI. 10K writes/sec is not impressive.
- opendais 12y agoHow would you build out at a metrics cluster that needs 1M writes/sec? 285 nodes that are easily automated to create/destroy/monitor doesn't seem like a management pain to me, personally. Just depends if you have the write internal tools built.
- arrryarr 12y agoOk, so you're super-admin. And about the ROI aspect? What's the cost for achieving the 1M write/sec metric in that mammoth cluster?
- opendais 12y agoI asked you how you would architect something that needs 1M writes/sec if you weren't using Cassandra. Ignore the fact that netflix is making this into a single cluster. I'm not saying I'd do it by myself. I'm saying with the write tools its doable and not unreasonable.
- Retric 12y agoI have seen single servers do 250k writes/second. So it really depends on the data and access patterns more than some arbitrary one size fit's all solution.
- rsynnott 12y agoI suspect systems where the dataset, or at least the indexes, largely fit in memory?
- 12y ago
- JonoBB 12y agoCouldn't help but notice: $398.70 per hour = $9568.80 per day = ~$3.5m per annum. They obviously get a discount...but still. What kind of discount do guys like this get?
- alex_sf 12y agoI'd be surprised if anyone got discounts as deep as Netflix considering their usage. For what it's worth, 3.5m/yr is about 0.08% their revenue.
- JonoBB 12y agoWell, that puts things in perspective.
- rsynnott 12y agoIf they were doing this for real, they'd be using reservations; with one-year reservations, the i2.xlarge bit (the servers) cost $905k/year.
- cpayne 12y agoFor this type of "report" the vendor will usually chuck in the hosting for free. It would have conditions like "can only be used for this test (not anything else" and "results must be published on your blog" etc. etc. Microsoft do this all the time - its great publicity. It would be interesting to see if they (Microsoft) have something showing similar results...