5 ms·
How many companies can possibly afford the cost and management pain of running a 285 node database? Few would have 285 servers of any type? So if this is what
by arrryarr 12y ago
How many companies can possibly afford the cost and management pain of running a 285 node database? Few would have 285 servers of any type? So if this is what it takes to get 1M writes on Cassandra that is some poor ROI. 10K writes/sec is not impressive.
- opendais 12y agoHow would you build out at a metrics cluster that needs 1M writes/sec? 285 nodes that are easily automated to create/destroy/monitor doesn't seem like a management pain to me, personally. Just depends if you have the write internal tools built.
- arrryarr 12y agoOk, so you're super-admin. And about the ROI aspect? What's the cost for achieving the 1M write/sec metric in that mammoth cluster?
- opendais 12y agoI asked you how you would architect something that needs 1M writes/sec if you weren't using Cassandra. Ignore the fact that netflix is making this into a single cluster. I'm not saying I'd do it by myself. I'm saying with the write tools its doable and not unreasonable.
- Retric 12y agoI have seen single servers do 250k writes/second. So it really depends on the data and access patterns more than some arbitrary one size fit's all solution.
- rsynnott 12y agoI suspect systems where the dataset, or at least the indexes, largely fit in memory?
- Retric 12y agoMost of the active data-set fit in ram, but considering you can get 100+GB of RAM that's generally not much of an issue. Also a single mid range SSD can break 120K writes per second so 0+1 RAID arrays can get crazy fast for less money than you might think.
- rsynnott 12y agoWell, I mean, you _can_ get 100+GB of RAM, but you'll certainly pay for it. The machines used in the example have 30GB of RAM and 800GB SSD. For a fairly random access pattern and most of the storage used, the RAM's not going to be helping that much. Let's say they're using 600GB of the 800GB, replication factor of three, that's 57TB over their cluster. It's pretty big.
- Retric 12y agoUhh, you can get 128GB on a basic dell server that costs less than 4k. It's 512+GB of ram that starts getting pricey now days and 1TB is actually an option. Though, I agree if you dataset is large enough and you need random access it's not going to help much.
- arthursilva 12y agoWell, it's not unreasonable (the maintenance part) with Cassandra.
- threeseed 12y agoI don't understand your point. The type of companies who could afford this are the types of companies who need 1M writes/second. Which are few and far between. And yes 10K is not that impressive but 1M is. And with Cassandra you could continue to improve that number just be rolling out more nodes.
- rsynnott 12y agoThey probably mean 10K writes/sec/node, which would be correct assuming a replication factor of three. It doesn't sound huge, but if it were sustained, with a key set vastly larger than memory, it's not too bad.
- jandrewrogers 12y agoIt is quite cheap to average a million writes per second. I've done it with five servers on AWS, and that was spatially indexing billions of GeoJSON polygons through storage while running queries against the index. Many companies need far in excess of a million writes per second. Basically, most machine-generated data sources, whether it is personal location data or any other kind of telemetry. Many companies that do not generate that data themselves buy and consume it. I know of companies doing over a billion writes per second. Cassandra is pretty good for this type of thing among open source software but it is not nearly as efficient as it could be in terms of write throughput. If the storage engine is correctly designed, you should be able to drive 10GbE all the way through storage -- call it 1 GB/sec per node. However, that does mean you can't do things like mmap()-ing files; those interfaces are slow due to poor scheduling by the OS when the throughput is very high.
- corysama 12y agoOut of curiosity, what would a correctly designed storage engine do to get better throughput than mmap()ing files?
- jandrewrogers 12y agoOn Linux, you would use io_submit + O_DIRECT on a small number of large files allocated in large chunks. In short, you become the I/O scheduler and buffer manager instead of the operating system, while removing most of the implicit context switches. It does require much more code than mmap()-ing and fairly sophisticated code at that. If you are doing it well, 3-5x throughput improvement seems to be average upside in my experience, which is huge. The scheduler behind mmap() simply does not have enough context about the workload to make good paging decisions leading to a lot of suboptimal or wasted I/O, and this is magnified when the storage I/O is under pressure. In principle, if you write your own I/O scheduler you can always make sure that the optimal I/O operation is executed at the optimal time.
- larsmak 12y agoThe replication factor is set to 3, meaning that all data is stored on tree separate nodes - in different availability zones. So in practice the writes per sec is 3x.