4 ms·
TPC-C's intent is to make a tradeoff between two competing goals: benchmarking total data stored and processing transactions on the stored data. It's one point
by arjunnarayan 8y ago
TPC-C's intent is to make a tradeoff between two competing goals: benchmarking total data stored and processing transactions on the stored data. It's one point in the spectrum of possibilities, but one that has withstood scrutiny for many decades. I agree that it would be nice to store 100s of terabytes in a cluster, but one then also has to run transactions over that stored data.
But for perspective, TPC-C over a 100TB dataset would be 1 million warehouses, and 12 million tpmC. Only one TPC-C benchmark EVER has hit this scale, Oracle, in 2010, and had a hardware cost of 30 million dollars.
We'd like to get to that scale one day, and benchmark it, just as we have today. But it's not where the industry is at right now. CockroachDB keeps up with TPC-C 10,000: this means 800GB, replicated 3 ways across a 30 node cluster. Aurora RDS doesn't keep up with the transaction rate at a mere 800GB of data (the 10,000 warehouses number). They claim to store up to 64TB of data, but what's the point of storing all that data if you can barely run any transactions on it?
- qaq 8y agoIndustry for what type of databases? 100+ TB DBs are very common thing in the RDBMS world. Doesn't have to be TPC-C just some benchmark with reasonable dataset size. 2010 was 8 years ago in modern world an x86 box can have 8 CPUs with close to 200 physical cores 6TB RAM and even consumer grade SSDs can store 800 GB multiple times on a singe drive. Outside of very niche scenario why would I use distributed database if I am not sure it can handle scale that plain old RDBMs can handle? The interesting use case is dataset size that a plain RDBMs would not be able to handle.