6 ms·
(MemSQL CTO here) 1. MemSQL is running with synchronous replication in all these benchmarks. All data is stored on a 2nd machine before any transaction is ack
by AdamProut 7y ago
(MemSQL CTO here)
1. MemSQL is running with synchronous replication in all these benchmarks. All data is stored on a 2nd machine before any transaction is acknowledged as committed. You’re right this is not as strong as running with both synchronous writes to disk and over the network. MemSQL supports this as well and results in about a 40 to 50% performance hit depending on the disk speed. Very few of our customers run in this configuration so we didn’t include it (the edge case of multiple machines losing power is not worth the performance hit for them).
2. Can you point me to what you’re describing in the TPC-C specification? I have never heard of what you’re claiming. TPC-C has maximum allowed latency requirements for the 5 transaction types it runs and also requirements around the mix of those transactions in the workload. The goal of the benchmark is still to run as many "New Order" transaction per minute while maintaining the latency requirements of the other unmeasured transactions running in the background (this is what tpmC stands for). We used the Percona TPC-C driver for MySQL to handle this (with a few small bug fixes).
3. The main thing we wanted to show is that our performance on TPC-DS is similar (better at some scale factors, slower on others) to data warehouses that specialize in running these types of queries. We likely should have provided more details (per query break downs and what not).
4. We used the Percona MySQL TPC-C driver with some changes to make the initial data loading faster. That driver uses the “FOR UPDATE” clause in MySQL instead of running in serializable isolation level.
I know you did a lot of work on CockroachDB. The point of the blog post was not to attack cockroach (I personally didn’t want to mention it at all), but to show how MemSQL is different. We are one of the few distributed SQL databases with competitive results on all 3 major TPC benchmarks.
- arjunnarayan 7y ago1. Comparing numbers from one system (Cockroach) that adheres to strict durability requirements to another that does not (MemSQL) is apples to oranges, especially, as you point out, you see a 2x performance hit when you impose that requirement. 2. What you're looking for is the 'Think Time' mentioned in the TPC-C spec[1] (table in 5.2.5.7). From 5.2.5.2, I quote: > for each transaction type, the Keying Time is constant and must be a minimum of 18 seconds for New- Order, 3 seconds for Payment, and 2 seconds each for Order-Status, Delivery, and Stock-Level. Chapter 4 is pretty thorough on elaborating on this. The comment under section 4.1.3 explicitly states: > Comment: The maximum throughput is achieved with infinitely fast transactions resulting in a null response time and minimum required wait times. The intent of this clause is to prevent reporting a throughput that exceeds this maximum, which is computed to be 12.86 tpmC per warehouse. Again, CockroachDB numbers are right up against this limit - because the database is waiting, as required! It's within ~99% of the maximum allowed. No bar is allowed to go more than 1% higher! So stacking a bar chart next to it that goes 10x higher is pretty misleading. 3. I'm pretty impressed that you can run all the TPC-DS queries. That's pretty impressive. But performance wise, there really isn't enough fleshed out, and given that the TPC-DS authors explicitly disavow the single metric that you use (power test numbers), is simply too little to claim parity to existing databases. That said, in this conversation I'm an OLTP guy; I'll let others more experienced with Data Warehouse benchmarking take this up, e.g. [3] 4. This one I'll concede that you are doing the appropriate thing as per spec (SELECT FOR UPDATE ensures serializability), but it's the single part of the spec that's not held up over time - the paper "Making Snapshot Isolation Serializable" is a great explanation of just what lengths you have to go to to prove that a set of transactions only provide serializable histories when run in a degraded isolation mode. That said, fair enough, no anomalies will be present due to Alan Fekete's proof. But do note that CockroachDB is doing a lot of extra work (work that MemSQL can elide, since it's simply not checking for isolation anomalies) to ensure that histories are always serializable[4]. 5. While I don't work there, I did a lot of work specifically on benchmarking CockroachDB, and would like to politely request that you take down those bars for CockroachDB, since you're taking numbers that are shackled to the THINK TIME maximum and comparing them to a system that is not. [1]: http://www.tpc.org/tpc_documents_current_versions/pdf/tpc-c_v5.11.0.pdf http://www.tpc.org/tpc_documents_current_versions/pdf/tpc-c_... [2]: https://dl.acm.org/citation.cfm?id=1071615 https://dl.acm.org/citation.cfm?id=1071615 [3]: https://twitter.com/gregrahn/status/1128448156180422656 https://twitter.com/gregrahn/status/1128448156180422656 [4]: I'll shamelessly plug my blog post on this for the reader interested in more about transaction isolation levels: https://ristret.com/s/f643zk/history_transaction_histories https://ristret.com/s/f643zk/history_transaction_histories
- dkhenry 7y agoThat "Think Time" that you are referring to is supposed to emulate users running transactions on the database. So its not the database waiting its the driver waiting. While I do understand the reason for putting that in, you know very well that violating that limit doesn't artificially give CockroachDB or MemSQL an advantage when you are talking about 100,000 warehouses and a random distribution off transactions. If CockroachDB is concerned about THINK TIME enough to ask for the numbers to be removed, this would be a great opportunity for them to remove that limit and see exactly how much they could push the benchmark.
- knz42 7y agoYou should read up on why this think time exists. It has nothing to do with "emulating slow clients" and everything to do with not claiming "I have a fast database" by running gazillion txn/sec on just 1MB of data in RAM. TPC-C requires that you increase the amount of "live data" if you want to display/advertise more performance. That's the benchmark's rule. If you want to benchmark something else, that's fine, but then 1) don't call it "TPC-C" 2) don't compare with databases that play by the rules.
- dkhenry 7y agoI know why it exists, and my point is they are not trying to show a gazillion txn/sec on 1MB of data. You are looking at a dataset that is several TB's. They have far surpassed the point where a vendor is trying to cheat by putting all the data in L1 cache and claiming to be fast.
- AdamProut 7y agoThank you for the the details. Its pretty clear at this point that its not a fair comparison. The TPC-C driver we used (Percona's MySQL TPC-C driver) is pushing MemSQL as hard as it can and isn't following the "think time" part of the spec that artificially slows the driver down. So, this understandably gives us higher throughput numbers. We'll remove the comparison and make it clear the driver we are using isn't obeying the "think time" part of the spec. Again, our goal here is not to have some showdown with cockroach. We don't really compete with each other. Our goal is to show the breadth of workloads MemSQL can run (fast in-memory point queries as well complex OLAP queries over large tables). None the less, we should have caught this before we published the article. I appreciate the correction.