28 ms·
CockroachDB 2.0 Performance Makes Significant Strides
- baconomatic 9y agoI'd love to hear from someone who has implemented this in production. Seems like really cool tech, but haven't had a chance to use it on a project yet.
- smnscu 9y agoWorks great, just a tad slow. Hopefully this improves things. Deploying with Kubernetes is pretty seamless as well.
- welder 9y agoUsing it in production currently with dual-write and dual-read to compare perf. I'll do a write-up showing how Cockroach performs to Citus and Cassandra for my use case.
- some_account 9y agoPlease do, that would be a very popular read I think.
- qeternity 9y agoWe use Citus and Memsql (big data analytics use cases). How does Cockroach handle joins and other OLAP style queries?
- arjunnarayan 9y agoCockroachDB is not (yet) ready for use on OLAP-style workloads. Our performance work has focused on OLTP workloads so far. That said, we do great on OLTP joins (which is a stressed in the TPC-C workload).
- manigandham 9y agoYou're not going to get better performance for OLAP than MemSQL's columnstore and in-memory rowstore for reference tables to join. Citus is great if you want the Postgres interface but is still using standard rowstore tables. CockroachDB is similar with rowstore performance but with added distributed consensus overhead. They are both much better for OLTP and sharding. CockroachDB also provides easy high-availability and replication.
- qeternity 9y agoYeah this is what we do. Citus is our single source of truth and powers a few interactive apps and admin panels. These sync hourly to our Memsql cluster which is cstore + ref tables and works amazingly.
- ams6110 9y agoMemSQL is one of those "ask for a quote" products. What are some rule of thumb estimates for what it costs?
- ddorian43 9y ago$25K/year/box from previous comments
- manigandham 9y agoLicensed by total RAM of all nodes. $25k/year minimum license now, but you should still talk to them if you're a small company. Regardless of price, I highly recommend the product as one of the most polished data warehouses available for on-prem/self-managed operations.
- shaklee3 9y agokdb+?
- manigandham 9y agoSure, but it's far more expensive and not as generally usable as the mysql-flavored MemSQL for common data warehouse scenarios. Performance will be similar but there are differences in functionality like kdb's asof joins that can't really be compared. kdb+ is much better for numeric/financial analysis apps, especially when used with the integrated query language and interpreter environment.
- xstartup 9y agoWe use clickhouse cluster with 1000 nodes and 50000 GB clickstream data.
- no1youknowz 9y agoI'm using MemSQL's columnstore myself and the performance is nothing short of amazing. I migrated from Citus DB to it for OLAP workloads. But for what reasons are you using Citus as well? Would like to know if I am missing something or hear another perspective. Can you explain your use case? Thanks
- baconomatic 9y agoPlease do!
- didip 9y agoPlease post your write-up on HN! I would love to read CockroachDB performing in real world.
- api 9y agoWhat drug was the person who drew that graphic on?
- 2474 9y agoAsk him. https://www.dalbertbv.com/about/ https://www.dalbertbv.com/about/
- GrayShade 9y agoWindows 7 had some similar wallpapers (Scroll down): https://blogs.msdn.microsoft.com/e7/2009/05/02/a-little-bit-of-personality/ https://blogs.msdn.microsoft.com/e7/2009/05/02/a-little-bit-...
- evrydayhustling 9y agoHeavy doses of Hieronymous Bosch? https://en.wikipedia.org/wiki/Hieronymus_Bosch https://en.wikipedia.org/wiki/Hieronymus_Bosch
- some_account 9y agoCongratulations to the cockroach team for putting out an awesome product :) Would be great to see how it compares against postgres in similar scenarios.
- evrydayhustling 9y agoGreat stuff. I appreciated being educated about TPC-C, and the whole spirit of not focusing on vanity benchmarks!
- itsdrewmiller 9y agoSame here, but in educating myself more I found that TPC-C seems to be a somewhat obsolete metric compared to TPC-E (see https://stackoverflow.com/questions/9246939/what-is-the-difference-between-tpc-c-tpc-e-and-tpc-h-benchmark https://stackoverflow.com/questions/9246939/what-is-the-diff...). Why use the old one here? edit: Looking into it even further, I agree with the co-author's response here that TPC-C is still an appropriate metric. TPC-E is different and newer but still not as widely used.
- arjunnarayan 9y agoI don't think it's true to claim that TPC-C is obsolete and subsumed by TPC-E. They are both different OLTP benchmarks, with different characteristics. TPC-C is more write heavy, TPC-E is far more read heavy. It's true that TPC-E is newer, but doesn't deprecate TPC-C (the way TPC-A, for instance, is now deprecated). We chose TPC-C because it's far more understood than TPC-E in 2018. We wanted to provide understandable benchmarks that can be put into context with other databases. Other databases report TPC-C numbers, so we choose to do so as well.
- tyingq 9y agoIt seems not used much anymore. Follow that link (http://www.tpc.org/tpcc/results/tpcc_results.asp?print=false&orderby=tpm&sortby=desc http://www.tpc.org/tpcc/results/tpcc_results.asp?print=false...) and sort by either score, or price performance. The vast majority of top results are a decade old or more. I couldnt find anything less than 5 years old without going to second/third pages. And the top results are usually crazy high number of cores clusters. The Sun example was over 1700 cores.
- makmanalp 9y ago
- deleted 9y ago[deleted]
- ngsayjoe 9y agoI will never use this product no matter how good it is due to its name. Hope you can understand that many ppl like me suffer phobia from cockroaches. (PS: As expected ppl will downvote this, but just trying to give a very valuable feedback to ppl who don’t understand cockroach phobia.)
- unethical_ban 9y agoThat's unfortunate, but I hope your post isn't suggesting they shouldn't have named their product as they did. Should people who don't like large numbers not use Google? Should people who fear fire not use Firebase? Should people who don't like coffee not use Java? Moreover, should those people suggest the names be changed due to their phobia? We could number all database servers. Server 1, Server 2, Server 3, Server 4... but 4 is unlucky in China, so we can't use that.
- ngsayjoe 9y agoNo not at all im just giving my feedback of why I couldn’t use a great product sadly due to my phobia suffering from its name that many ppl might face the same but don’t bother to give feedbacks. (PS: Yes if your target customers are Chinese you should avoid using number 4 especially in real estate)
- atomical 9y agoAre there people or cultures who have a fondness for cockroaches?
- soperj 9y agoWall-E
- tdb7893 9y agoAlmost no one fears coffee or large numbers in the same way people fear cockroaches and fire also doesn't generally elicit near the same reaction (there are even types of fires that people react positively to, like camp fires, and many people like the smell of fire). They are allowed to name it whatever they want but it's hard to imagine a name with more negative feelings associated without getting vulgar or ridiculous.
- Rafuino 9y agoHow much and what kind of memory and storage (SATA SSD, NVMe SSD, HDD?) is included in the 3 nodes used for testing? This benchmarking is really interesting but the next level is to understand the cost per tmpC measured. Memory especially and storage is a big component of cost these days.
- arjunnarayan 9y agoShort answer: 3 n1-highcpu-16 GCE VMs with Local SSDs attached. We're working on a complete disclosure document, with comprehensive reproduction steps to replicate all our numbers. This document should be out in a couple of weeks. We want to walk you through, command by command, on how to reproduce these numbers, and verify the results for yourself.
- Rafuino 9y agoThanks for the short answer. Would be good to know how many local SSDs are attached though for the 850 warehouse scenario. The TPC-C documentation says each warehouse maintains 100,000 items in their stock, but I can't surmise from that how much storage is required to hold 850 warehouses' worth of data. I'm impatient though so let me try to work through the #s myself. I'm using GCP's monthly reserved pricing in the US-Iowa region as a reference as of today's pricing. A n1-highcpu-16 GCE VM costs $289.84/month. Local SSDs are added at 375GB per drive, and they cost $30/month at $0.08 per GB. I highly doubt you could fit the ~1250 warehouses (what got you the peak TPM-C) on 375GB local SSD, but I have to make assumptions here! So, now you're paying $319.84 per instance per month, or $949.52 for 3 of these instances. At 16,150 TPC-C, you're paying roughly $0.06 per TPC-C, or, looking at it the other way, you're getting 16.83 TPC-C per dollar spent each month. Is that good? I don't know! Now, the really interesting question is, is that TPC-C/$ on CRDB 2.0 actually better than TPC-C/$ on CRDB 1.1? The answer lies in how many local SSDs you have to provision to reach that peak throughput. Peak is at ~1300 warehouses on CRDB 2.0, and ~800 warehouses on CRDB 1.1. Does anyone with more knowledge here know how much storage you need per warehouse in the TPC-C test?
- jordanlewis 9y ago
- qaq 9y agoHow is this meaningful without detailed setup description? http://www.tpc.org/tpcc/results/tpcc_results.asp?print=false&orderby=tpm&sortby=desc http://www.tpc.org/tpcc/results/tpcc_results.asp?print=false... Looking at this list of results one wonders what those results actually mean?
- ElijahLynn 9y agoAt the bottom of the article it says that information is coming: "Note: We have not filed for official certification of our TPC-C results. However, we will post full reproduction steps in a forthcoming whitepaper."
- qaq 9y agoI guess will have to wait. The progress they made is obviously impressive but would really help if one could understand the overhead vs conventional RDBMS 5X might be OK 20X not so much.
- espadrine 9y agoThe minimum latency will still be at least ~5ms (worse if nodes are very far apart) and especially bad for contended reads (because of the way they use clocks for serializable operations). A traditional RDBMS does not have to worry about split brain decisions, but it can hardly do multi-master in the intuitive way.
- pinars 9y agoI think you can still drive some insights. I clicked on the TPC-C results you shared and read their executive summaries. The Oracle on SPARC cluster (at the top, 2010) performs 30.2M qualified tx/min vs the 16K tx/min in this blog post. The Oracle cluster also costs $30M, which is clearly higher than the Cockroach cluster's cost. That said, the TPC-C benchmark is new to me. Happy to update this comment if I'm misreading the numbers. (Edited to incorporate the reply below.)
- 9y ago
- strict9 9y agoBefore clicking the comments link, I always know what to expect in HN comment section for a CDB post announcing their latest milestone or feature: A lot of congrats and excitement, questions about who uses it in a production environment, very specific use-case questions, and of course the name. Weird how predictable the response to one company/tech always is.
- latenightcoding 9y agoPeople complaining about the name and how they are never going to be able to use it in production because of how gross cockroaches are is definitely the most recurring point. I think it worked well for them, since everyone remembers the name, specially with all the distributed stores coming out lately.
- ngsayjoe 9y agoSometimes i wonder the evolution of cockcroaches' grossness has something to do with its high survivability?
- johnmarcus 9y agocame here to comment on the name.
- beamatronic 9y agoI came here to upvote comments about the name
- misterbowfinger 9y agoSo. For me, personally, I don't care about the name. I generally care that it's great tech, and it clearly has a great team behind it. However.... If I worked at CockroachDB, and I saw the negative feedback around the name, I'd take it to heart. At the end of the day, the name is marketing for the hard work of their engineers, and marketing for the engineers that want to use this DB (remember, they need to sell it to their managers who may not be technical). This issue can show up in unexpected ways. For example, for cloud providers like Compose (IBM company), would they be comfortable with putting "CockroachDB" on the front page? They might if it's good enough, but it's at least a consideration (i.e. another meeting, another stakeholder to convince). Or how about an enterprise company that's going through due diligence, and when their client asks them about their tech stack do they say "CockroachDB" or do they obfuscate the name by saying "It's a high-performance distributed database". That's a crucial moment to market CockroachDB, and it could get lost. As sad as it is, saying that you're using MySQL "because Oracle" is a point of leverage for some sales people. Is the name worth it? Asking honestly.
- johnmarcus 9y agoI will not use this product based on it's name alone, it give me jeepers. Petty? Damn straight it's petty. Doesn't make it less real though.
- farseer 9y agoCame here to say the same! I am sorry but this name alone prevents me from exploring this product. Whoever chose that has probably never encountered a cockroach....flying.
- jrs95 9y agoIt's just a pun about it being nearly indestructible. Pretty good name imo
- wufufufu 9y agoShame. We've got a CockroachDB infestation in both our East and West datacenters and it's been critical in reliably scaling thus far. Sometimes a CockroachDB egg sac will go down, but we can spawn another which will hatch very quickly. The only downside is when our ORM burrows deeply into human ears, causing pain and hearing loss.
- hellofunk 9y agoNice pun there. Cockroaches do indeed have a habit of making significant strides.
- segmondy 9y agoI like what Cockroach is doing, I'm rooting for them to grow and survive. Unfortunately the only time I hear about it is when they post blogs. I never hear about it from other people.
- AlimJaffer 9y agoraises hand we're using them extensively. They're our database of choice that we've paired with Nakama[1] which is an open-source, distributed server. Have nothing but great things to say about the database itself in terms of growing performance and the team behind it :). They've been great to us since day-1. [1] github.com/heroiclabs/nakama
- netghost 9y agoWhat kind of workload are you using it for? What's been your biggest win while using it?
- AlimJaffer 9y agoA couple of our use-cases include: good KV access (stored user data etc.) and listing blocks of data that has been pre-sorted on disk at insert time (leaderboard records, chat message history etc.). As well, the clustering technology is particularly useful at scale. We work in the games space with some very large games in production, which allows us to spread the load across multiple database nodes and offers us peace of mind regarding redundancy.
- skybrian 9y agoI wonder how far apart those three nodes are and how much the latency between them matters?
- ardit33 9y agoAnd their fucking stupid name..... if they fail, they surely will be remembered as the company with the idiotic name. I can't see serious engineers working on a company named "Cockroach".
- dexterdog 9y agoCockroaches are considered pretty durable right? In the 80s I remember the line was always that after the nukes landed there would only be cockroaches and twinkies left. That's not a bad thing to say about a database.
- jrs95 9y agoAnd TwinkieDB would be a copyright violation! ;)
- awalton 9y agoBut you could also call it, "BombproofDB", "NukesafeDB", "GeodistDB" or something that gets the same idea across.
- crispinb 9y agoThis is the sort of obtuse insistence on narrow denotational semantics that makes everyone avoid engineers at parties ;)
- sheeeep86 9y agoI dont like when companies are not transparent about the pricing of their product. If you have a price page, show the price, so that Í can decide if this is relevant for me or not ...
- ahmedalsudani 9y agoThe only thing the Enterprise offering gives you is priority access.
- mjibson 9y agoEnterprise allows access to various features like distributed backup and restore.
- ahmedalsudani 9y agoAh, my mistake. I stand corrected.
- true_religion 9y agoIt's not relevant to you. Enterprise pricing generally basically scales with the size of your company/budget and how much trouble they think you'll be worth as a customer. As a rule of thumb, it starts at just above 1000 USD per unit, and goes up from there. Many contracts are bespoke orders especially when you're dealing with a small company, so you can't have transparency since there isn't a single product.
- nhumrich 9y agoI would usually agree with you, but cockroach is so new, I doubt they have any type of fixed price. They probably work it out on a 1-by-1 basis.
- wilbeibi 9y agoThe thing I really don't get is why CockroachDB is avoid benchmarking with it's rival tidb (https://github.com/cockroachdb/docs/issues/1412 https://github.com/cockroachdb/docs/issues/1412). tidb already pretty mature, used in many big companies (Let's say, Didi, which on the similar scale data with Uber, and banks). Even if I like CockroachDB's pg sql more, it would be helpful to have the comparison/benchmark to show something more.
- manigandham 9y agoTiDB is much more complicated to run with several moving pieces. Also missing a lot of standard relational features as you can see from the roadmap: https://github.com/pingcap/docs/blob/master/ROADMAP.md https://github.com/pingcap/docs/blob/master/ROADMAP.md
- jinqueeny 9y agoAs a distributed database system, the highly-layered architecture of TiDB is not a disadvantage because it makes TiDB easier to debug and diagnose when issues occur. The complicity resulted from the separated modules can be easily tackled by automatic deployment tools and microservice framework. For example, TiDB can be easily deployed either using Ansible: https://pingcap.com/docs/op-guide/ansible-deployment/ https://pingcap.com/docs/op-guide/ansible-deployment/ or Docker Compose: https://pingcap.com/docs/op-guide/docker-compose/ https://pingcap.com/docs/op-guide/docker-compose/ TiDB has been widely adopted by many users (https://github.com/pingcap/docs/blob/master/adopters.md https://github.com/pingcap/docs/blob/master/adopters.md) in production because it support the best features of both RDBMS and NoSQL. It is quickly evolving and iterating based on users’ requirements which are prioritized and listed on the Roadmap.
- manigandham 9y agoSure, I get the advantages of layers in architecture, especially the growing trend of separating compute and storage. However there's a difference between architecture and deployment, and having everything in a single package makes operations much easier. CockroachDB also uses a KV storage layer (using RocksDB) with SQL on top.
- etaioinshrdlu 9y agoProject idea: globally hosted / managed CockroachDB that lets developers quickly start building small apps cheaply or free using this database. This database has the potential to dethrone Spanner in a major way.
- joris 9y agoThat’s on their roadmap: https://www.cockroachlabs.com/docs/stable/frequently-asked-questions.html#does-cockroach-labs-offer-a-cloud-database-as-a-service https://www.cockroachlabs.com/docs/stable/frequently-asked-q...
- SoulMan 9y agoMy Org/team is too conservative to use this is they have to hire ops and too froogle to use spanner.
- kevincox 9y agoWhy do you think that hosted cochroachdb would be cheaper then hosted spanner? Google has been optimizing spanner performance for years so I would expect that it will be cheaper to run for quite a while. Of course the markup can be different but I wouldn't expect it to make a huge difference.
- tyingq 9y agoMaybe training? Spanner's lack of UPDATE/INSERT/DELETE requires a team that's trained pretty specifically on how it works.
- ZeikJT 9y agoFrugal, froogle was the old name for google shopping.
- SoulMan 8y agoThanks. Frugal is exactly what I meant.
- elvinyung 9y agoSince you only have 3 nodes, doesn't that mean every range is replicated to every node? Doesn't that make joins trivial (i.e. no different from non-distributed joins)?
- d4l3k 9y agoYeah, though from what I understand this benchmark is measuring both transactional read and write performance rather than just join performance. Transactional writes are likely the slowest thing since they need to talk to all replicas.
- deleted 9y ago[deleted]
- elvinyung 9y agoActually hmm, do reads need to talk to all replicas in this case (serializable isolation)?
- ComputerGuru 9y agoAs I understand it, reads only need to talk to the lease owner, which in turn bypasses raft since write consensus guarantees atomicity of after completion of write intents. Cockroach tries best-effort to have the lease owner and the raft leader be the same.
- atombender 9y agoLooks very promising! We've looked at Cockroach for a particular project, and we've been concerned that performance wasn't good enough. Cockroach performance seems to scale linearly, but single-connection performance, especially for small transactions, seems rather dismal. Some casual stress testing against a 3-node cluster on Kubernetes showed that small transactions modifying a single row could take as much as 7-8 seconds, where Postgres would take just a few milliseconds. The documentation recommends that you batch as many updates as possible, but obviously that doesn't work for low-latency applications like web frontends that need to be able to do small, fine-grained modifications.
- tyingq 9y ago" small transactions modifying a single row could take as much as 7-8 seconds" That's surprising. I wasn't expecting CockroachDB to be really fast, given the constraints they work within. But that sounds more like a bug or config error. Unless perhaps you mean a really high number of processes trying to update the same row at the same time? Like a global counter or something?
- lobster_johnson 9y agoIndeed, the stress test updates just one row, which mirrors certain write patterns in our application. I just started this testing, so we'll see what happens when I extend it to more than one row.
- jetrink 9y ago7-8 seconds seems extremely long. Human beings performing the raft consensus algorithm using paper and pencil over Skype wouldn't be much slower than that. Are you sure everything was working correctly?
- lacker 9y agoI don't know about you, but it would take me a lot longer than 7 seconds to perform the raft consensus algorithm with paper and pencil.
- pieterhg 9y agoGreat stuff but this name really doesn’t work. Make it a name with positive connotations.
- as1mov 9y agoWhile we are changing names for petty reasons, let's rename Python to something else, since a sizeable part of the population has a phobia of snakes.
- expliced 9y agoI think it's a great name once you get the.. uh.. pun behind it.
- arbitrage 9y agoWhat is the pun?
- mmilano 9y agoI agree, the unfortunately clever/punny name will detract from potential consideration based on illogical human subconscious thinking.
- Asdfbla 9y agoDoes someone have more information about how they implemented serializable in such a way that, as they claim, performance isn't negatively impacted? Seems pretty hard to achieve that.
- d0ugie 9y agoHadn't heard of Cockroach but based on the article, this thread and the rest of their site it sounds at least worth installing on a few hobby nodes if only to get familiar with the behavior and configuration should a need arise - like Cassandra was years ago when I had already on my own learned the gist of it, sort of a road not taken relative to my then-firm's usual prescriptions (MySQL and Mongo), it turned out to be perfect for my team's needs (paperwork to get permission to use it notwithstanding). Thanks for posting and good luck!