20 ms·
CockroachDB 1.0
- nathell 9y agoI read the announcement, got all excited, then clicked "What's inside CockroachDB Core?" and got rewarded with a 404. Ouch! This itches.
- deleted 9y ago[deleted]
- orangechairs 9y ago[cockroachdb here] Yeah, we're experiencing some caching issues.
- cwisecarver 9y agoCue the comments stating that no one will use this because the name is bad.
- Gurrewe 9y agoI love their name, it clearly states what they want to achieve.
- toxican 9y agoSpread disease, infest, smell, make my skin crawl?
- jachee 9y agoAre you suggesting the name is a feature, not a bug?
- dsimms 9y agoliterally a bug! Also, I enjoyed the beta announcement title: CockroachDB skitters into Beta. Good work y'all! I hope Cockroach Labs continued success!
- notacoward 9y agoMaybe it's both.
- Double_a_92 9y agoBut the more important question is "Does it webscale?" /s
- mirekrusin 9y agoJust call it C9DB if you don't want to use the c-word.
- cholantesh 9y agoWouldn't it be C8DB?
- JetSpiegel 9y agoStylized as C8=>DB
- s_kilk 9y agoMaybe they want to keep all the startup webscale dweebs out.
- jabl 9y agoJust like, err, GIMP? What is it with Spencer Kimball and naming things that gets people so upset? It's not like other company or product names are that good; we're just used to them. Some high profile tech companies: - Google: some propellerhead big number joke (hey, I have a Phd and I don't know offhand how big a googleplex is...) - Alphabet: Really? That out of ideas? - Amazon: Some hot snake and insect-infested jungle? Why should I go there? - Microsoft: At least it gives a hint what the company does, but really... (cue the penis jokes) - Yahoo: WTF, some slang term I've never heard of before.. - Apple: Mmmm, are they organic and locally produced? Oh, they sell computers and phones? WTF?! (Yes, I've heard the backstory about Alan Turing and the poisoned apple which I guess puts me in a very small minority)
- Mizza 9y agoI mean the big issue is really the inherent positive/negative association. I wouldn't over-think it. Amazon - big, marvelous jungle. Yahoo - an expression of Joy. Apple - a delicious sweet fruit. Cockroach - a disgusting plague insect.
- cocktailpeanuts 9y agoThere's bad name, and then there's repulsive name. All the examples you mention fall under "bad name", and it's not even objectively bad, I actually think they're great names, so it's subjective. And NONE of them are repulsive. Then again, if you insist cockroaches are lovable creatures I have nothing more to say.
- jabl 9y ago> All the examples you mention fall under "bad name", and it's not even objectively bad, I actually think they're great names, so it's subjective. And NONE of them are repulsive. My argument was not that they are good or bad, but rather that we've come to associate positive things with the companies in question, and then we post-hoc come up with explanations why they names are good etc. > Then again, if you insist cockroaches are lovable creatures I have nothing more to say. I don't think they are lovable, no. But they are an evolutionary success story; they've been around for hundreds of millions of years, long before humans. And they'll be here after we humans have extincted ourselves in some nuclear holocaust/massive environmental disaster/pick your favorite apocalyptic scenario/. And if you manage to squish one, there's hordes of em left; just like I'd like my DB to be, so actually I think it's a very good name! :)
- nathan_f77 9y agoI think it's one of the worst names I've ever heard. Both because of the bug, but also because of the first syllable. That's not going to stop me from using it if I have to, but I'm certainly less interested in trying it out. And I would be embarrassed to put "Cockroach Expert" on my resume.] Disclaimer: I've come up with my fair share of bad names. SCM Breeze [1] makes me cringe now. [1] https://github.com/scmbreeze/scm_breeze https://github.com/scmbreeze/scm_breeze
- korzun 9y ago> I think it's one of the worst names I've ever heard. Both because of the bug, but also because of the first syllable. You sound like you would get offended by a slight breeze of air.
- nathan_f77 9y agoI'm sorry it comes across that way, but I'm not really "offended" by it. I just think it's a gross name and it gives me a bad feeling. It's not something I can really control.
- irfansharif 9y ago> I would be embarrassed to put "Cockroach Expert" on my resume well, now that you mention it... https://irfansharif.io/irfan-sharif-resume.pdf https://irfansharif.io/irfan-sharif-resume.pdf
- dang 9y agoPlease don't do this here—it's just as bad as the bad thing. Worse in fact, because it's smug too.
- cwisecarver 9y agoSorry dang. I know your job is hard enough.
- Gurrewe 9y agoCongratulations to the team on the relase! Everything under "The Future" really excites me, especially the geo-partitioning features. That is something that I'm really looking forward to be using!
- jazoom 9y agoThat might end up being an enterprise feature though.
- dis-sys 9y agoI really like the fact that the CockroachDB team recently did a detailed Jepsen test with Aphyr. The follow up articles from both CockroachDB and Aphyr explaining the findings are very interesting to read. For those who might be interested - https://www.cockroachlabs.com/blog/cockroachdb-beta-passes-jepsen-testing/ https://www.cockroachlabs.com/blog/cockroachdb-beta-passes-j... https://jepsen.io/analyses/cockroachdb-beta-20160829 https://jepsen.io/analyses/cockroachdb-beta-20160829
- dmix 9y ago> CockroachDB is a distributed, scale-out SQL database which relies on hybrid logical clocks I was curious what "hybrid logical clocks" meant and found the linked paper a bit over my head. I found this more layman description: http://muratbuffalo.blogspot.ca/2014/07/hybrid-logical-clocks.html http://muratbuffalo.blogspot.ca/2014/07/hybrid-logical-clock... Apparently Google used GPS/atomic clocks to keep time synced: >> To alleviate the problems of large ε, Google's TrueTime (TT) employs GPS/atomic clocks to achieve tight-synchronization (ε=6ms), however the cost of adding the required support infrastructure can be prohibitive and ε=6ms is still a non-negligible time. And CockroachDB created more of a hybrid version that works on commodity hardware. Distributed systems programming sounds endlessly challenging as you are always balancing trade-offs.
- irfansharif 9y agoYou might find our post[1] on atomic clocks, rather having to do without them, partially interesting. [1]: https://www.cockroachlabs.com/blog/living-without-atomic-clocks/ https://www.cockroachlabs.com/blog/living-without-atomic-clo...
- socmag 9y agoHey guys, I'm a fellow developer of distributed systems here. First of all I think what you are doing is great. My question is what's the point of clocks at all? The current time is a very subjective matter and I'm sure you know this, the only real time is at the point when the cluster receives the request to commit. Anything else should be considered hearsay. Specifically the time source of any client is totally meaningless since as you say further in the discussion that client machine times can be off by huge margins. If you accept that then one has to accept the fact that individual machines within the cluster itself are prone to drift too, although one can attempt to correct for that I appreciate. Wouldn't you think though that what is more important is that the order is more based on the bucketed time of arrival (with respect to the cluster). I don't see how given network delays anyone can be totally sure A is prior to B, atomic clocks or not. What is important is first to commit. [edit] Yes would love to talk privately about this topic @irfansharif
- nik736 9y agoWhat advantages do I have using Cockroach compared to Postgres, Cassandra, Rethink or MongoDB? (I know that all of them are completely different, that's part of the question)
- irfansharif 9y agoWe have an comparison page[1] that might potentially be what you're looking for. [1]: https://www.cockroachlabs.com/docs/cockroachdb-in-comparison.html https://www.cockroachlabs.com/docs/cockroachdb-in-comparison...
- elmalto 9y agoDo you have any performance comparisons as well?
- arjunnarayan 9y agoSo performance is complicated. Right now, we’re performance testing CockroachDB regularly, and everything is out in the open. Everything we do is tracked with a GitHub issue with the “perf:” prefix, if you want to follow along. Here are all our issues that track performance: https://github.com/cockroachdb/cockroach/issues?utf8=%E2%9C%93&q=is%3Aissue%20is%3Aopen%20perf%3A https://github.com/cockroachdb/cockroach/issues?utf8=%E2%9C%... Here’s our open source repository where we keep our load generators: https://github.com/cockroachdb/loadgen https://github.com/cockroachdb/loadgen A blog post (well, many) are in the works outlining our performance benchmarking. The situation on the ground is changing fast - our performance has improved rapidly over the past months, and each time we sit down to write a blog post, it gets quickly obsoleted. So, trust that we will have a blog post talking about performance very soon. Anecdotally, our customers are not finding performance to be a bottleneck. I encourage you to set up a Cockroach cluster, and try the various load generators (we've got the standards and a couple other homegrown ones in the repository).
- chuckledog 9y agoFrom the linked website: "CockroachDB provides scale without sacrificing SQL functionality. It offers fully-distributed ACID transactions, zero-downtime schema changes, and support for secondary indexes and foreign keys". Significantly, CockroachDB has had extensive design dedicated to surviving adverse network conditions (see Jepsen references in other posts)
- misterbowfinger 9y agoCan someone give a brief pros/cons between Cockroach DB Core and Google Cloud Spanner?
- bpicolo 9y agoOpen source vs not open source. Cockroach still in it's infancy vs spanner. I'm sure there are a variety of things here, but they mostly aim to solve a similar problem with a slightly different approach. Some of the big details relate to not requiring atomic clocks: https://www.cockroachlabs.com/blog/living-without-atomic-clocks/ https://www.cockroachlabs.com/blog/living-without-atomic-clo... Here's their comparison chart, though naturally it's biased for things-cockroach-does: https://www.cockroachlabs.com/docs/cockroachdb-in-comparison.html https://www.cockroachlabs.com/docs/cockroachdb-in-comparison... (I guess you can't write to Spanner with SQL? That seems like a big difference. No INSERT/UPDATE?)
- dianasaur323 9y ago[cockroachdb here] Thanks for the great response, bpicolo!
- greenshackle2 9y agoI'm confused. What's the difference between 'Yes' and 'Optional' in the 'Commercial Version' row on the comparison chart? To me 'Yes' suggests there is only a commercial version, but clearly that's not true for CockroachDB.
- dianasaur323 9y agoThanks for pointing that out! We will fix that to optional for us :)
- amq 9y agoCan someone explain how is/can it be better than MariaDB Galera or MySQL Group Replication?
- dis-sys 9y agoYou can't deploy your MariaDB Galera/MySQL Group Replication systems across the Pacific and then expect it to further scale from there.
- sergiotapia 9y agoIs Cockroach DB intended for just "big-data" companies? Would a small project run really well with Cockroach DB? Of course a small database probably won't need a lot of the unique features, but is this aiming to replace PG/MySQL in the small/mid-size projects?
- dianasaur323 9y ago[cockroachdb here] Yes! In addition to being highly scalable, CockroachDB also comes with built-in replication. That means that even with a smaller project that hasn't scaled yet, you still get the benefit of a more resilient database. Also, CockroachDB is super easy to install and get started with!
- eblanshey 9y agoI've come across many projects that are easy to get started with, but the main stuff to look for is in the details. Although MySQL might be easy to get into, for example, it takes time to learn the intricacies for query optimizations, and importantly, what to do when SHTF, like when a table gets corrupted. My question is, in your opinion, what does it take to become proficient in CockroachDB sufficiently enough to be comfortable using it in a high volume, high-uptime-required environment? Thanks.
- zokier 9y agoI can't speak for others, but at least for me the main attraction of CockroachDB is getting foolproof HA straight out of the box. That is something I think anyone can appreciate regardless of their dataset size. Note that I haven't actually ran CockroachDB yet, so I can't confirm if it really delivers on that promise, but I'm hopeful.
- raybb 9y agoWhat is HA?
- deleted 9y ago[deleted]
- niceperson 9y ago>Cockroach What were they thinking?
- dkersten 9y agoCockroaches are highly resilient creatures. The name, I assume, is alluding to the goal of this database being a highly resilient system. Whats the problem?
- camus2 9y agoCockroaches are disgusting, the name hurts to the product. I'm confident they will acknowledge that fact sooner or later and change the name.
- deleted 9y ago[deleted]
- johnwheeler 9y agoI think the name "Cockroach" was a really poor decision from a marketing standpoint. The team intended to convey durability, since cockroaches can live through anything. But when I think of a cockroach, I think, gross, disgusting, etc.
- scandox 9y agoIt's memorable. So if the product is really excellent and is needed by customers - then I think it could be a boon. I mean Mongo has very bad associations for me in terms of childhood taunts and Blazing Saddles...but now the name really relates more to the product than to the original meaning.
- gm 9y agoThe difference is that "mongo" does not have a universally-known meaning. Cockroaches and known throughout the world, and are disgusting throughout the world.
- chimeracoder 9y ago> I mean Mongo has very bad associations for me in terms of childhood taunts and Blazing Saddles...but now the name really relates more to the product than to the original meaning. The difference is that the word "mongo" is an issue of the same word having different meanings in different dialects. Whereas, with "cockroach", it's the same intended meaning, but with different connotations.
- ericb 9y agoCan Cockroach be plugged into a Rails app where mysql was? I'd be interested in hearing: - the backup story - the replication/failover story - horizontal scaling story (is it plug and play)
- irfansharif 9y agoNot mysql, but we've tested and recommend the Ruby pg driver and the ActiveRecord ORM[1] (CockroachDB supports the PostgreSQL wire protocol). It should be 'plug and play' insofar as you simply point to any node in the running cluster when setting up ActiveRecord::Base.establish_connection. As for our backup story, our doc page[2] on the subject should shed more details. [1]: https://www.cockroachlabs.com/docs/build-a-ruby-app-with-cockroachdb-activerecord.html https://www.cockroachlabs.com/docs/build-a-ruby-app-with-coc... [2]: https://www.cockroachlabs.com/docs/backup.html https://www.cockroachlabs.com/docs/backup.html
- arjunnarayan 9y agoI have ported a MySQL-based ActiveRecord Rails app that was somewhat complicated to Postgres, and then on to CockroachDB. It works pretty well, so I'd give it a go. We're also committed to supporting ActiveRecord via the Postgres connector, so if you run into any bugs, we would do our best to fix them. I am personally invested in ActiveRecord support myself. At this point ORM support on CockroachDB is driven mostly by usage so please try it! Your other questions are better answered on the blog post, but quickly: * CockroachDB core comes with a `dump` command to backup your databases. CockroachDB Enterprise has blazingly fast _incremental_ cloud backup and restore, the kind that you might want for a very large deployment. * Replication is managed under the hood by sharding the data into many ranges that are each 64mb in size. Each range is replicated using Raft, and if a node goes down, the other replicas scattered across the cluster seamlessly take over and upreplicate a new replica to "heal" the cluster. * The horizontal scaling is indeed plug and play - just add more nodes to the cluster and they'll automatically rebalance replicas across the cluster with no downtime and no additional configuration.
- brightball 9y agoHow does it compare to Couchbase with it N1QL?
- ansible 9y agoThe main difference is the consistency model: https://blog.couchbase.com/10-things-developers-should-know-about-couchbase/ https://blog.couchbase.com/10-things-developers-should-know-... Whereas CockroachDB aims to be strongly consistent. This makes life for the application developer much easier.
- Perignon 9y agoName still sucks and is disgusting af.
- socmag 9y agoClocks are meaningless under load. The higher frequency the transactions the more you get into quantum physics. In reality, nobody cares if T-Mobile debited your account 0.01ms before WalMart. [edit] what is important is isolation and consistency of the transactons.
- socmag 9y agoInstead of just downvoting, how about refuting my claim? I'm seriously curious what is the disagreement. These guys already established atomic clocks are unnecessary. Very interested in which use cases require them.
- deleted 9y ago[deleted]
- theptip 9y agoRead and learn: https://research.google.com/archive/spanner.html https://research.google.com/archive/spanner.html Serializability is all about ensuring a single consistent ordering of events. Lots of algorithmic shortcuts you can take if all your nodes' clocks are precisely in sync.
- socmag 9y agoI'm very familiar with the literature since I'm a distributed database developer. If you investigate high frequency trading you will understand that the quantum phenomena that I'm talking about is not just me high on mushrooms but a real world thing. The only "time" relevant is the time when the cluster agrees an atomic, isolated transaction is time to commit from its own perspective.
- nocman 9y agoAm I wrong in remembering that the HN guidelines used to say that you should not downvote someone's comment simply because you disagreed with it? I went looking, and I don't see that in the current guidelines. I could be wrong about it being there before, but I was almost certain that it was at one point. Seems like it used to say that you should only downvote comments that you think don't contribute anything of value to the conversation. Just curious, because it seems to me that for quite a while now there have been a lot of comments that appear to get downvoted just because people don't agree with what the person said (and often there are no responses to counter, the person just gets downvoted).
- ccallebs 9y agoFirst, this is awesome! Congrats to the team for reaching this milestone. Secondly, I think the name is memorable and conveys exactly what it should. If I were ever on an engineering team that chose not to use CockroachDB due to being "grossed out" by the name, I wouldn't be on that engineering team for long. Perhaps someone can explain the knee-jerk reaction to it for me.
- nathan_f77 9y agoI hate the name, and that does put me off from experimenting with it. But I would use it if it was the right tool for the job.
- gervase 9y agoI might be an interesting case. I had previously been a big supporter of their name, agreeing with some other posters that it promotes the durability of the system. However, after a move last year, I was forced to live with cockroaches for approximately 6 months, after never encountering them prior to that. Since then, I've completely switched camps. Can't see the name without being skeeved out. The reality of cockroaches is so absolutely repulsive that it completely changed my view 180º. I moved out of that place in November, and haven't seen once since; I'm curious if my aversion will fade over time.
- 5ourpu55 9y agoI'm amazed at how visceral a reaction the term "cockroach" receives haha Though I AM a country kid, so maybe I'm just a bit less squeamish when it comes to these things.
- v3ss0n 9y agoCongrats Ben Darnell and team! I am fan of his work on Tornado web server!
- bdarnell 9y agoThanks v3ss0n!
- anthonylebrun 9y agoSince there's a little side riff about the name going on I thought I'd throw in my 2 cents. Personally I love the name. I think it does a great job of conveying the spirit of the project and provides unlimited pun opportunities. Plus it's memorable, just like a real life roach encounter. Unfortunately I'm sure some people will discriminate against your DB on the basis of name alone. That's ludicrous, but that's our species for ya.
- dboreham 9y agoAt first when I saw yet more name comments on this thread I felt disappointment that people can't leave the subject alone. But then I realized that as someone who doesn't care about the name, even positively enjoys it, I have a competitive advantage over those people. Now I feel good again.
- api 9y agoChoosing technologies based on first-hand review and first principles rather than things like Gartner magic quadrants, big company brand recognition, feature lists, and "serious" sounding names is a competitive advantage that startups often have over big businesses. The latter are forced by their procurement departments and other forces to use old, inferior, and more costly technology. On the flip side though if I were in charge of CockroachDB I would look at doing something about the name. Maybe rename it something like "Resilient" as part of the "exit from beta" milestone. It's going to be a serious liability for them selling to the kinds of customers I described above, and unfortunately that's where most of the money is in these devops/infrastructure markets. The key to success is to make a superior product and then figure out how to sell it to pointy haired bosses. The latter often means making it look more boring than it actually is. Fun factoid: scientists sometimes do this with grant proposals. I've had two scientists independently tell me that they often take cool, fascinating research proposals and "make them boring" to sell them to bureaucrats. "You have to hide all the interesting stuff and make it sound like you are doing boring incremental research. If you talk about anything 'revolutionary' you will never get funded."
- thraway2016 9y ago
- wmfiv 9y agoAre there published benchmarks for multi-key operations and more complex SELECT statements? I apologize if I missed them. I'm trying to determine whether there's a place for Cockroach within what I think are the constraints in the database space. * Traditional SQL Databases - Go to solution for every project until proven otherwise. - Battle tested and unmatched features. - Hugely optimized with incredible single node performance. - Good replication and failover solutions. * Cassandra - Solved massive data insert and retention. - Battle tested linear scalability to thousands of nodes. - Good per node performance. - Limited features. It seems like many new databases tend to suffer from providing scale out but relatively poor per node performance so that a mid-size cluster still performs worse than a single node solution based on a traditional SQL database. And if you genuinely need huge insert volumes, because of the per node performance you'd need an enormous cluster whereas Cassandra would deal with it quite comfortably.
- arjunnarayan 9y ago[Cockroach Labs engineer here working on performance benchmarking] We have load generators for YCSB (just raw key-value ops in a firehose) and TPC-H (very complicated read-only queries) running right now, and we're about to start running TPC-C queries (moderately complex queries in large volume) as well. You can follow along on our progress here: https://github.com/cockroachdb/loadgen https://github.com/cockroachdb/loadgen In the context of your dichotomy, we want to bridge that gap. We want the linear scalability of your second group along with the full feature-set of the first group. We will be publishing our performance numbers, but we haven't so far because the product has improved rapidly, and our numbers have been quickly obsoleted, but rest assured, we will be publishing a series of blog posts very soon. Anecdotally, our beta customers are not finding that they need very many more CockroachDB nodes than their existing database solutions, even with something as high-performant (but inconsistent) as Cassandra.
- wmfiv 9y agoThat's great. Thanks for the response and I'll keep an eye out for the blogs.
- deferredposts 9y agoIn a couple of years, I suspect that they will rebrand their name to just "RoachDB". It conveys the same meaning, while not being that awkward to discuss with users/clients
- chimeracoder 9y ago> "RoachDB". It conveys the same meaning, while not being that awkward to discuss with users/clients "Roach" has some other connotations as well[0], which may not help with selling to larger and enterprise clients. [0] https://www.urbandictionary.com/define.php?term=roach https://www.urbandictionary.com/define.php?term=roach
- deleted 9y ago[deleted]
- api 9y agoAbout nine months ago we made the decision to go with RethinkDB for our infrastructure in place of PostgreSQL (at least for live replicated data), but if this existed at the time we'd have seriously taken a look. We're pretty happy with RethinkDB but I plan on still taking a look at this so we have a backup option.
- dianasaur323 9y ago[cockroachdb here] We are big fans of RethinkDB, but also glad to hear that you'll explore CockroachDB. Let us know how it goes, and definitely file any issues / feature requests in our GitHub repo!
- wtf_is_up 9y agoDoes CockroachDB have a streaming API a la RethinkDB changefeeds? This is a killer feature, IMO.
- arjunnarayan 9y agoNot yet, but it's on our roadmap.
- ralusek 9y agoJust out of curiosity, do you mind elaborating a little bit on why not? It strikes me as something that would be very easy to implement in a database, is there a reason why so few databases have a mechanism to do this? If it's about maintaining an open connection in order to notify the client, that part makes sense, but at the very least the changefeed itself should be toggleable and easy to query in any DB.
- state_machine 9y agoOne of the challenges for us in implementing something like LISTEN/NOTIFY comes from our distributed nature: since a table is likely broken up across many nodes, you somehow need to aggregate changes from all of them back into a single change feed wherever the listener is, and in such a way that it doesn't create a single point of failure.
- MichaelBurge 9y agoIt probably scales but how is the performance? If I need to load a couple billion rows and do a dozen joins in some analytics, is that one machine, a dozen, or 100? Is it more for web apps, analytics, or what? When would I consider switching from e.g. Postgres to CockroachDB?
- arjunnarayan 9y ago[Cockroach Labs engineer here] For just a couple billion rows and a dozen joins, a single node will suffice (with the caveat that you really want at least 3 nodes because CockroachDB is built for replication and fault-tolerance and you're not getting that with a single node cluster), but you'll get linear speedup as you add more machines. Your performance on a single node should be on the same order of magnitude as doing this in Postgres right now. We are rapidly closing that gap, and intend to close it completely for TPC-H style queries, while retaining the linear performance speedup with more nodes. The reason this gap isn't already closed is we've been focused on transactional performance in distributed, fault-tolerant situations rather than analytics performance, for 1.0. There are lots of optimization low hanging fruit that we haven't focused on in analytics scenarios that we are just getting started on.
- MichaelBurge 9y agoThanks for your response. It sounds like CockroachDB might be an alternative to setting up an RDBMS for read replication once you need many connections.
- gflarity 9y agoHi Cockroach Labs Engineer here, On the feature FAQ joins are describe as 'functional' which doesn't inspire a lot of confidence but maybe it's just a perception thing. What exactly does functional mean? A SQL db without joins sounds a lot like just a NOSQL db with a familiar query dialect.
- arjunnarayan 9y agoIf you are using Joins in an OLTP setting, everything should work absolutely as you might expect. "Functional" is our caveat that if you run Joins across your data in an OLAP setting, it will work, but it may not be the most performant Join possible. For example, our query planner does not currently plan Merge-joins even if the appropriate secondary indices exist. So after a point (joining ~billions of rows of data) it no longer is as performant as it could be. Now we expect to roll out this particular fix within 6 months. However, optimizing 4 or 5-way nested Joins in OLAP-cube style settings isn't something we're going to be performant at for years. We need a lot more infrastructure built up before we start solving the kinds of problems revealed by, say, the Join Order Benchmark paper (http://www.vldb.org/pvldb/vol9/p204-leis.pdf http://www.vldb.org/pvldb/vol9/p204-leis.pdf).
- therealmarv 9y agoDoes this work theoretically interplanetary (just asking because for science) ?
- arjunnarayan 9y agoNo. Once your latency goes beyond single digit seconds, performance will probably collapse. Too many subsystems would time out. in theory it could be made to work (with terrible performance, and extremely long commit-waits due to having to wait until the remote planets get back to you), but I wouldn't architect a planetary spanning distributed database this way. We probably would have to go back to the drawing board and start from scratch.
- therealmarv 9y agoThanks for the long answer. Much appreciated. The question came into my mind when reading some of graphics and specifications.
- pishpash 9y agoYou'd need to give up on consistency, because there is no such thing when the time of communication is long compared to interval of events. In the long run, ACID is dead.
- state_machine 9y ago[cockroachdb employee] Short answer: no. Long answer: at their closest earth and mars are about 54m km apart, at the furthest it's over 400, with an average of around 225m km, so theoretical latency is varies between 4 and 24 minutes. CockroachDB uses synchronous replication via raft, and that latency would cause problems as would some other setting like our window sizes and their interaction with timeouts.
- politician 9y ago> CockroachDB uses synchronous replication via raft Deep space aside, I wish the announcement just said that! I came back to HN for insight into the paragraph about "multi-active availability... an evolution in high availability from active-active replication". Marketing... sometimes... I tell you what.
- v_elem 9y agoIt looks like there is still no mechanism for change notification, which in our particular case is the only missing feature that prevents using it as a postgresql replacement. Does anybody know if this feature is planned in the short or medium term ? https://github.com/cockroachdb/cockroach/issues/6130 https://github.com/cockroachdb/cockroach/issues/6130 https://github.com/cockroachdb/cockroach/issues/9712 https://github.com/cockroachdb/cockroach/issues/9712
- arjunnarayan 9y agoThis feature is planned, but I cannot give you a concrete timeline. We want to do this right, and we need other parts in place to do this with high performance, in a transactionally consistent fashion, in the face of high contention, and for arbitrarily complicated "views". I will say that this is the single feature that I personally am most invested in at the company, so it will happen.
- triangleman 9y agoName doesn't bother me. It's memorable and I'd definitely consider using it, whether in a startup or enterprise. Better than "Postgres" -- how do you even pronounce that?
- bfrog 9y agoShould've gone with tardigrade instead as a name, those little bastards can live in space!
- Svenskunganka 9y agoPardon the nature of my question, but I'm really interested in what your experience has been so far building a database with Go? Has its runtime (the GC for example) posed any issues for you so far? Looking at other RDBMS's, languages with manual memory management like C or C++ seems to be the go-to choice, so what were the reasons you chose Go? I'm quite frankly amazed that Go's runtime is able to support a database with such demanding capabilities as CockroachDB!
- arjunnarayan 9y agoWe have a post on why we chose Go, from a year and a half ago: https://www.cockroachlabs.com/blog/why-go-was-the-right-choice-for-cockroachdb/ https://www.cockroachlabs.com/blog/why-go-was-the-right-choi... More technically, here's a somewhat random set of thoughts on the subject: The Go GC is performant and predictable, unlike the JVM GC. We do have some very memory-allocation-conscious code patterns to minimize the performance impact of working in a garbage-collected language runtime, but in the end it's not as bad as you might expect if your expectations are coming from the JVM world. Library support is good. To quote our CEO, "Most of us on the team have done extensive work with C++ and Java in the past. At Google, C++ was the standard for building infrastructure and there are a lot of good reasons for that. It's fast and predictable. It would be a good choice for Cockroach, except that in the world outside of Google, in open source land, the supporting libraries for C++ are either terrible, incredibly heavyweight, or non-existent. We didn't want to rebuild everything which you take for granted at Google from scratch. It turns out that Go has many of the necessary libraries, and they're straightforward and very well written." Basically, if Google's internal C++ libraries, tooling, style guides (and the tooling to enforce them) were available externally, we might have gone with C++. Some of us are fans of Rust, but Rust sadly did not exist in a stable state when CockroachDB started. I'm not sure we would pick Rust were we to start today (tooling is still a concern there), but it would certainly be part of the discussion. The native support for concurrency in Go is a huge plus. We use thousands of goroutines in CockroachDB, and that's been a huge blessing. I can answer any more specific questions if you have them.
- Svenskunganka 9y ago
- sixdimensional 9y agoHow does Cockroach efficiently handle the shuffle step when data is on many nodes on the cluster and has to move to be joined? Does Cockroach need high capacity network links to function well? I always see companies making the claim of linear speedup with more nodes but surely that can't be the case if the nodes are geographically disjointed over anything less than gigabit links? Perhaps linear speedup with more nodes is only possible over high speed connections? How high is that exactly? Congratulations to the team on the release! Introducing this kind of database is no easy task - thank you and great job, keep up the good work!
- arjunnarayan 9y agoThe short story is we do need high capacity network links to function well. By "high capacity" I mean at least double digit megabit links between your datacenters. A query that inherently requires shuffling because the data is geographically distributed can't get past the bandwidth needs of performing the shuffle. At the very least, with the literal simplest query plan, you're going to need all the raw data to be transported to a single node/datacenter, and I doubt there's a query and network setup where that's more efficient than doing networked shuffles themselves. I don't think you need gigabit networks, but you're certainly going to want at least 10 megabit links. We have not tried to benchmark scenarios where we are bandwidth constrained, so I can't tell you precisely what the minimums are. All the cloud scenarios we've tested (on GCE, Azure, AWS, DigitalOcean) are constrained on other dimensions (i.e. CPU cores, memory, disk IO). And thank you :)
- sixdimensional 9y agoThat makes sense- I think part of the reason such types of databases are well suited to cloud operations is the guaranteed throughput of the cloud providers own network backbone, which is almost impossible for any single "regular" organization to match, at least for the price. I think we are at a point where doing business without the cloud will become nearly (but not completely) impossible at huge scale with all these features. Thank you very much for your detailed answer and good luck with the continued rollout!
- 9y ago
- daliwali 9y agoCockroachDB looks like a great alternative to PostgreSQL, congrats to the team for doing so much in such a short time. The wire protocol is compatible with Postgres, which allows re-using battle-tested Postgres clients. However it's a non-starter for my use case since it lacks array columns, which Postgres supports [0]. I also make use of fairly recent SQL features introduced in Postgres 9.4, but I'm not sure if there are major issues with compatibility. [0] https://github.com/cockroachdb/cockroach/issues/2115 https://github.com/cockroachdb/cockroach/issues/2115
- hd4 9y agoI'm basically here to ask a similar question, whether this is aimed as an modern alternative to Postgresql, since they don't clearly state this on the OP news announcement.
- tyingq 9y agoTo me, at least for now, it seems more like a SQL enabled etcd or similar. They aren't currently claiming performance numbers that make it sound suitable for general purpose relational database scenarios. A SQL aware etcd like thing has a lot of appeal though, and I assume the performance work is coming.
- jordanlewis 9y agoI'm an engineer on the SQL team at CockroachDB. We're very aware of our missing support for array column types - and in fact beginning to add support for arrays is one of my team's priorities for the next release cycle. What kind of other recent SQL features introduced in Postgres 9.4 do you use? Postgres has a ton of features, as I'm sure you're aware, and while we strive for wire compatibility with Postgres it's not a goal of ours to implement support for every Postgres feature out there.
- octernion 9y agowhat's the story for change data capture with CockroachDB? Postgres 9.4 added logical replication, which is incredibly useful for this use case. Also, we use JSONB fairly extensively -- I see the tracking issue here https://github.com/cockroachdb/cockroach/issues/2969 https://github.com/cockroachdb/cockroach/issues/2969 but no movement.
- newsat13 9y agoVery disappointed with HN turning into a 4chan/reddit style trolling board about the name. Guys, we get it that you don't like the name. Can we please stop bike shedding and move on? The people at cockroachdb have obviously seen all your messages but decided it's worth keeping the name. What more is there to talk about? Why not talk about the relative technical merits of this DB?
- SilasX 9y agoIt's not bikeshedding when the bikeshed's color will actually have concrete effects on adoption. Most people -- i.e. in procurement, management, finance, and others you need to appeal to -- don't want anything to do with cockroaches. The idea disgusts them at a gut level, not something you can talk away. HN users are giving vital advice, for free. Those who ignore it will have only themselves to blame. As I say every time this comes up, would you be so dismissive about critics of naming a product PubesDB? Or GonorrheaDB? Or [n-word]DB? Then you agree that disgust-invoking connotations of the name matter, and we're just haggling over the details. Ubuntu, Mongo, Swagger (edit: Hadoop also) ... they're weird, sure, but they don't evoke the visceral feeling of disgust that cockroaches do.
- networked 9y agoOn the other hand, if I heard of a database called "CockroachDB" gaining ever-greater adoption, I'd pay close attention to it because it was clearly succeeding despite a marketing handicap.
- thraway2016 9y agoprocurement, management, finance, and others you need to appeal to They don't need to appeal to any of these suits. Just the technical decision-makers, whose express job it is to choose solutions on their technical merits, not their spurious emotional reactions.
- siculars 9y agoYou can't possibly think that's true. If you do you can't have had much experience in buying or selling technology.
- singularjon 9y agoHow does the speed compare to that of Postgresql and MongoDB?
- gog 9y agoSlightly offtopic, but what do you use for your blog and documentation pages?
- toddmorey 9y agoThere was a great session with Spencer Kimball (CockroachDB creator) and Alex Polvi (CoreOS) at the OpenStack Summit. It's a good overview and demo: https://youtu.be/PIePIsskhrw https://youtu.be/PIePIsskhrw
- irfansharif 9y agothere's a second part to this presentation[1] running cockroachdb across 16 (!) cloud vendors. [1]: https://www.youtube.com/watch?v=nBXXLNIwAoo https://www.youtube.com/watch?v=nBXXLNIwAoo
- v3ss0n 9y agoWill there be a rethinkdb style REALTIME Changefeed or PostgreSQL's Listen Notify ?
- ralusek 9y agoI'd also like to know this. PG notify and triggers in general. Any equivalent to DB link?
- state_machine 9y agoYes, change feeds (and triggers) are on the roadmap (though not yet in active development).
- v3ss0n 9y agoWill there be a rethinkdb style REALTIME Changefeed or PostgreSQL's Listen Notify ?
- whatnotests 9y ago/me forks the damned repo, renames it, wins the Internet.
- knz42 9y agohttps://github.com/tschottdorf/bikesheddb https://github.com/tschottdorf/bikesheddb
- vtomasr5 9y agoI think this is the DB Project of the year in the open source community. Cockroachlabs has done an incredible effort to develop and test a new Database and these guys are giving it for free (I read about the series B raise too ;)), for us to use it. Thanks for doing this. You're very much appreciated. (BTW I love the name and the logo!!)
- sandstrom 9y agoI think it's an excellent name! Also, biologists would argue that cockroaches is a magnificent creature, highly adaptable and very fit (in 'survival of the fittest' terms). I would pay for and deploy a cockroach db — because of its name.
- apognu 9y agoI've been following CockroachDB for quite a while. Great job on 1.0. I've had a question for quite some time though (and I think there is an RFC for it on GitHub): do we still need to have a "seed node" that is run without the --join parameter, or can we run all the nodes with the same command line, with the cluster waiting for quorum to reconcile on its own?
- bdarnell 9y agoCurrently, you need to run one node without --join for the initial bootstrapping (as soon as this bootstrapping is complete, you can and should restart it with --join to get everything into a homogenous configuration). I was hoping to make some changes here so you could start every node with --join from the beginning, but it was trickier than anticipated so it didn't make the cut for 1.0. Watch for improvements here in a future release.
- apognu 9y agoThank you for your answer. That's okay, for now, I run a simple StatefulSet where each pod checks whether the Service is reachable on port 26257 to determine if it should join or init the cluster. It's not as nice as if it was handled by Cockroach itself, but it does the job.
- bdarnell 9y agoThis bootstrapping problem is tricky. We publish kubernetes templates at https://github.com/cockroachdb/cockroach/tree/master/cloud/kubernetes https://github.com/cockroachdb/cockroach/tree/master/cloud/k... that contain our current best solution for the join/init problem.
- acd 9y agoCongrats to bringing out 1.0 bern following the project and look forward to try it out!
- raarts 9y agoOn a three node cluster will it survive two nodes going down?
- irfansharif 9y agoshort answer: nope. cockroachdb replicates data for availability and in order to guarantee consistency across the replicas, it uses Raft[1] internally. Raft necessitates a majority of the replicas remain available in order to operate. it ensures that a new 'leader' for each group of replicas is elected if the former leader fails, so that transactions can continue and affected replicas can rejoin their group once they're back online. [1]: https://raft.github.io/raft.pdf https://raft.github.io/raft.pdf
- novembermike 9y agoWhat are the recommended configurations then? If I want to survive multiple node failures could I have 9 replicas?
- irfansharif 9y agoraft is premised on overlapping majorities, so to speak. in order to tolerate up to `n` node failures you'd need to run `2n + 1` instances (for nine nodes you'd tolerate up to four node failures).
- xmichael99 9y agoNow if we could get a 1.0 of TiDB ???
- ngaut 9y agoAlmost there.
- rantanplan 9y agoIn an era where hot air and hip DB technologies prevail, I'd like to emphasize the fact that the CockroachDB engineers are consistently honest and down to earth, in all relevant HN posts. This builds up my confidence in their tech, so much so that even though I had no real reason to try this new DB, I'm gonna find one! :D
- nicwagenaar 9y agoExactly! The confidence that the devs inspire by taking the time to explain the choices behind the tech, makes me want to find a project to test it out on.
- nhumrich 9y agoDoes the replication work cross-region, say US-East and US-West? or even cross continent? It sounds like the timing requires very short latency and might not work in these scenarios
- d4l3k 9y agohttps://forum.cockroachlabs.com/t/cross-region-sync-replication-deployment-of-cochroahdb/415 https://forum.cockroachlabs.com/t/cross-region-sync-replicat...
- dis-sys 9y agoJepsen test results basically show that latency caused by replica distance won't screw your data. On the other hand, clock drift can stop your system, or even potentially corrupt your data, depending on how fast such incident can be detected/handled and what is your workload/what you are doing.
- a-robinson 9y agoYes, it works. Your latency will just be correspondingly higher (due to the speed of light). We are constantly testing a cross-region (i.e. US-East and US-West) cluster and have periodically run tests on cross-continent clusters (US to Asia-Pacific). In these cases you can help the cluster out by following some of the advice on the "Recommended Production Settings" page (https://www.cockroachlabs.com/docs/recommended-production-settings.html https://www.cockroachlabs.com/docs/recommended-production-se...) around specifying which `--locality` each node is in.
- bish2 9y agoI'm struggling to understand how this company has raised $50 million dollars when db companies with paying customers like RethinkDB and FoundationDB had to shut down. They are gonna earn back $50 million by selling...a backups tool?
- swsieber 9y agoI think one major difference is that it's a drop in replacement for certain SQL products, plus a major selling point of NoSQL - good horizontal scaling. RethinkDB and FoundationDB are great, but require a paradigm shift I think.
- swaraj 9y agoFree open source ops tools + enterprise support is a pretty solid business. For recent-ish DB companies see Mongo, Elastic, Redis, MemSQL, etc. I'm excited to track this project!
- daxfohl 9y agoCurious why Mac is better supported than Windows. This is obviously something you'd run on a server. Do orgs run Mac servers? Is it just to support dev work for people too lazy to launch a VM? Sorry, Windows/Linux ops person here with very little awareness of Mac ecosystem.
- jpgvm 9y agoIt's not so much a matter of Mac > Windows but rather Mac+Linux+*nix > Windows. This just comes down to the fact that Windows is a special snowflake that does everything differently. Sometimes for good reasons, but usually not for good reasons.
- gred 9y agoVery interesting. I have to admit I've seen the product name a few times, but never took the time to have a look. I do have a few questions, though, if any of the engineering team are still around watching the discussion :-) From the high availability page [1] in the docs: > Cross-continent and other high-latency scenarios will be better supported in the future. Do you have a specific timeline in mind? I've been working on an application that needs to be highly-available, and which uses Oracle right now. It seems like you can add all sorts of tools to the mix (RAC, DataGuard, etc), but there are always significant caveats around the capabilities of the resultant system. We're talking 1 to 2 TB of data total, tables of up to 100 million rows with 1 million rows added per day, distributed across three data centers (US, EU, Asia). And regarding high availability in the context of application deployments, is there any documentation on the locking characteristics of DDL statements? I'm interested in the ability to modify the schema during an application deployment without having to bring down the system or implicitly locking users out. Apologies if I missed it somewhere on the website! [1] https://www.cockroachlabs.com/docs/high-availability.html https://www.cockroachlabs.com/docs/high-availability.html
- radub 9y agoI don't have a specific timeline but it is something we will be focusing on in the following releases. Regarding DDL statements, this blog post [1] has details. In a nutshell, online schema changes are possible; the changes become visible to transactions atomically (a concurrent transaction either sees the old schema, or the fully functional new schema). [1] https://www.cockroachlabs.com/blog/how-online-schema-changes-are-possible-in-cockroachdb/ https://www.cockroachlabs.com/blog/how-online-schema-changes...
- ncrmro 9y agoAny support for postgres trigram searches?
- doanerock 9y agoSay you scaled up to 100 nodes for the holiday season, is there any way to tell how many/much storage/nodes you have to keep running in order to keep 3 backups and maintain your new post holiday load?
- BramG 9y agoWe don't have any auto scaling for either up or down scaling, but if you're using a deployment tool such as Kubernetes, I don't see why it wouldn't be fairly easy. And it might be a good idea to add a message in the admin UI if you all of your nodes are experiencing a high load. By just looking at your max load over the last 24h or perhaps week, it would be pretty easy to see when to down scale. That being said, as long as you remove the cockroach nodes one at a time , it's pretty easy to down scale a cockroach cluster.
- doanerock 9y agoSince CockroachDB is Eventually Consistent Reads then how would that affect my SaaS multiuser application? How long on average would I have to wait for them to become Consistent?
- a-robinson 9y agoCockroachDB reads are strongly consistent, not eventually consistent. You don't have to wait at all.