11 ms·
Call Me Maybe: MongoDB Stale Reads
- chatman 11y agoApache Solr has done very well at Jepsen tests.
- sylvinus 11y agoAnother instance of Kyle's amazing research! You may want to catch him on stage with other great minds at dotScale on June 8: http://dotscale.io http://dotscale.io
- jamescostian 11y agoI'm so glad to see the Jepsen series re-instated. Thank you so much Stripe
- bkeroack 11y agoIf you are a database author and you get a bug report from Kyle, spend a long time thinking about it before closing the issue as invalid.
- BinaryIdiot 11y agoFor context (for those who didn't read the whole article): https://jira.mongodb.org/browse/SERVER-17975 https://jira.mongodb.org/browse/SERVER-17975
- erewh0n 11y agoIndeed, database vendors should aim to have their software "Jepsen certified".
- cheeseprocedure 11y agoAgreed! As an example, HashiCorp included Jepsen test results in their Consul documentation: https://consul.io/docs/internals/jepsen.html https://consul.io/docs/internals/jepsen.html
- takeda 11y agoAs Kyle himself mentioned (I think it might even be in relation to this document), Jepsen cannot prove that your software is safe, only prove that it isn't. He even called them out for modifying timeout value to pass the tests[1]. [1] https://aphyr.com/posts/316-call-me-maybe-etcd-and-consul https://aphyr.com/posts/316-call-me-maybe-etcd-and-consul
- GordyMD 11y ago> Jepsen cannot prove that your software is safe, only prove that it isn't. This doesn't just apply to Jepsen but to all software tests. You only ever test a finite set of scenarios, so you can't really ever guarantee your software is 'safe'/bug-free - only that it does not fail in common/expected scenarios.
- tel 11y agoDepending on how large "software tests" is in your mind this can easily be not the case. Kyle often mentions the tool TLA+ which enables complete and total formal analysis of certain systems. This can be a test which is complete and therefore a positive proof of correctness. It's also not difficult to test smaller components of your system if you note that the state space here is small enough to be exhausted. This is a very important thing to look in to in serious, important components of a software system. Finally, though it's futuretech today still, dependent types offer a general system for embedding proofs of correctness directly into code thus elevating positive proof to a computational artifact just like any other.
- hibikir 11y agoType systems and proofs have their limitations, if just because for a proof to be worth anything, I need a perfect description of what I want, and that doesn't make any sense outside of trivial cases. I once worked with a relatively well known, prove everything developer that you might know by name. He's written books and everything. He built an algorithm to make a distributed system of equal peers figure out when it needed to start more nodes, or could shut some of them down. He wrote a proof, in code. He wrote a paper. When in production, the system would not work as advertised, and he blamed it on other pieces, because the algorithm was proven correct! So the problem stayed there for months. After he left the company, I decided to figure the problem out, so I read through the proofs, the paper, and the code: All the single letter variables you could possibly want. I figured out that yes, the algorithm was flawless, as long as every operation in the system was atomic and instantaneous. Instead of the proof, I built a small simulator that didn't have such flawed assumptions, and got the exact same behavior as the production system. So the proof was perfect, as long as we made assumptions that are impossible in our universe. And the entire algorithm was less than 200 lines of code. So whenever we have a reality that is difficult to model (and let me tell you, distributed systems fit the bill), dependent types will not save you, haskell or no haskell. Proofs will always be limited by your assumptions. So while the tools you mention are nice. They hit the same limits as everything else we build. Whether to write a proof in idris, use generative testing, or just some example testing, is really all a tradeoff, but you will never escape from bad specification, as all specifications are bad.
- jodah 11y agoThey're not the only ones: https://github.com/elastic/elasticsearch/issues/2488#issuecomment-89509176 https://github.com/elastic/elasticsearch/issues/2488#issueco...
- sciurus 11y agoIt looks like that specific issue was closed because the other problems reported are being tracked in other issues.
- craigching 11y agoActually elasticsearch took their jespen results pretty seriously and even have an ongoing status of their resiliency. See my post on this: https://news.ycombinator.com/item?id=9418318 https://news.ycombinator.com/item?id=9418318
- hueving 11y agoEspecially since many people will judge the competence of the engineering based on the response. It really makes me nervous when a database company doesn't even understand the problem for many days and then pretends like it's expected behavior afterwards even though it flies in the face of their documentation, advertising, and tech talks.
- pje 11y agoupvoted for the Look Around You link alone.
- deleted 11y ago[deleted]
- threeseed 11y agoShame this wasn't done with the latest version 3.0. Although given that improvements are scheduled for 3.1 I would imagine it might be still an issue. Nice writeup either way though. Would like to see a similar article for Couch* and MySQL.
- jxf 11y agoThe most interesting lessons from the Jepsen series: * You should never trust, and always verify, the claims made by database manufacturers. * Especially when those claims relate to data integrity. * Super-especially when every safety level provided by the manufacturer that includes the word "SAFE" is actually unsafe.
- threeseed 11y agoActually the broader lesson would be to assume the worst in your application layer and try and remediate/verify wherever possible. If you look at his articles: Redis, PostgreSQL, Cassandra, ElasticSearch etc all had data consistency errors. And none of those have vendors making any claims. It's pretty sobering to say the least.
- chatman 11y agoApache Solr's jepsen tests have had very good results.
- ak4g 11y agoUm, this is the postgres article: https://aphyr.com/posts/282-call-me-maybe-postgres https://aphyr.com/posts/282-call-me-maybe-postgres There were no acknowledged writes lost. The only unacked-but-successful writes resulted from a connection while a commit ack was in-flight. That doesn't qualify as a data-consistency error, it means the client has to check if the data is present after reconnecting. But in no cases would the client reconnect to find that there were acknowledged-as-committed records that were missing or stale. In no cases would the client find that responded-as-rolled-back data was actually committed. This is very, very different than what is seen with MongoDB.
- hendzen 11y agoI don't think the results of Aphyr's MongoDB and postgres experiments are directly comparable. In the OP, MongoDB was run in a 5 node replicated configuration. In the post you reference, the experiment was run against a single postgres node. Furthermore, the postgres experiment only checked that no writes were lost. As Aphyr acknowledges, MongoDB did not lose any writes with "majority" write concern. The postgres experiment did not include the verification of a linearizable history of reads, which is what the bulk of the OP is about. I'd like to see a similar experiment run against a replicated postgres configuration with auto-failover.
- deleted 11y ago[deleted]
- dantiberian 11y agoThere's a lot going on here, but the summary is: "What Mongo actually does is allow stale reads: it is possible to execute a WriteConcern=MAJORITY write of a new value, wait for it to return successfully, perform a read with ReadPreference=PRIMARY, and not see the value you just wrote." https://jira.mongodb.org/browse/SERVER-17975?focusedCommentId=892980&page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel#comment-892980 https://jira.mongodb.org/browse/SERVER-17975?focusedCommentI...
- posnet 11y agoWould the use of wired tiger as a storage engine affect these results?
- misframer 11y agoMy guess is no. This is more about the behavior of the database as a distributed system and not just a storage engine.
- tupshin 11y agoDefinitely not
- victorhooi 11y agoHmm, I can't think of any reason why it would make a difference in this case.
- cpks 11y agoPeople really underestimate the value of Occasional Consistency. Occasionally Consistent databases, like MongoDB, are great for approximation algorithms, sublinear time algorithms, and similar applications.
- jdcryans 11y agoThe issue isn't that MongoDB is eventually consistent, it's that the documentation claims that in some cases it's strictly consistent[1] while Kyle found that: "MongoDB, even at the strongest consistency levels, allows reads to see old values of documents or even values that never should have been written." 1. http://docs.mongodb.org/manual/reference/glossary/#term-strict-consistency http://docs.mongodb.org/manual/reference/glossary/#term-stri...
- cpks 11y agoI didn't say Eventual Consistency. I said Occasional Consistency. MongoDB is has hard Occasional Consistency. Indeed, it is the most occasionally consistent database I know of. I once wrote a few million records into Mongo. It was consistent before the write, but never again after. Great for sub-linear time algorithms! At that point, all my algorithms ran at less than O(n) on the size of the data I had written in. From a business perspective, Occasional Consistency is also a very nice property if you are storing audit data for certain types of organizations. It gives complete plausible deniability about rule compliance.
- jdcryans 11y agoHeh, Poe's Law, etc :) I guess you could also call it Quantum Consistency.
- digitalzombie 11y agoOccasional Consistency sounds like a made up word. I can't even google this term without it auto correcting to Eventually Consistency. The wiki article, https://en.wikipedia.org/wiki/Consistency_model https://en.wikipedia.org/wiki/Consistency_model, doesn't even have such a term. Doing a hard search term on google reveals nothing.
- lobo_tuerto 11y agoI think it would be great to see one of these done for RethinkDB :)
- timmaxw 11y agoRethinkDB engineer here. RethinkDB currently doesn't support automatic failover, so this test couldn't be performed for RethinkDB yet. But when we implement automatic failover we're planning to test it against Jepsen. That will probably be sometime in the next few months.
- tootie 11y agoDB engineers living in mortal fear of Kyle is where we want to be.
- chucksmart 11y agoMaybe we should listen to Larry Ellison when he say "gimme my money!"
- Roboprog 11y agoHas Oracle been run through this test battery? Or would publishing the results of doing so bring down an army of Larry's lawyering henchmen? "You violated the EULA, now you must pay! You will only wish you were dead when we are done with you, bwahahhahahaha!"
- Roboprog 11y agoI guess the first rule of "Benchmark Club" is that we don't talk about Benchmark Club? Now off to deface a piece of corporate art... :-)
- jxf 11y agoQuestion: How do I actually run Kyle's tests to see this for myself? (Not that I don't believe him, I just want to play around a bit.) When I run `lein install` and then `lein test`, I get: ╰─▶ ψ lein test Exception in thread "main" java.io.FileNotFoundException: Could not locate jepsen/db__init.class or jepsen/db.clj on classpath: , compiling:(mongodb/core.clj:1:1) at clojure.lang.Compiler.load(Compiler.java:7142) at clojure.lang.RT.loadResourceScript(RT.java:370) at clojure.lang.RT.loadResourceScript(RT.java:361)
- jodah 11y agoUpdate the version of the "jepsen" dependency in project.clj to 0.0.3. You'll need to do this for each of the test projects you want to run. FWIW, this was filed the other day: https://github.com/aphyr/jepsen/issues/52 https://github.com/aphyr/jepsen/issues/52
- jxf 11y agoAh, that explains it! Thanks.
- caipre 11y agoCan't answer your question, but I'm curious how you managed to include an image in your comment. I didn't think embedded HTML was possible?
- jimktrains2 11y agoWhat image? The arrow and psi are characters.
- jxf 11y agoThat's from my command prompt that I wrote -- the arrows are just Unicode characters. You can see it here if you like: https://github.com/fj/dotfiles/blob/master/home/.config/shell/prompt.sh#L140-L143 https://github.com/fj/dotfiles/blob/master/home/.config/shel... Edit: On further reading, I think you seem to be thinking that the paste of my console is an image. It is just text. To make fixed-width text on HN, indent each line by four spaces, like this: This text has four spaces at the beginning of its line.
- narrator 11y agoI knew something was funny with Mongo when all the api calls defaulted to writes not being guaranteed to sync to disk. Maybe for a use case like aggregate statistics gathering it would be ok to risk missing a few updates in a crash for the sake of speed, but to make that the default??
- threeseed 11y agoYou know this configuration change was changed in November 2012. Do you think it's still relevant to be bringing this up ?
- lazyloop 11y agoActually, the defaults are still unsafe, just the marketing language has changed. Take the Node.js driver for example, it defaults to w=null and j=false. http://mongodb.github.io/node-mongodb-native/2.0/api/Db.html http://mongodb.github.io/node-mongodb-native/2.0/api/Db.html
- edejong 11y agoWould you want to miss the most important post-recovery data: when your system is under duress?
- addisonj 11y agoMongo absolutely nailed creating a database that is easy to get started with and even do things that are traditionally more 'hard' such as replication. It is still super attractive for me to pick it up for small projects, even after dealing with its (many) pain points both in development and operational settings. Given this, it is so tragic to see how dismissive they have been in regards to the consistency issues that have plagued the db since the early days. Whether it was the stupidity of bad defaults in drivers to not confirm writes, or easily corruptible data in the 1.6 days, or now with not seriously looking at the results of jepsen, the mongodb organization has never taken the issues head on. It would be so refreshing to see more transparency and admitting to the faults rather than wiggling around them until eventually pushing a fix buried in patch notes. I often feel like a mongodb apologist when I admit that I don't mind using mongo for small (and not important) projects and while the mongodb hate can be a bit extreme at times, the companies treatment of these sorts of issues may justify some of it.
- meritt 11y agoAfter MongoDB published their write speed benchmarks based entirely on unacknowledged writes (e.g. how fast can you write to a socket?), it's been a long downhill ride with an immense amount of inexplicable ignorant support.
- tootie 11y agoIf they are making money, why would they care? There's lots of shitty software raking in huge license fees based on misplaced reputation.
- hendzen 11y agoCan you post a link to these unacknowledged write benchmarks? I can't find them.
- meritt 11y agoNeed to find an archived version but it caused a lot of arguments in 2009/2010: e.g. http://rethinkdb.com/blog/the-benchmark-youre-reading-is-probably-wrong/ http://rethinkdb.com/blog/the-benchmark-youre-reading-is-pro... references similar benchmarks Can also link simply to the HN discussion from back then too: https://news.ycombinator.com/item?id=1496035 https://news.ycombinator.com/item?id=1496035 > Full disclosure: I work for 10gen. > We did this to make MongoDB look good in stupid benchmarks.
- rdtsc 11y agoI still don't get it. MongoDB can't possibly call itself a database. I can understand MongoScratchStorage, MongoPorbabilisticDataEngine but not MangoDB.
- gasping 11y agoMongoWeakReference
- ccleve 11y agoDoes anyone have any references on how you could write a distributed database that met all ACID properties? Surely there's an academic paper that says that if you do A then B then C, you are guaranteed a certain level of consistency. We've developed a type of distributed database at my company, and I think it's pretty solid, but I need a broader familiarity with the available theory.
- tylertreat 11y agoConsensus and atomic broadcast. Paxos, raft, zab. Formal methods.
- teraflop 11y agoAside from reading papers, it's a good idea to look through the syllabus of a distributed systems course to get a broad idea of what the problem space looks like. Academic papers will talk about "minimal" problems like consensus, or desirable properties like sequential consistency, and expect you to already know why those concepts are important. If your experience is mostly hands-on, it may not be obvious how it all applies to real-world systems. Say you have a complex distributed database. Forget all the bells and whistles: can it solve the problem of allowing a set of processes to reliably agree on a single Boolean value? If so, then you're trying to provide the same consistency guarantees as Paxos/Raft. So if your architecture is substantially simpler than Raft, then either you've come up with something really ingenious, or you've missed some edge cases.
- kasey_junk 11y agoOr you are obviously violating proven theory...
- holon 11y agoConsensus is the main hurdle - if you have multiple nodes that can be read from, then any values that are successfully written to the system must guarantee that those same values can be read sequentially. Issues arise when network partitions interrupt communication between nodes; even if you require all nodes to send acks when writing, how do you deal with those acks not being received?
- Kiro 11y agoThis article is too technically advanced for me. As a casual MongoDB user, how do these problems affect me?
- rockdoe 11y agoAt several points real world scenarios are described, so just skim ahead to those. Users could end up seeing private information in each others' accounts, for example.
- shawabawa3 11y agoIn very rare cases, you could see confirmed writes being rolled back, or reads returning data from before a confirmed write How these problems will actually affect you depend on your applications. It could for example allow 2 users to be created with the same email address
- tel 11y agoIn replicated Mongo scenarios higher write volume increases the probability of inconsistent reads. What this means is that there's a chance that on some data—no matter how safely you attempt to write it—you'll end up in a totally inconsistent state for your system. The actual impact of an inconsistent state is very hard to judge. It could be as minimal as having two different users just see something weird on their screen for a moment and then it goes away. It could even be totally avoided if your application handles inconsistent data well. At the same time, it could also cause complete and nearly untraceable complete corruption of all data in your system. Who knows? It'd be a bit like building a bridge using metal with a known defect. It'll probably work fine for a long time and depending on how and where that metal was used you might be alright. Or you might have a complete structural integrity failure at any moment once stress starts ramping up and you'll just have to blame it wholesale on using bad materials.
- agopaul 11y agoSo, now I'm wondering: why is Stripe using Mongo at all? Maybe they are planning to migrate to another DBMS?
- ivanb 11y agoSo what should users of MongoDB do? I'm asking because it is the main database used in Meteor and I'm very interested in Meteor. Should the general advice just be "store in MongoDB everything that doesn't require consistency and use Postgresql for everything else"?
- edejong 11y agoThe general advice should be: use PostgreSQL in case you are uncertain what to use. Watch some Youtube video's with Michael Stonebreaker (2014 Turing Award winner) and start getting disillusioned by the NoSQL hype. Then, try to understand the mess Edgar Codd tried to fix in the '60s and '70s.
- AdrianRossouw 11y agoDo you have any specific videos we should watch?
- edejong 11y agoIn the top-hit on Youtube, Michael starts to discuss the solution-space @31m15s (https://www.youtube.com/watch?feature=player_detailpage&v=OYGJe1z97VI#t=1875 https://www.youtube.com/watch?feature=player_detailpage&v=OY...)
- uptownJimmy 11y agoI think Meteor is great, except for that one thing. I won't be touching it again until there if full support for one of the SQL technologies.
- tel 11y agoYou can write a DDP backend which is backed by Postgres. In such a case it ought to feel like Mongo is just serving materialized views of the genuine, consistent data. If you treat the Minimongo data that way—just consistent enough to show an image once—and verify all the writes on the server then you ought to be able to get by.
- 11y ago
- rustsucks 11y agoWhat I love here on HN is that even though everyone here seems to hate on MongoDB, it keeps on rolling, improving, and getting used in real companies. There may be "better" (for spurious defn of better) products like rethink or whatever, but folks just aren't using them. Get over it!
- bsaul 11y agoI seem to remember from a foundationDB talk that they first spent two years building a simulation environment to control everything from network to persistance for testing scenarios. Does anyone know of any open-source project that would aim at doing the same, so that future NoSQL DB can finally be built on strong foundations ?
- Maro 11y agoFrom 2009 to 2012 I had a distributed database startup that competed with MongoDB. We used Paxos for replication and built the database with on-disk consistency guarantees --- like the ones this article looks for and rightly obsesses over --- in mind. https://github.com/scalien/scaliendb https://github.com/scalien/scaliendb Outcome: you've never heard of ScalienDB; MongoDB brilliantly won by winning the hearts and minds of hackers and coders who don't care about such issues, but were able to get started quickly with Mongo (and got cool free cups at meetups). It turns out that's most engineers out there, definitely the initial critical mass to target for a database startup like Mongo. Btw. the story behind Oracle is similar: early versions were basically write-only; read Ellison's book 'Softwar'. Of course there are other ways to get started: for example DBs coming out of academic research like Vertica seem to avoid this problem; in that case initial funding is basically provided by the gov't and when they create the company to commercialize they're already shooting for Enterprise contracts, skipping the opensource/community building phase of Mongo.
- dublinclontarf 11y agoI don't know, I got started with Mongo because ... it was so easy to start with, but come deployment time got bitten by a LOT of issues (which essentially negated any advantage in using Mongo). As a result of this experience I almost exclusively use PostgreSQL, and I've never EVER been burned by taking this approach. Sometimes I do use another DB but there has to be a seriously good reason for it.
- nazka 11y agoSame thing. I was a software dev, then I went to the back end with MongoDB and NodeJS, and I learnt what ACID means... Also about MySQL, many people forget it isn't ACID compliant either and just look at benchmarks or use MySQL "because Facebook uses it". I am sure that if any startup would use PostgreSQL it would avoid many problems on the road.
- coolgeek 11y agoI think that PostgreSQL is, in almost every aspect, a superior product to MySQL. But MySQL not being ACID compliant is flat out wrong (assuming that you're using InnoDB)
- geowa4 11y agoSince Postgres added a JSON type and Docker made running it simple in development, I haven't had a need for anything else. Call me old school, but I prefer starting with a relational database and changing when it's no longer appropriate.
- nazka 11y agoOld school? It's the best thing to do in my opinion. An ACID relational database that can do even more than that! I think it's one of the best DB for startups.
- bakhy 11y agoI must admit, I always feel like I am missing something in these discussions. Like I didn't get some memo... I just don't expect a DB like MongoDB to guarantee consistency. The whole story around NoSQL and the likes was to enable crazy horizontal scaling needed for the web. Phrases like "eventual consistency" flew around. It seems so logical - you lose consistency, gain scalability. But somehow, people simply started using them everywhere? Assuming that these DBs are just like any other? And now, we're all bashing on MongoDB because it is - not consistent? What happened here? :) NB that I do not wish to attack the OP - if MongoDB now claims to be consistent in any way, that deserves scrutiny. And these analyses are always a really interesting read. But the general tone in the developer community about MongoDB seems a bit irrational.
- Jweb_Guru 11y ago"Eventual consistency" has a very particular meaning (when it's not being used as a buzzword). "Read uncommitted" doesn't even come close to the sorts of guarantees that people expect from an AP database. More importantly, MongoDB doesn't advertise itself as an AP database, it advertises itself as something you can use as the primary datastore for important information. Kyle has analyzed AP databases like Cassandra and Riak as well, and evaluates them according to their claims.
- fennecfoxen 11y ago> It seems so logical - you lose consistency, gain scalability. ... And now, we're all bashing on MongoDB because it is - not consistent? What happened here? :) There are ways to do "eventual consistency" responsibly. Mind you, it's obnoxiously tricky to do it right, even when someone has provided an underlying implementation that works exactly as promised. But if you design your data access patterns in the right way, the system can provide guarantees so that even if it doesn't have all your data at the moment, you can still ask questions about the the state of the data that is available, and get meaningful responses back that conform to a certain set of guarantees. What happened here -- why we make fun of MongoDB -- is that it doesn't provide many promises like that, and even when it does, its implementation does a very, very bad job of delivering them ... and it doesn't even do a good job of delivering scalability. (It's basically a mmap()'d series of b-trees of BSON documents, so as soon as you run out of RAM, you're at risk of having the kernel swap out all your indicies instead of your data, whereupon performance craters. Oh, and the much-mocked global write lock has finally been replaced with a per-database write-lock in recent versions.) In short, you sacrifice everything and gain... a modestly convenient API for document-storage, maybe.