9 ms·
Startup Engineers and Our Mistakes with MongoDB
- nemild 9y ago(I’m the author of this series) Eliot Horowitz (HN: @ehwizard), MongoDB’s current and founding CTO, reached out after my last post - and spent two hours providing feedback last week in Palo Alto. It was an expansive discussion, and Eliot was reflective and eager to understand the perspectives I had heard. He noted how much it mattered to him what HN thought. I left with tremendous empathy for the challenges of building a venture-backed database company - while we also disagreed in a few key areas. As engineers, I hope we continue to have thoughtful, if spirited discussions, like Eliot and I had - where we both are open to being wrong and have a desire to understand differing viewpoints. In the interest of length, I didn’t write the whole interview into my series, but wanted to share key parts of our discussion on HN: - NoSQL: I felt that NoSQL was overhyped in the early 2010s - and that 10gen’s marketing claims were overblown. Eliot argued that many of NoSQL’s benefits have been realized and that 10gen’s early marketing accurately reflected the changes to come. For example, while he doesn’t love the term NoSQL for its impreciseness, he feels that both the JSON-like data structure and horizontal scaling are here to stay and were in fact the key changes NoSQL led to (not the SQL DSL). In that sense, he argues that Amazon RDS is a form of NoSQL today - and that NoSQL had a powerful impact on the roadmap of existing SQL databases. - Data Loss: I generally have stayed away from talking about the controversial examples of MongoDB data loss in the earliest days (there are several of HN posts that note this). I know personally of only one team that lost data, but it’s always hard to understand if this was a database issue - or their own mistakes. I did ask Eliot explicitly about this, and he said that these were exceedingly rare. - Defaults: I feel like it’s playing with fire to set bad defaults in a database - with numerous data breaches due to 10gen’s early decisions on authentication, remote login, and encryption (see for example, https://snyk.io/blog/mongodb-hack-and-secure-defaults/ https://snyk.io/blog/mongodb-hack-and-secure-defaults/ ). For auth, Eliot argues that developers need to take responsibility for exposing MongoDB on public servers - and that the SLA for a self-hosted instance is different than a managed instance (at minimum, I have issues with users having their data exposed to the world through no fault of their own). He disagreed with 10gen’s decision to turn on auth by default in later self-hosted versions once MongoDB ignored remote connections by default (but thought this was the right choice for the managed Atlas service). (But before 2014, the default behavior was no auth - and accepting all remote connections, see https://snyk.io/blog/mongodb-hack-and-secure-defaults/ https://snyk.io/blog/mongodb-hack-and-secure-defaults/ ; Eliot notes that this took a while because changing the default would have caused issues for existing customers) I do have concerns when 10gen explicitly targets junior developers (to be fair, he could never have predicted the growth of Node and the interest in the backend for frontend engineers). What he says makes sense say 20 years ago, but with 25% of new software engineers coming from coding bootcamps with non-engineering backgrounds, I worry that defaults matter ever more in dev tools (and even seasoned engineers may mess this up, if they’re coming from a database with different defaults). We discussed analogies like seat belt lights versus the responsibility of passengers to know better. He also argued that waiting to get all this right - not just auth - would impact database innovation, while I think there’s a balance that gets us a lot of the low hanging fruit (like security). - Mistakes with MongoDB: Eliot felt that certain posts (such as Sarah Mei’s popular HN post about Diaspora: https://news.ycombinator.com/item?id=6712703 https://news.ycombinator.com/item?id=6712703 ) misunderstood how to architect MongoDB. Regardless of the particulars, I think that new technology is especially susceptible to issues like this until the community develops broader knowledge. The failure states are well known in many relational databases and there is a broad base of knowledge about how best to architect relational data models. As startup engineers, we have to weigh tradeoffs of new vs old tools, with some new tech being game changing for startups (see PG’s much cited post about the benefits of new languages like Python: http://www.paulgraham.com/gh.html http://www.paulgraham.com/gh.html ) while many others are hyped new tech that damages startup productivity. - SLAs are hard in databases: Early products make choices that customers get used to (unacknowledged saves), and you can’t quickly just change the default behavior, esp in a database. 10gen’s early customers loved not waiting for the writes, and only later did 10gen realize that this was an issue for others (the default is very different in nearly every other database). To migrate their first customers without dislocation, they had to hold off on changing the default behavior for longer than they wanted to. In Eliot’s view, their competitors would unfairly argue that this default was a way to juice benchmarks, hoping to stoke anger at MongoDB and cut into their growth. (I tend to have sympathy for 10gen’s perspectives, with the controversial, mistaken benchmark that I referred to in part 1 looking to me like an honest mistake) Generally, he felt that the issues with MongoDB were few and far between - and that the benefits of MongoDB was game changing for so many startups. Many of them would not have survived without MongoDB. I’ll add some more notes from the interview in part 3, where I go into 10gen’s early marketing. I also let Eliot know that I would be happy to share/publicize any broader responses/critiques he writes, so that we can have a thoughtful debate that benefits others (and he can point out issues in my arguments). (Apologies in advance if I’ve made mistakes in representing Eliot’s views)
- notduncansmith 9y ago> To migrate their first customers without dislocation, they had to hold off on changing the default behavior for longer than they wanted to. This does seem like a sticky situation, but a potential solution springs to mind. Maybe there's a reason this would have been infeasible, but why not introduce a "MongoDB Legacy" product line which would be a fork with unsafe defaults, secondary to the main product/branch with safer defaults? That way the old customers would have a clear upgrade path every release, just like the customers on the safe product, at the small expense of 10gen having to cut 2 releases each time and mind the diffs around options and defaults. Maybe this would have been more expedient than waiting until version 2.6 of MongoDB?
- eip 9y ago> In that sense, he argues that Amazon RDS is a form of NoSQL today Lol. What? RDS is literally hosted RDBMS SQL databases.
- nemild 9y agoI had a similar reaction.
- nasalgoat 9y agoHaving personally talked to Eliot many times on the phone going over our substantial issues with Mongo at scale back in 2012, including data loss, I find him saying it was "exceedingly rare" rather amusing.
- cmenge 9y agoI'm a long-time MongoDB user and largely a fan of it because I believe the interface is superior to some text-based SQL statements. But whatever DB you like, if you move the state of your application to another application (i.e, a database) you better make sure you really understand how it works. For SQL databases, many people think they know how they work, but misconceptions seem widespread. Essentially, many beginners believe that SQL databases guarantee serializability everywhere and all the time, relying on the database (and OR-mappers) a lot / too much in my opinion. Choosing a system that only offers very few guarantees forces you to think about them more explicitly. On the other hand, if you never bothered to understand SQL, you probably won't bother understanding anyDB's restrictions and guarantees either, and then everything goes to sh*t. Some databases might be easier to understand than others, but I feel MongoDB is on the 'easier' end here. YMMV.
- danpalmer 9y ago> Some databases might be easier to understand than others, but I feel MongoDB is on the 'easier' end here. YMMV. I agree with most of this, relational databases have a lot of moving parts that many developers (myself included) don't fully understand. However, I believe relational databases (Postgres is what I have experience with) has defaults that are basically correct, and unlikely to cause significant issues, whereas in my experience MongoDB did not have this. MongoDB might be easier to understand on the surface, but I think that's deceptive.
- adjkant 9y ago+1 - Mongo is much easier to misuse. Not to mention that many users of SQL these days use ORM's that abstract many of the complex SQL features with a professional library that actually knows how to handle the small details and only exposes a basic subset of functionality that is good enough for most use cases.
- Bartweiss 9y ago> I believe relational databases (Postgres is what I have experience with) has defaults that are basically correct, and unlikely to cause significant issues, whereas in my experience MongoDB did not have this. Strongly, strongly agreed. Learning every moving part in InnoDB (as an example) is quite a task. But not knowing about those features will rarely burn you, and it's usually apparent when you need to know something you don't. Meanwhile Mongo... well, Mongo says a replicated write is complete as soon as it's queued to be written. They eventually updated the default settings to be slightly less catastrophic, but still far from good. What you don't know about Mongo can and will hurt you, which arguably makes the appearance of ease a drawback. http://hackingdistributed.com/2013/01/29/mongo-ft/ http://hackingdistributed.com/2013/01/29/mongo-ft/
- flavio81 9y agoToday in 2017, there is no real reason to use MongoDB other than in prototyping. I am happily waiting until the final nail is put on the coffin of this overhyped, flawed document store.
- jamroom 9y agoWhy? I don't personally care if MongoDB succeeds or fails, but if they can listen to their users, fix their issues and come out the other end with a database that works as advertised what is wrong with that? Why is it so important to you that they fail?
- mikestew 9y agoWhy is it so important to you that they fail? If it fails, then I can quit having the same discussions repeatedly as try to talk my co-workers out of using it.
- optimuspaul 9y ago"Today in 2017, there is no real reason to use Oracle other than in prototyping. I am happily waiting until the final nail is put on the coffin of this overhyped, flawed relational store." Point being you can say just about the same thing about any other database. They are all flawed for one reason or another. I happen to like Mongo quite a lot, but I rarely use it. You need to understand what situations it works best in and it's unfair to both it and yourself to generalize and wish it dead.
- flavio81 9y agoOracle RDBMS isn't flawed. It is expensive, complex to configure/mantain/extend, expensive in licensing, expensive in storage, and also expensive. Did i mention expensive? PostgreSQL is eating Oracle's customers little by little. RDBMS aren't flawed at all, this is technology that has been perfected for the last 40 years. What is flawed is to think that relational data can be easily stored on a document store (typical mistake). I think document stores are a great thing; the only problem is that MongoDB isn't a good document store.
- pryelluw 9y agoI feel the best thing to come out of MongoDB is that Postgres now handles JSON.
- STRML 9y agoYes, this was an incredible development, and lead to hilarious projects like ToroDB (MongoDB API on top of Postgres), which ended up being faster and safer, with relations when you need them. JSON columns can be incredibly useful and easy to understand for things like user preferences, without the hassle of a metatable.
- tinix 9y agoCall this hilarious if you want, I think it's rather cool, thanks for the pointer. Once this is out of beta it will definitely be useful, if not already.
- flavio81 9y agoNo doubt the best thing that has happened lately to Postgres as well! And Postgres can do JSON with vastly more performance than Mongo: https://www.enterprisedb.com/postgres-plus-edb-blog/marc-linster/postgres-outperforms-mongodb-and-ushers-new-developer-reality https://www.enterprisedb.com/postgres-plus-edb-blog/marc-lin...
- metheus 9y agoIf you're going to get on the "MongoDB was overhyped" bandwagon, you should probably not post links to performance comparisons you pulled off of an EDB Postgres website.
- pritambaral 9y agoAt least EDB used an inspectable and reproducable test. I think posting links to performance comparisons is fine when the performance comparison can be independently verified.
- MikeKusold 9y ago> By far the most consistent mistake was choosing a non-relational database, when your data was strongly relational. Mongoose's ODM made this mistake surprisingly easy to make, which led to issues down the line. This mirrors my experiences with Mongo as well. The vast majority of data is relational. Mongoose allows people to make the mistake of structuring their data relationally. But all you are doing is pushing all the joins to the web server (which in Node land is single threaded and compounds the mistake of choosing NoSQL). There may be some microservices that Mongo is great fit for, but it should not be your core data store.
- mrtksn 9y agoIf data needs to be processed, an option would be to use MongoDB for collecting the data in bulk and later decide how you need to structure it for your needs. When it comes to "read" the data, you read it only from processed database that can be anything. Since you can do your processing completely independently of your web server, your are not necessarily pushing your computational load from the DB to the web server.
- emodendroket 9y agoIt that point it's harder for me to understand why I'm not just writing out JSON files or something.
- mrtksn 9y agoYou definitely can but MongoDB provides convenience over storing and managing JSON files on you filesystem by your own efforts.
- emodendroket 9y agoSure, but it is also another service to keep running. Depends on what scale you're operating at.
- 9y ago
- adjkant 9y agoThe reality of MongoDB is that it's a very specific use case of nonrelational data, which is very rare these days. I'm sad that so many people get looped into Mongo with a MEAN/MERN stack when those apps are almost always the CRUD apps that would benefit from SQL. Why force yourself to maintain a schema implicitly for relational data when you can get error messaging and explicit schemas? I think MongoDB is a great fit for the right problem. I've been coding for 4 years and have yet to find a problem that fit Mongo best. As far as quick prototyping goes, this article articulates well why that really isn't the case. Going back to the authors post here on HN, that's a point I would really disagree with the CTO on as well.
- scarface74 9y ago"I've been coding for 4 years and have yet to find a problem that fit Mongo best." I'm going to leave that right there.....
- adjkant 9y agoI can't read the implication of this, but to clarify, I haven't personally encountered a problem that fit Mongo best. I can imagine one.
- hitgeek 9y agonothing much new here. mongodb marketing/PR was a problem, advertising a general purpose RDMS alternative without the technology to support it. users were also a problem, using a technology without understanding the architecture and trade-offs. to be over critical of "mongodb" though is unproductive. there were a lot good people working on new software to solve hard problems and I'd like that to continue. lets just not repeat the same mistakes. don't get caught in marketing hype, and carefully evaluate technology decisions.
- slap_shot 9y agoIt terrifies me to see this quote from their CTO: "MongoDB's CTO disagrees with this statement arguing that nearly 90% of database installations today would benefit from being replaced with MongoDB." I used to attend "office hours" at MongoDB's office where guests ask MongoDB employees for help. Most of my questions involved very complex aggregation queries (that would have been trivial in SQL) that even MongoDB employees could not solve. While I waited to be helped I would listen to other people ask the same questions over and over and over: "How do I join these two collections? How can I <perform a transaction> and ensure both writes succeed/fail?" Rather than say "MongoDB doesn't offer this functionality" and maybe advise them that this database isn't what they needed for the specific project, the engineers spent a majority of these office hours explaining they don't need schemas, transactions, relations and advised them how to hack together something that worked. I don't think MongoDB is appropriate for 90% of the database installations out there. I don't know the number, but it isn't 90%, it isn't a majority, and doubt it's a significant number. MongoDB really needs to explain what it does exceedingly well versus the other distinct offerings in databases (RDBMS, Key Value stores, MPP/warehouses, HDFS, etc) and market that. What they do right now is somewhere between disingenuous and down right negligent. Of course its the responsibility of a company and its staff to choose the right tool for the job, but these postmortems are becoming tiring.
- greenshackle2 9y agoMongoDB is web-scale.
- justusw 9y agoSo is /dev/null.
- valarauca1 9y agoDoes /dev/null support sharding?
- kakwa_ 9y ago
- spchampion2 9y agoThis brings back some memories from 2012. I was updating an internal tool at the company where I worked, and after a lot of thought I decided to replace a tried and true MySQL database with Mongo. It really seemed like a good idea at the time - we were storing loosely related documents where not having a fixed schema was an advantage. Any advantage I got was lost by the time I built all the logic to handle the joins I did need to make (those join keys weren't so schemaless after all). And then I needed logic to make sure every record had the right keys, that they made sense, etc. This was all stuff that MySQL had been doing for me. But the real kicker was that Mongo made my performance worse. When I migrated my database (roughly 10 million records), I compared before and after sizes of the actual files. Mongo's were 2-3x the size. I didn't realize it before starting, but Mongo's preallocation model gave me huge files with sparsely written records that had room for future updates. But I didn't need that extra room because my data rarely changed. So I ended up with larger files that took longer to scan for unindexed queries (mostly for reporting), which meant I had to index more stuff, which increased my memory usage, and forced me to upgrade to larger servers. Had I stuck with MySQL, I would have been fine. Many expensive lessons were learned from this, which I guess was valuable.
- eberkund 9y agoWhat about for IoT devices? I could see NoSQL still having a place there, although I could also see it working just as well with a SQL database.
- mbreese 9y agoWhat is special about IoT devices that make them more suited towards a NoSQL solution? Scale? Is that really an issue for the majority of IoT deployments?
- kazen44 9y agomost IoT devices do not really need to be scalable themselves. The backend/processing side of IOT needs to be scalable. And it already mostly is doing that on both network side (ipv6) and datastore side (TSDB?).
- eberkund 9y agoScale and also the data structure is often times very flat; usually just a bunch of records from sensor readings.
- yinso 9y ago> Schemas: Schemaless does not mean no schema; instead, it means an implicit schema in the app (a particularly challenging misnomer for anyone outside our industry) Still a challenge for many in the industry.
- meesterdude 9y agoA really great writeup. A revealing exploration of the perils of mongoDB, and of the greater issue of selecting technology for a project. There is never any mention of usecase, in companies. Everyone is still trying to be google or facebook, and uses "their tools" without a use case for why. It's just because that's what everyone else does. All technology has a place somewhere for some use - but we need to know why we are choosing them. For example, I choose Ruby on Rails because of it's gains in productivity, reliability, security and pleasure of use. I can quickly sling maintainable code that can scale up to millions of users. And most importantly, I don't have to make a lot of decisions. Best practices are defaulted into the framework. Some people shy away from frameworks, favoring "simpler" tools - but that only means they have to re-solve all the problems the framework has already solved. Likewise, with MongoDB, all those useful things SQL databases do for us need to be resolved. But, that's a database decision, and many people making the decision are just developers in search of a data store. Transactions and ACID are often not even on their radar. Overall, i think this is a symptom of a lower barrier for entry. Where before you needed a CS degree to be a programmer, now you just need to go to a bootcamp. People are not properly trained, aware (or interested) in the underpinnings like databases, networking or security - and to their own peril and the peril of the companies they work for.
- deleted 9y ago[deleted]
- mnm1 9y ago'One 10gen engineer made this point in analogizing SQL to Cobol, arguing that "SQL is Annoying":' Why wasn't this the tagline for Mongo to begin with? At least it shows the arrogance, ignorance, and outright stupidity behind the database. This is not an engineer's comment. It cannot possibly be now, decades after SQL has become a de facto standard for some very good reasons (outlined in the article and elsewhere, won't rehash here). It is a marketer's comment, probably spoken by an engineer who isn't qualified to build simple demo apps, let alone anything as complex as a database. Mongo's CTO seems to fit this bill of marketer more than anything else or how would he be able to make the claims he makes with a straight face? Our industry has a problem with fads and decisions based on feelings. Even if SQL is annoying, it's incredibly stupid to base your tech choices on feelings. It's even dumber to choose a product whose engineers build the product based on feelings. Last I checked, I thought we were an industry of engineers, trying to apply scientific principles and some human ingenuity to build software. Where does 'annoying' fit into that? Or other adjectives I often see thrown around on these boards like 'clean' and 'slick'. I know it's hard to quantify the quality of software and speak about it with any coherence, but this is beyond incoherent. 'Annoying' as applied to SQL tells me nothing. 'Annoying' as applied to the engineer making this incredibly stupid comment tells me that this engineer is either lazy, gullible, or just downright stupid. Have we seriously lost our ability to discern hype from reality that we let people and companies like this dictate our technology choices, throwing out decades of solid research in computer science for what some idiot at a company more specialized in marketing and PR than engineering tells us? I'm willing and ready to hear well-thought out criticism of SQL and RDBMS. What I'm not willing and ready to hear is some idiot's feelings about SQL or RDBMS. You want to make a claim? Measure it. Present a report with data. But goddamn, if one of my reports came to me with this kind of stupidity, I'd give them a chance, but if they persisted, I'd fire them. This isn't revolutionary thought. It's not evolutionary thought. It's lazy idiots who don't want to learn SQL, a language so easy even business people with essentially no computer skills can pick up.
- kenwalger 9y agoThese "Why technology X is worthless" types of posts and statements always give me a bit of a laugh. As with most technologies, there is a) a learning curve, and b) best practices to follow. Many of the comments here seem to discuss a few things which, in my opinion, lead me to believe that the implementation is incorrect or decisions are being made on old data or based on older versions of MongoDB. Making arguments against any product based on previous versions seems to be counter productive. Most of the posts here don't make reference to a specific MongoDB version, but many reference their experience with the product in 2012. Assuming that these experiences were at the end of 2012 with the most current at the time version, we are still talking about version 2.2.2. The current stable version is 3.4.6. As you can imagine, there have been many advances to the product in five years. Basing one's knowledge and opinion of MongoDB on old versions doesn't seem logical. Even back in 2012 though, there was a lot of misinformation about MongoDB. A blog post from that time period (https://blog.serverdensity.com/does-everyone-hate-mongodb/ https://blog.serverdensity.com/does-everyone-hate-mongodb/) argues some of that information at that time. Many other comments were made based on what seems to be a lack of spending time with learning MongoDB. It is a non-relational document store. It is NOT a relational database. It requires a different way of storage design and thinking and attempting to force MongoDB to be a SQL-like database is silly. SQL is indeed a popular option and can work well. However, non-relational databases work well too. There are lots of implementations of it across a lot of companies. A few relatively recent posts are good examples https://engineering.snagajob.com/mongodb-in-aws-does-it-really-work-a705b3da07a2 https://engineering.snagajob.com/mongodb-in-aws-does-it-real... and https://mongomikeblog.wordpress.com/2016/04/29/why-we-went-with-mongodb/ https://mongomikeblog.wordpress.com/2016/04/29/why-we-went-w... That doesn't account for many of the large companies out there using it, like Expedia, Facebook, Forbes, to name a few. If we want to have a discussion about specific, current, features of MongoDB, that's great. There are a lot of them. Many have been mentioned. Some have been mentioned as a negative due to what appears to be poor implementation. If we want to debate pros and cons of any product, we should be sure that we are talking about how things are intended to be implemented and not some hacked together approach. Perhaps the marketing spin was too aggressive. I'm not a marketing person. I'll leave it to Mr. Horowitz to backup his claims. But if we are going to have a discussion about the pros and cons of MongoDB, can we at least agree to not talk about old versions? I mean, I had a bad experience with Windows 3.1, so should I not use Windows anymore? ;-)