22 ms·
What MongoDB got right
- bryanlarsen 11y ago"So while MongoDB today may not be a great database, I think there's a good chance that the MongoDB of 5 or 10 years from now truly will be." Either MongoDB will be, or other databases that have learned the lessons, both good and bad, of MongoDB. RethinkDB appears to have captured the "MongoDB done right" mindshare, and PostgreSQL has gained JSON and is gaining better replication in order to cover the same niches.
- infomofo 11y agoI agree- I was a huge fan of MongoDB when it came out because of the unique data structures it enabled easily. However when it came time to select a new database for my new project, I found that the JSON support that PSQL had added gave me all the flexibility I needed while still in a somewhat relational form, and additionally it is dead simple to spin up postgres RDS instance in AWS, and it's a pain to use Mongo there.
- annnnd 11y ago> ...instance in AWS, and it's a pain to use Mongo there Why is that? Quick google search doesn't hint at problems, but rather at pretty slick marketing pages: (which doesn't mean much, I know) https://aws.amazon.com/blogs/aws/mongodb-on-the-aws-cloud-new-quick-start-reference-deployment/ https://aws.amazon.com/blogs/aws/mongodb-on-the-aws-cloud-ne...
- cwmma 11y agoIt's more then Postgres is so easy to use there with RDS which does all the setup for you.
- cwmma 11y agoIt's more then Postgres is so easy to use there with RDS which does all the setup for you.
- rspeer 11y agoRunning MongoDB effectively on AWS is very expensive. It's the opposite of what AWS is optimized for. It requires a dedicated amount of RAM that scales with your database, and it requires permanent storage that's about 100x the size of the actual data. It's very much unlike a nice, bursty, CPU-bound web server. Also, you can't use the big selling point of AWS, which is that you can scale it "elastically". Okay, you could, but it would be a terrible experience with lots of unavailability.
- threeseed 11y agoHow can MongoDB be a pain in AWS ? It is the easiest database in the world to setup. Download and run ./mongod. I've setup plenty of them in AWS and had zero issues. And if we are talking about managed databases then it's equally dead simple to spin up a Compose/MongoLab instance in AWS.
- damon_c 11y agoLast I checked, the recommended production config of mongo, in the simplest case, required at least 3 separate servers configured in a master, slave, and arbitrator cluster. Compared with setting up AWS multi zone replicated RDS SQL servers (push a button), it is very much a pain. Yes compose.io helps. Edit: https://docs.mongodb.org/manual/tutorial/deploy-config-servers/ https://docs.mongodb.org/manual/tutorial/deploy-config-serve...
- threeseed 11y ago> RethinkDB appears to have captured the "MongoDB done right" mindshare Mindshare is irrelevant. MongoDB is killing it in the enterprise right now. They have integration with Oracle, Teradata, Hadoop and countless partnerships with other vendors. You can guarantee MongoDB will still be around in 20 years the way it is positioning itself. Can't say the same about RethinkDB (as great as it is). > PostgreSQL has gained JSON and is gaining better replication in order to cover the same niches The PostgreSQL replication story is pretty pathetic given how old/mature it is. And I've seen nothing to suggest that anything is really improving in this area. There are a range of addons none of which are supported or built in. Basic replication is confusing, the documentation non existent in parts and good luck getting any support. You compare it to MongoDB (or really any of the newer NoSQL databases) and it's like night and day. It takes minutes to setup a replica set and there is plenty of documentation and official support for any issues.
- bryanlarsen 11y agoToday's mindshare is tomorrow's market share. It's not assured, but there's a strong correlation. Conversely, lack of mindshare doesn't really hurt sales, but it does hurt growth. PostgreSQL 9.4, 9.5 and 9.6 all introduce foundational changes to make eventually enable replication, but none of it is really exposed to the end user. They are working on it, but they are being very conservative.
- threeseed 11y ago> Today's mindshare is tomorrow's market share. This is just nonsense. There are plenty of hyped startups/products who went nowhere. Unless you understand how to market and execute you're going nowhere. MongoDB has demonstrated they are seriously good at it and given how well 3.0 has been received (write lock gone, extremely fast performance, Call Me Maybe test fixed) they have a lot of momentum. > They are working on it, but they are being very conservative Conservative being the operative word. You would think sometime in the last 20 years they would've tackled it.
- habitue 11y ago
- Thaxll 11y agoPostgreSQL is nowhere near Clustering / HA / sharding features. Afaik it's only a master / slave architecture by default.
- jpgvm 11y agoSure, but the underlying primitives are better. Specifically it supports quite flexible binary replication including the new logical replication features. With this you can build things like Manatee[1] which enable effectively seamless HA. The platform I am currently working on Flynn[2] uses a variant of this state machine to implement an effectively maintenance free Postgres cluster. In future though there will be support for bi-directional replication in Postgres, i.e true multi-master support. [1] https://github.com/joyent/manatee https://github.com/joyent/manatee [2] https://flynn.io/ https://flynn.io/
- tracker1 11y agoAgreed, and now that RethinkDB supports automagic failover, it's pretty much a no brainer... and while I like Mongo's query interface slightly more for most queries, RethinkDB avoids some of the weirdness when you have more interesting queries. And server-side collation/joins is a really nice feature in a document-centric database. PostgreSQL really needs a MUCH better replication/sharding/failover story... While I would use PostgreSQL in a situation where all I need/want is a single server, where multiple servers are needed for HA/failover, I'd probably just defer to MS-SQL, only because pg is so convoluted in that regard. As to MySQL/Maria... I haven't touched it in years, and every time I have some weird behavior drives me nuts. I find it funny that people can love mysql, and bash on JS. I'd also like to acknowledge ElasticSearch and Cassandra... ES is wonderful to work with for what it does best, search, and C* is a champ when you need a really good distributed table/kv store, though I think that RethinkDB is a better option today, if you don't need more than 10-20 nodes (which is a LOT).
- rwmj 11y agoJust a note that in PG'OCaml (an OCaml interface to PostgreSQL), you can write: "insert into foo (col1,col2,col3) values ($a, $b, $c)" and it creates the safe prepared statement with ? placeholders. At compile time. Type-checked against the database to make sure your program types match your column types. http://pgocaml.forge.ocamlcore.org/ http://pgocaml.forge.ocamlcore.org/
- annnnd 11y agoI would be very careful with such SQL statements. I am guessing it relies on some intrinsic fields' order? That could change anytime. Order of fields shouldn't have any impact on you app, but I think in your case it does.
- rwmj 11y agoThe "..." wasn't literal. I have amended the post to make this clear.
- emilburzo 11y agoI have to agree with the author, especially since the points he raises are the ones that helped me greatly on my first "serious" personal project[1]. Coming from postgresql land I would have never thought you can have such great replication with automatic failover. I've had literally 100% uptime for the past year. And that's on commodity servers (one of them being in a room in my apartment, the other two in a proper datacenter) going through the usual upgrades, downtime, reboots, going from mongo 2 to mongo 3 and such. Speaking of which, the migration from mongo2 to mongo3 was another pleasant surprise: they've made it backwards compatible. So I could do the upgrade on the servers, one by one, checking everything was ok and after that I could focus on updating the drivers and rewriting the deprecated queries, no need to have everything ready at once. The accessible oplog was another gem that fit my project really well. Gone was the need to poll the database, I could just "watch" the oplog. That, coupled with long polling on the browser side meant I'd have very little chatter between the db/server/web client when idle. Websockets would have been nice, but adoption wasn't high enough that I'd be comfortable going forward with it. And all this considering MongoDB was my first NoSQL experience. I agree it doesn't fit every project, but when it does, it's a really nice experience. [1] https://graticule.link/ https://graticule.link/
- ngrilly 11y agoI agree that MongoDB has a great replication story. But I don't understand that part: > The accessible oplog was another gem that fit my project really well. Gone was the need to poll the database, I could just "watch" the oplog. Coming from PostgreSQL, you could do the same using LISTEN/NOTIFY?
- emilburzo 11y ago> Coming from PostgreSQL, you could do the same using LISTEN/NOTIFY? I have to admit I was not aware of this feature. However, from the docs[1]: > Commonly, the channel name is the same as the name of some table in the database, and the notify event essentially means, "I changed this table, take a look at it to see what's new". From what I understand, you just know that something has changed, the actual change is not included in the event, so you need at least another query to see what changed. Did I understand correctly? In MongoDB you get the operation (insert, update, delete), the document and another few details right in the event. [1] http://www.postgresql.org/docs/9.4/static/sql-notify.html http://www.postgresql.org/docs/9.4/static/sql-notify.html
- s_kilk 11y ago> Let's start with the simplest one. Making the developer interface to the database a structured format instead of a textual query language was a clear win. I think this is the most significant factor, by far. With Mongo it's turtles (or at least Maps/Hashes) all the way down, without a strange pseudo-english layer near the bottom that forces you to translate back and forth. For some devs that's a big deal. For the last while I've been experimenting with bringing the same feature to PostgreSQL (http://bedquiltdb.github.io http://bedquiltdb.github.io), turns out it's very do-able, but I don't have enough time to make it as featureful as it needs to be.
- collyw 11y agoSQL is still one of the most readable languages in my opinion. Its the one language where I find it easier to read queries than write them.
- s_kilk 11y agoI should clarify, I'm a big fan of SQL and relational DBs, but it's hard to deny that some/many devs find the abstraction boundary a bit weird. And really, when you think about it, it is a bit weird, (we send sentences of almost-english toward the DB, then get structured data back) and there's nothing wrong with that. But many devs do value the ability to stay in hash-map land all day without needing to think about how to cross an abstraction boundary into the DB, and MongoDB actually has a pretty cool solution there.
- pmelendez 11y agoSQL is fine, no problem with it unless you are embedding queries in a OO application and you are forced to hack mapping layer between the two interfaces, which gives you two options, either compromise performance for the sake of development easiness or vice versa.
- halostatue 11y agoI’ve played a bit with Jeremy Evans’s Sequel (http://sequel.jeremyevans.net http://sequel.jeremyevans.net), and I’m not sure that ORMs are necessarily constrained along this dichotomy (easy to use but with poor performance vs great performance but hard to use). Sequel seems to be about making the hard things easy and reducing the amount of time you need to drop to pure SQL to almost zero. As someone who has built with Oracle’s Pro C(++) and with various ORMs, Sequel is an ORM that makes me reasonably happy and gives me the expressiveness of embedded SQL without compromising my object model.
- ngrilly 11y agoI agree that the three areas outlined in the article are things that MongoDB got right: a structured query language (instead of a textual query language), replica sets, and the oplog. But the lack of transactions over multiple documents (in the same shard at least) and the lack of joins over multiple collections are a big showstopper for the kind of applications I develop. I note that solutions like YouTube's Vitess provide something similar to MongoDB's replica sets. I also note that PostgreSQL's logical decoding provide the same functionality than MongoDB's oplog tailing.
- progx 11y agoAlways wonder what kind of simple apps most people must write, if they not need joins? I will be happy if i got such simple tasks :)
- maze-le 11y agoThere is always the possibility to make a map/reduce over multiple datasets. It is not exactly a replacement for join, but you can cover aggregation over multiple datasets with a output-collection. Well... in my opinion this is much more complicated and error-prone than a join-operation on a relational database.
- threeseed 11y agoHow exactly do you think eBay, GMail, Facebook etc work ? They aren't relying on relational database joins. If you want to write a truly scalable application you structure everything such that you do joins in your application layer. http://highscalability.com/ebay-architecture http://highscalability.com/ebay-architecture And in the case of MongoDB you avoid joins since it is a document database. You embed data instead.
- ngrilly 11y agoNot everybody works at eBay, GMail or Facebook scale. Most applications fit very well in a single server. For example, Stack Overflow runs on a single instance of SQL Server, replicated to a slave in another data center. In such a case, the convenience of joins and transactions is priceless. And even at scale, it makes sense to rely on joins and transactions. The perfect example is AdWords that runs of F1 and Spanner: "Our users needed complex queries and joins, which meant they had to carefully shard their data, and resharding data without breaking applications was challenging." http://static.googleusercontent.com/media/research.google.com/fr//pubs/archive/41344.pdf http://static.googleusercontent.com/media/research.google.co...
- _yy 11y agoRethinkDB took all the good parts of MongoDB and added proper engineering. https://www.rethinkdb.com/ https://www.rethinkdb.com/
- ngrilly 11y agoBut still no transactions over multiple documents (at least in the same shard)?
- jmakeig 11y agoACID transactions in a highly available distributed system are hard and often fail in subtle ways when done wrong at the edges. Any implementation will take years to mature in the lab and in actual production usage. This isn’t a knock on the Rethink guys; their product looks pretty awesome and is moving quickly. For a solution today, MarkLogic is a transactional distributed document database. Cross-document and cross-partition transactions have been a key tenet of the architecture from the beginning (like, 2002 beginning). Take a look at https://developer.marklogic.com/blog/how-marklogic-supports-acid-transactions https://developer.marklogic.com/blog/how-marklogic-supports-... for details. Full disclosure: I’m a Product Manager at MarkLogic.
- krisdol 11y agoI don't understand the recent backlash against NoSQL here. First off, almost all of the complaints would have been valid years ago. Secondly, there is so much more choice out there today if mongodb wasn't the right answer for your project, and so many NoSQL stores have had time to mature and get polished APIs and docs. We use various data stores for different purpose across microservices, mostly ES, couchbase, and datomic, and "use the right tool for the job" and "do one thing and do it well" feels like the right approach to take. For most applications, a SQL DB feels like a really big hammer that is put to a lot of things that don't look like nails.
- gedrap 11y ago>>> use the right tool for the job" and "do one thing and do it well" feels like the right approach to take Absolutely. However, database is a sort of an extreme example. A lot in the modern software (especially Web) relies on the database, and often migrating to completely different one (because requirements change and it might not be the right tool anymore) is a huge task. So you want to use something flexible enough. Also, you want to hire people, people leave the jobs, people change teams, etc. If you use some exotic, less common DB, it adds a lot of overhead. And if you apply the "right tool" to an extreme and have a few completely different DBs flying around, your maintenance cost increases a lot. See, SQL might not be a perfect, most elegant choice, but most often it is just good enough. A lot of people have used it, a lot of people have scaled it. If you run into an issue, often enough,other people did too and blogged about it, etc. Hiring / getting help will be much easier than $insertNoSQLDBName. And, let's be realistic, relatively few companies have hundreds of gigabytes or terabytes of data that typical relational DBs can't handle. My rule of thumb is that if you're in doubt, use SQL/relational store (I realize that they are different things but often used as synonyms and mean MySQL/PostgreSQL/etc).
- jeffdavis 11y ago"Do one thing and do it well" is problematic for things that store a lot of data. Especially for things that are supposed to be an authoritative source. Getting storage right is very hard. Either it's too low-level, and it's hard for applications to coordinate complex operations without corrupting data; or you end up putting a lot of features in and end up with a SQL dbms; or everything does its own storage and you have a mess.
- sriku 11y agoWhen we chose MongoDB for a project, a dominant criterion was out of the box geo queries. It helped that the storage and query approach had good impedance match with NodeJS. From a query perspective, we wouldn't have benefited much from SQL anyway, since much of the reading is free text or social graph or location based search which we moved to Solr.
- bsg75 11y ago> You can argue, and I would largely agree, that this is actually part of MongoDB's brilliant marketing strategy, of sacrificing engineering quality in order to get to market faster and build a hype machine, with the idea that the engineering will follow later. Author nearly lost me here with this logic. Placing Marketing ahead of quality in something that is supposed to store a very valuable asset (data) is near insanity. I get the mindset of "break fast", "release often", etc. in terms of customer facing features, but in something that is supposed to be a core part of your foundation, stability is if utmost importance. Otherwise nothing else works - and you lose customers, business, opportunities - because you can't look them up later. Its not "brilliant marketing", its just marketing.
- smacktoward 11y agoThis is all true, but the success of MySQL shows pretty clearly that just because something is insane doesn't mean it's not good business.
- bsg75 11y agoI think the success of MySQL is due to there being fewer options for a period of time (the "dot.com boom"), and thus it became a popular choice to avoid commercial RDBMS costs. I'm no MySQL fan when things like PostgreSQL are an option, but its probably more sane than some other currently popular choices.
- franzwong 11y agoIt becomes much simpler to setup replication in PostgreSQL than before. reference: https://www.digitalocean.com/community/tutorials/how-to-set-up-master-slave-replication-on-postgresql-on-an-ubuntu-12-04-vps https://www.digitalocean.com/community/tutorials/how-to-set-...
- angelbob 11y agoI love the point about the Oplog. There are a few equivalents for common SQL DBs (see LinkedIn's Databus for Oracle and MySQL), but in general, getting access to the write log is really hard. Even though it's sitting there! It would be wonderful if there were some kind of established API or library that would let you parse the MySQL write log without doing hideous, fragile operations that change from version to version. Sure, change the format, but at least version and document it!
- yummyfajitas 11y agoCounting arguments very carefully? Nearly every SQL library does this for you. cur.execute("INSERT INTO a (b,c) VALUES (%(a)s, %(b)s);", { 'a' : a, 'b' : b }) Also, SQL is typed, so even if you did fail to count arguments there is a good chance you'd just detect it the first time you ran it. The article acts as if treating the DB like native structures is somehow innovative and new - it's not. https://en.wikipedia.org/wiki/Object_database https://en.wikipedia.org/wiki/Object_database We mostly abandoned object databases because they sucked. SQL was a huge improvement over them. SQL is a great way to organize and preserve the integrity of a lot of business data. It's also a fantastic way to avoid repeated trips to the DB: SELECT * FROM employees AS e WHERE e.department_id = (SELECT id FROM departments WHERE name = "engineering"); In Mongo, I'm pretty sure you need to first lookup engineering, then lookup the employees in engineering. That could be O(# employees in engineering) queries rather than 1.
- acjohnson55 11y ago> In Mongo, I'm pretty sure you need to first lookup engineering, then lookup the employees in engineering. That could be O(# employees in engineering) queries rather than 1. Or, you could denormalize, and give yourself all sorts of future headaches maintaining data integrity.
- lloyd-christmas 11y ago> In Mongo, I'm pretty sure you need to first lookup engineering, then lookup the employees in engineering. That could be O(# employees in engineering) queries rather than 1. The problem with that summary boils down to bad architecture. The point of document storage is storage with purpose; the intent being to make querying EASIER. This could easily be structured to be a single query. You can structure a document countless ways to represent that query, all of them would likely be different based on the purpose of the app.
- yummyfajitas 11y agoWhereas with SQL there is more or less a single canonical way to do it and it's mostly independent of the app. I.e. the data design is minimally coupled to the specific use cases. Right now I'm building a data store and I don't know the app(s) that's are going to be built on it. It would be really great if computing could stop forgetting it's history. Object databases failed for a reason.