42 ms·
MongoDB queries don’t always return all matching documents
- Animats 10y agoNot when they're changing rapidly, anyway. Well, that's relaxed consistency for you. Does this guy have so many containers running that the status info can't be kept in RAM? I have a status table in MySQL that's kept by the MEMORY engine; it's thus in RAM. It doesn't have to survive reboots.
- doubleorseven 10y agoMongo, in one word: sucks. Couchbase, does not.
- bioinformatics 10y agoI use RethinkDB in most of my production things. I recommend it.
- ubercore 10y agoAgreed, I've had nothing but positive experiences with it.
- bioinformatics 10y agoI started with Mongo too, had some performance issues and started used RethinkDB when they first released it. It does get better with every update.
- jwr 10y agoSame here. It has its quirks, but it works.
- louthy 10y agoHaving not used it, I'd be interested to know what those quirks are?
- vegabook 10y agoand when I grow big, I will use Cassandra.
- jordz 10y agoI absolutely love Cassandra, assuming you have the right use case for it. We had some aches and pains with anti-entropy ops back in the day (7TB + on too small a ring) but it's my favorite storage system :)
- tracker1 10y agoRethinkDB has been really mindful of its' consistency, has an in the box story for observable collections, and a really great UI out of the box. I really hope I get to use it for more than toying around with. I'm not sure that I'd choose Mongo over alternatives these days, if I'm on AWS and can use RDS, I'd go PostgreSQL mostly, but RethinkDB, ElasticSearch, C*, and others also have their places.
- holografix 10y agoMongo in on word: popular. Couchbase, not so much. I'd hazard a bet that you'd find some dirt under that couch if you went looking but not enough people have looked at it yet.
- skjhn 10y agoAt least you get SQL and JOINs with Couchbase.
- lossolo 10y agoI've just migrated one project from mongo to postgresql and i advise you to do the same. It was my mistake to use mongo, after I've found memory leak in cursors first day I've used the db which I've reported and they fixed it. It was 2015.. If you have a lot of relations in your data don't use mongo, it's just hype. You will end up with collections without relations and then do joins in your code instead of having db do it for you.
- Joeboy 10y ago> don't use mongo, it's just hype I'm kind of curious as to where this hype is. I've almost never heard anybody say anything positive about mongodb. All I ever see is people saying it's terrible / hilarious for various reasons.
- api 10y agoIt's massively successful in the entry level web coding world due to a combination of good marketing and the belief that it gives you unlimited scale 'for free' and everything 'just works.' Not saying those are impossible goals but so far no database has managed to deliver that. Building a large-scale endlessly-scalable database is still very hard and very detailed and easy to screw up.
- tim333 10y agoTheir website is quite hype prone: >The Standard for Modern Applications >MongoDB 3.0 features performance and scalability enhancements that place MongoDB at the forefront of the database market as the standard DBMS for modern applications. >Also included in the release is our new and highly flexible storage architecture, which dramatically expands the set of mission-critical applications that you can run on MongoDB. >These enhancements and more allow you to use MongoDB 3.0 to build applications never before possible at efficiency levels never before attainable.
- 010a 10y agoThe people who wrote that are not the same people who write the database.
- jitix 10y agoWhat storage engine are you using? I wonder if the same issue comes in wiredtiger MVCC engine.
- im_down_w_otp 10y agoSaid it before, will say it again... "MongoDB is the core piece of architectural rot in every single teetering and broken data platform I've worked with." The fundamental problem is that MongoDB provides almost no stable semantics to build something deterministic and reliable on top of it. That said. It is really, really easy to use.
- eloff 10y agoAs a guy who works on ACID database internals, I'm appalled that people use MongoDB. You want a document store? Use Postgres. Why on earth would you use a database that makes so little in the way of guarantees about what results you get from it? I think most people have really low load and concurrency, so things seem to work. When things get busier you're in for a world of pain. Look I get that's it's easy to use and easy to get started with, but you're going to pay for all of that later.
- PeCaN 10y agoMongoDB: Because /dev/null doesn't support sharding.
- nameless912 10y agoBut it does support sharting. Your entire app. Into the ether.
- brianwawok 10y agoAt webscale
- leothekim 10y agoHa ha, you said "sharting."
- deleted 10y ago[deleted]
- 10y ago
- apeace 10y agoTL;DR During updates, Mongo moves a record from one position in the index to another position. It does this in-place without acquiring a lock. Thus during a read query, the index scan can miss the record being updated, even if the record matched the query before the update began.
- nameless912 10y agomongo developer 1 Man, it's just taking too damn long to run this update. What should we do? mongo developer 2 Uh....remove all the locks? mongo developer 1 Oh, yeah, that makes sense, let's do that. mongo developer 3 Hey, what are you guys up to? mongo developer 2 Making it web-scale! mongo developer 3 Good job, keep it up!
- xyience 10y agoSeriously. I was looking at DB usage statistics recently and was appalled MongoDB is still so popular. I thought it was done, nail in the coffin, when https://www.youtube.com/watch?v=b2F-DItXtZs https://www.youtube.com/watch?v=b2F-DItXtZs came out 6 years ago, I haven't followed it much since then apart from the occasional post like this whose content is just "you thought it was bad already? haha it's worse."
- ruw1090 10y agoWhile I love to hate on MongoDB as much as the next guy, this behavior is consistent with read-committed isolation. You'd have to be using Serializable isolation in an RDBMS to avoid this anomaly.
- prodigal_erik 10y agoThis is worse than read-committed because you're not even seeing the old state of the document. If an update moves a document around within the results, and it ends up in the portion you've already read, you just don't see it at all.
- ht85 10y agoThe article suggests that tuples being moved to different storage locations can cause them to not show up in a table scan. No such thing can happen in a sane RDBMS, no matter the transaction isolation level.
- the_mitsuhiko 10y agoWith read-committed you see the old state.
- anarazel 10y agoIn postgres (and a fair number of other databases) you'll not see that anomaly, even with read committed. Usually you'll want to have stricter semantics for an individual query, than for the whole transaction.
- teraflop 10y agoI think this is incorrect, but it's not as simple as the other replies are making it out to be. Under read-committed isolation, within a single operation, you must not be able to see inconsistent data. So if you do "SELECT <star>" on a table while rows are being updated, you're guaranteed to always see either the old value or the new value. But if you do two separate statements, "SELECT <star> WHERE value='new'" and "SELECT <star> WHERE value='old'" in the same transaction, you may not see the row because its value could have changed. Serializable isolation prevents this case, typically by holding locks until the transaction commits. It gets messy because the ANSI SQL isolation levels are of course defined in terms of SQL statements, which don't map perfectly to the operations that a MongoDB client can do. Mongo apparently treats an "index scan" as a sequence of many individual operations, not as a single read. So you could argue that it technically obeys read-committed isolation, but it definitely violates the spirit.
- jsemrau 10y agoWeird to see that Mongo is still around. We started to use them on a project ~4 years ago. Easy install, but that's where the problems started. Overall terrible experience. Low performance, Syntax a mess, unreadable documentation. They seem to still have this outstanding marketing team.
- vegabook 10y agoI have moved from Mongo to Cassandra in a financial time series context, and it's what I should have done straight from the getgo. I don't see Cassandra as that much more difficult to setup than Mongo, certainly no harder than Postgres IMHO, even in a cluster, and what you get leaves everything else in the dust if you can wrap your mind around its key-key-value store engine. It brings enormous benefits to a huge class of queries that are common in timeseries, logs, chats etc, and with it, no-single-point-of-failure robustness, and real-deal scalability. I literally saw a 20x performance improvement on range queries. Cannot recommend it more (and no, I have no affiliation to Datastax).
- pixelmonkey 10y agoGenuinely curious: when you say "it brings enormous benefits to a huge class of queries that are common in timeseries", what are you referring to, exactly? I run Cassandra in production and I love its operational simplicity, scale-out design, and write performance. But I think its support for time series is perhaps over-hyped. To me, it seems the only queries you can run in Cassandra is a key lookup (partition key row get) and a column slice (partition key row get filtered by an ordered range of columns). This allows for a certain time series use case e.g. where each row represents exactly one series, and where the only thing you want to do with a series is to get its raw values. But it doesn't allow for many of the things I personally think of when I think about "time series queries", e.g. resampling, aggregates, rollups, and the like.
- vegabook 10y agoI am referring to anything that resembles a range query, ie, where you require a bunch of contiguous information queried on a single key. Think "give me all of this person's chat entries from x time to y time", or indeed "give me all this topic's comment entries from x time to y time" (but not both - only one of the above would be efficiently stored - you decide which it would be). Cassandra, as you know, forces a certain amount of "low level awareness" requirement on the programmer because to tap into its uniqueness, you need to know how you will query stuff, so that Cassandra will ensure that the most common range queries are contiguously stored in rows. All other databases hide the on-disk storage order from you in an abstraction, and you can find atomisation causing inefficiency. Cassandra forces you to think about it, and in return, guarantees contiguous storage order on disk along one of your keys so that along that key, retrieval is lightning fast as it requires only one pass. Basically, both spinning disks but also SSDs, are in essence, 1d media (ie, a lot in common with tape) in the sense that along one dimension you can read stuff massively fast, but as soon as you need to seek (ie start using dimension 2), even on an SSD, your performance dramatically declines. Cassandra forces you to think about your queries so that they will be "aligned" along the most efficient direction on disk. Now agreed that if your queries cannot be aligned along said direction, then Cassandra drops to being no better than all the others, and penalises you with some complexity. That includes some examples of aggregrates, resampling etc (though I would argue that the order of magnitude contiguous read still helps these). Some of this can be mitigated with denormalisation ie: storing stuff more than once, in transposed or sub-sampled orders, something that relational DB purists will hate, with some justification (potential for inconsistency). FWIW Riak TS sounds promising with automatic "blob" style storage etc and resampling capabilities which might take Cassandra on quite explicity and in a higher level, more convenient way. I am about to evaluate it because I agree with you that the resampling capability in particular could be better supported in Cassandra, though ultimately, both databases will still be limited by the underlying D1 v D2 "contiguous v seek" capabilities of the storage so I'm not expecting miracles from Riak. By the way, I'm not even touching on Cassandra's scale-out ease. More perf needed? Literally just add boxes though it would be unfair not to comment on the cost of this, which is Cassandra's node-level consistency tradeoffs for very recently added data, and which is, if I recall correctly, why Facebook went to Hbase. You can force consistency at the query level, but performance can suffer.
- jtchang 10y agoThis single issue would make me not want to use MongoDB. I'm sure there are design considerations around it but I rather use something that has sane semantics around these edge cases.
- wzy 10y agoDoes Meteor support a proper database system yet, a la. MySQL or Postgres?
- dan_ahmadi 10y agoYes - with Apollo/GraphQL (currently available as a technical preview): http://docs.apollostack.com/apollo-client/meteor.html http://docs.apollostack.com/apollo-client/meteor.html I recommend you check out the Apollo Meteor Starter Kit: https://github.com/apollostack/meteor-starter-kit https://github.com/apollostack/meteor-starter-kit
- wzy 10y agoNoticed how I referenced 2 proper RDBMS in my question? Then how you proceeded to introduce another flavour-of-the-month
- stickfigure 10y agoNoticed how I referenced 1 proper RDBMS in my question? Fixed that for ya. Yes yes I know, HN is not Reddit. But there's still something fishy about declarations of what is/isn't a "proper" RDBMS.
- sotojuan 10y agoI haven't looked at Apollo, but GP should've explained that GraphQL is not a database and can be hooked up to any backend, so with it I'm guessing you can use any kind of database in Apollo/Meteor apps. Still, kind of weird. Another reason why I never bothered with Meteor.
- wzy 10y agoI remember when Meteor was the JavaScript flavour-of-the-month and everyone was saying it will kill Rails. I wanted to believe so I looked into Meteor, then I saw its dependence on MongoDB... Nope!
- TimPrice 10y agoThe article is interesting, but title is fud. Besides, all this is not unexpected: > How does MongoDB ensure consistency? > Applications can optionally read from secondary replicas, where data is eventually consistent by default. Reads from secondaries can be useful in scenarios where it is acceptable for data to be slightly out of date, such as some reporting applications. https://www.mongodb.com/faq https://www.mongodb.com/faq
- ahachete 10y agoStrongly biased comment here, but hope its useful. Have you tried ToroDB (https://github.com/torodb/torodb https://github.com/torodb/torodb)? It still has a lot of room for improvement, but it basically gives you what MongoDB does (even the same API at the wire level) while transforming data into a relational form. Completely automatically, no need to design the schema. It uses Postgres, but it is far better than JSONB alone, as it maps data to relational tables and offers a MongoDB-compatible API. Needless to say, queries and cursors run under REPEATABLE READ isolation mode, which means that the problem stated by OP will never happen here. Problem solved. Please give it a try and contribute to its development, even just with providing feedback. P.S. ToroDB developer here :)
- nimrody 10y agoHow does ToroDB handles sharding across multiple instances?
- ahachete 10y agoRight now ToroDB handles sharding at the backend (RDBMS) level, with those dbs that support that. There's currently a Greenplum-based backend on the works, that obviously handles sharding by itself. Also CitusDB is on the roadmap. At a later release, we also plan to natively support MongoDB's sharding protocol.
- partycoder 10y agoThis use-case is not something that you would use MongoDB for. Try Zookeeper. This being said, I would feel embarrassed to post this on behalf of the engineering department of a company. This post is just a very illustrated way of saying "we have no idea about what we are doing and our services are completely unreliable". This is so bad that is more of an HR problem than it is an engineering problem.
- teraflop 10y agoDid you miss the part about how they're running a hosting platform that stores details about the status of containers for all their customers? Zookeeper is fine for things like service discovery that deal with a bounded amount of data. You don't want to use it for something where the amount of data depends on, say, how many containers your customers decide to start. Every ZK server keeps all of its data on the Java heap, so if your data gets too big, pow. How big is too big? Don't worry, you'll find out the hard way sooner or later! Plus, there's no sharding -- every write operation has to be acknowledged by a majority of nodes in your cluster. So for write-heavy workloads (which is what I would expect a service status dashboard to experience) your cluster actually gets slower if you try to add more machines.
- partycoder 10y agoZookeeper slows down when you add nodes since quorum/consensus is larger. You can mitigate some of this with non-voting nodes (observer nodes) but only up to certain extent. So yes, a single Zookeeper cluster won't scale horizontally. But that doesn't limit the amount of independent clusters you can have. The reason I suggested Zookeeper is because it offers you ephemeral nodes, which is convenient to mark stuff as unavailable.
- shruubi 10y agoSeriously, who looks at MongoDB and thinks "this is a sane way of doing things"? To be fair, I've never been much of a fan of the whole NoSQL solution, so I may be biased, but what real benefits do you gain from using NoSQL over anything else?
- avital 10y agoI believe this is solved by Mongo's "snapshot" method on cursors: https://docs.mongodb.com/v3.0/faq/developers/#faq-developers-isolate-cursors https://docs.mongodb.com/v3.0/faq/developers/#faq-developers...
- glasser 10y agoIf I understand correctly, this method says "only scan the built in _id index, not any other index". Which means that you will not hit this index-specific bad behavior, but also that you won't get the performance characteristics of using an index.
- d3ckard 10y agoI worked with MongoDB quite a lot in context of Rails applications. While it has performance issues and can generally become pain because of lack of relations features, it also allows for really fast prototyping (and I believe that Mongoid is much nicer to work with than Active Record). When you're developing MVPs, work with ever changing designs and features, ability to cut off this whole migration part comes around really handy. I would however recommend to anybody to keep migration plan for the moment the product stabilizes. If you don't, you end up in the world of pain.
- xenadu02 10y agoUse of MongoDB at PlanGrid is probably the single worst technical decision the company ever made. We've migrated our largest collections to Postgres tables and our happiness with that decision increases by the day.
- rjurney 10y agoMongo is hilarious. Ease of use is so important, we just don't much give a shit that it has all these gaping holes and flaws in it.
- fiatjaf 10y agoCouchDB is simple and reliable. You can understand it from day one. I can't imagine why it isn't being used.
- skeoh 10y agoI really want an excuse to build something with CouchDB and PouchDB (https://pouchdb.com/ https://pouchdb.com/). Can you expand on your experiences with it?
- mikekchar 10y agoI'm not the original poster, but I can give you some of my limited experience with CouchDB from an application I inherited. The original idea for the project still seems like a good idea to me. Basically they wanted to record events that came into the system and store them in a write only ledger. Then they wanted to version every change so that you have an audit trail. Finally they wanted to be able to create views of that ledger to create the kind of data that they would work with on a day to day basis. For this, CouchDB seems like a perfect fit. Unfortunately, it didn't work out as well as one might hope because the people who implemented the idea didn't seem to be able to resist using the DB the way they would use a relational db. Instead of maintaining the concept of a write only ledger, they started to use it as a data store for things that were ephemeral. Also, instead of replicating the db, using a view to create a new db that was optimal for certain queries, they wrote a huge number of views in the main db. Finally they organised the views by relation rather than by use, so you would have 60-80 views in the same design document that would have to be reindexed if one of them changed. The result was something with very poor performance and where the storage for the indexes was more than an order of magnitude more than the storage for the documents themselves. CouchDB is also not super speedy at the best of times. There is a lot of latency involved in serializing the documents and farming them out to view servers, etc. So it takes a good 10 minutes to process a million documents, but you will find that your CPU is chugging along at 30-40% utilisation. Having said all that, one of the things I want to try (but have only done some preliminary trials with) is to keep the concept of the write only ledger, but to replicate the db into several views of the data (some with severely restricted content). Then instead of building something like a rails application to farm out the data, make "couch applications" where you serve the HTML and JS directly from attachments on documents in the DB. In fact, I've written a React application to allow users to interact with portions of the data and it was quite simple. Then you can write a really small coordinating application to allow users to navigate to the parts of the system (really single page apps) that they want to use. Again, the nice thing about this is that you have a write only data store with versioning and the ability to audit history. You have views that allow you to interact with a small subset of the overall data. You can easily write single page applications where deployment is as easy as pushing a document to the DB. Replication is relatively cheap and you can move expensive view creation to restricted versions of the DB. You can stick the whole thing behind a load balancer and scale it as cheaply as setting up a new replication (again just another document in your DB). But, I will warn you. Don't use it like you would a relational DB, or else you will be in for a world of hurt. Especially you will see comments in this thread about migrations. If you are migrating your data, by definition you do not have a write-only-with-versioning application. Your application will have to deal with multiple versions of data or else you will not have the ability to audit history. If you do not care about this, then possibly there are better solutions than this.
- lath 10y agoA lot of Mongo DB bashing on HA. We use it and I love it. Of course we have a dataset suited perfectly for Mongo - large documents with little relational data. We paid $0 and quickly and easily configured a 3 node HA cluster that is easy to maintain and performs great. Remember, not all software needs to scale to millions of users so something affordable and easy to install, use, and maintain makes a lot of sense. Long story short, use the best tool for the job.
- ahi 10y agoThis has also been my experience. Millions of large documents on a single (beefy) node with a single user it's been fine. Although, the sysadmins had previously left me with flat file xml on shared storage so the bar was pretty low.
- danbmil99 10y agoHa ha, had to scroll all the way down to find a positive comment. It's actually a great paradigm; not every problem fits into a relational box.
- wizardhat 10y agoTLDR: He was reading the database while another process was writing to it. Why all the Mongo hate? I'm sure this would happen with other databases.
- Amezarak 10y agoNo, this does not happen with any relational database I've worked with.
- deleted 10y ago[deleted]
- insulanian 10y agoEver heard of transaction isolation?
- mynewtb 10y agoLol no, ACID does not allow nonsense like that.
- Osiris 10y agoI hear a lot about MongoDB's reliability issues. How do CouchDB or other document store database compare in terms of reliability and consistency?
- rdtsc 10y agoCouchDB is rock solid. Used it for 5 years now. Never got corrupted data. Has master-to-master replications. Really shines in sometimes offline operation mode (with re-sync on reconnect). I use that extensively to build custom replication cluster topoligies (overlapping rings, star, hierarchy) etc. Has HTTP interface so easy to build clients for. Transactions are per document only. So have to design your application to accomodate it. Raw single document write speed is not as fast as Mongo or Postgres. But I noticed in a concurrent environment, multiple connections writing it scaled pretty well. Moreover, CouchDB 2.0 will have built-in clustering from code donated by Cloudant. And it will also have a similar query language like MongoDB (instead of having to use Javascript / Python / Other map-reduce functions).
- paradox95 10y agoShould an infrastructure company be advertising the fact that it didn't research the technology it chose to use to build its own infrastructure? All these people saying Mongo is garbage are all likely neckbeards sysadmins. Unless you're hiring database admin and sysadmins, Postgres (unless managed - then you have a different set of scaling problems) or any other tradition SQL store is not a viable alternative. This author uses Bigtable as a point of comparison. Stay tuned for his next blog post comparing IIS to Cloudflare. Almost every blog post titled "why we're moving from Mongo to X" or "Top 10 reason to avoid Mongo" could have been prevented with a little bit of research. People have spent their entire life working with the SQL world so throw something new at them and they reject it like the plague. Postgres is only good now because they had to do some of the features in order to compete with Mongo. Postgres been around since 1996 and you're only now using it? Tell me more about how awesome it is.
- glasser 10y agoMy goal in writing this post was not to convince people to use or not use MongoDB, but to document an edge case that may affect people who happen to use it for whatever reason, which as far as I could tell was inadequately documented elsewhere.
- paradox95 10y agoOnly the first line was directed at you - and it was more in jest. Everything else was directed more at the other commenters and Mongo detractors in general.
- rgo 10y agoEverytime I hear arguments for going back to relational databases, I remember all the scalability problems I lived through for 15 years in relational hell before switching to Mongo. The thing about relational databases is that they do everything for you. You just lay the schema out (with ancient E-R tools maybe) load your relational data, write the queries, indexes, that's it. The problem was scalability, or any tough performance situation really. That's when you realized RDBMSs were huge lock-ins, in the sense that they would require an enormous amount of time to figure out how to optimize queries and db parameters so that they could do that magic outer join for you. I remember queries that would take 10x more time to finish just by changing the order of tables in a FROM. I recall spending days trying different Oracle hints just to see if that would make any difference. And the SQL-way, with PK constraints and things like triggers, just made matters worse by claiming the database was actually responsible for maintaining data consistency. SQL, with its naturalish language syntax, was designed so that businessman could inquire the database directly about their business, but somehow that became a programming interface, and finally things like ORMs where invented that actually translated code into English so that a query compiler could translate that back into code. Insane! Mongo, like most NoSQL, forces you to denormalize and do data consistency in your code, moving data logic into solid models that are tested and versioned from day one. That's the way it's supposed to be done, it sorta screams take control over your data goddammit. So, yes, there's a long way to go with Mongo or any generalistic NoSQL database really, but RDBMS seems a step back even if your data is purely relational.
- wvenable 10y agoI've been in the opposite situation and I couldn't disagree more. But I will say this, it's always possible to take an RDBMS model and de-normalize it and use it like a NoSQL database (like reddit does, for example) but it's not possible to go the other way.
- lloyd-christmas 10y ago> but it's not possible to go the other way Why not? We do exactly that. We prototype in mongo and then migrate to postgres when we're comfortable with where the app is headed.
- hardwaresofton 10y agoIf you're currently using MongoDB in your stack and are finding yourselves outgrowing it or worried that an issue like this might pop up, you owe it to yourself to check out RethinkDB: https://rethinkdb.com/ https://rethinkdb.com/ It's quite possibly the best document store out right now. Many others in this thread have said good things about it, but give it a try and you'll see. Here's a technical comparison of RethinkDB and Mongo: https://rethinkdb.com/docs/comparison-tables/ https://rethinkdb.com/docs/comparison-tables/ Here's the aphyr review of RethinkDB (based on 2.2.3): https://aphyr.com/posts/330-jepsen-rethinkdb-2-2-3-reconfiguration https://aphyr.com/posts/330-jepsen-rethinkdb-2-2-3-reconfigu...
- brightball 10y agoHow does it compare to Couchbase? That seems to be lighting the world on fire in that space lately.
- neumino 10y agoRethinkDB's query language is infinitely better (among many things).
- hardwaresofton 10y agoI'm not sure if lack of overbearing marketing speak counts for something, but RethinkDB definitely has that going for it. I'm not an expert on couchbase (and neither on RethinkDB, to be frank, though I am a huge fan), but here's what RethinkDB has going for it: - Changefeeds - easily open a persistent connection to the server and get updates when the results of an almost arbitrary query changes. - Joins - Expressive query language that is pretty functionally minded, really shines in their clojure/haskell drivers - Excellent client libraries, well maintained - Geospatial queries/objects - Amazing admin interface (it has been amazing for a long time, too, not a recent change) - First class consideration of replication & sharding (it is not a bolt-on in any way shape or form) - API-driven cluster configuration - API driven permissions management (this is relatively new) - Excellent, easy to follow documentation There are more things, but this is just what I can think of off the top of my head. The team at RethinkDB is also just great -- I've met them in person and gotten help from them and they're straight shooters. They've also got this great project coming up called Horizon: https://www.youtube.com/watch?v=Sb1lH5mvYmU https://www.youtube.com/watch?v=Sb1lH5mvYmU Also they have a video up with a member of the team building a realtime game with React Native: https://www.youtube.com/watch?v=xRK0SYSgVF0 https://www.youtube.com/watch?v=xRK0SYSgVF0 Maybe someone who is very familiar with couchbase can help make a list... I'll start it off: - Custom query language - First class consideration for scale -- replication and sharding
- mouzogu 10y agoIs MongoDB really that bad? I am someone just getting into Meteor Js and it seems like moving from MongoDB would make it Meteor trickier to learn. Is it difficult to switch to an alternative? Thanks
- nevi-me 10y agoIt's not, go ahead and use it, learn and gain experience. It's not a replacement for SQL databases. It doesn't have joins and the biggest issue academics and as sysadmins have is it's not fully ACID compliant, so no transactions for example. If I was writing this 2 years ago, I would say horizontal scaling is much easier. Add a node to your replica, watch it catch up, and continue. Have data stored in an array 4 levels deep? Mongo will find it for you. It's only difficult to switch to an alternative to the extent that you've convolved your schema in an unfriendly way. Most migration entails normalising your data into different SQL table and exporting it. Not rocket science as people make it seem to be. I use SQL at work, Oracle, SAS excuse, a bit of MySQL and sometimes Postgres - I'm a consultant. I have tried some NoSQL DBs but always come back to Mongo, for personal projects. I've done a few prototypes for clients using Mongo, but those are almost always for geospatial support.
- mouzogu 10y agoThanks
- acarrera 10y agoIf you were inserting changes in the status you'd have much better data and never incur in such issues.
- xchaotic 10y agoUnless you want to code every rdbms and enterprise feature in the application layer, don't use Minho, use Postgres or Use Marklogic. It is 'nosql', but it is acid compliant and uses MVCC so what the queries return is predictable.
- hendzen 10y agoActually, if this lack of index update isolation is correct, you can get the matching document zero, one or multiple times!
- vs2370 10y agoI am pretty excited about cockroachDb. Its still in beta so not suggested for production use yet, but its being designed pretty carefully and by a great team.. check them out cockroachlabs.com
- cachemiss 10y agoMy general feeling is that MongoDb was designed by people who hadn't designed a database before, and marketed to people who didn't know how to use one. Its marketing was pretty silly about all the various things it would do, when it didn't even have a reliable storage engine. Its defaults at launch would consider a write stored when it was buffered for send on the client, which is nuts. There's lots of ways to solve the problems that people use MongoDB for, without all of the issues it brings.
- zamalek 10y agoI really agree with your sentiments, that first paragraph is a great quote. I grew quite an adverse to MongoDB after researching it. While I never found this specific caveat, I found other very worrying decisions. > reliable storage engine By "reliable" I assume you mean "consistent?" While MongoDB claims that it's CP (which it's not, as per the article) there's nothing wrong with inconsistent databases (AP, e.g. CouchDB). Mathematically there is no reason for MongoDB to behave like this. It's fundamentally broken; it's neither AP nor CP.
- cachemiss 10y agoI actually mean reliable. Its probably different now, but at launch, the defaults were fsync'ing every 30 seconds or so. It would literally just apply the change to an memory mapped buffer and just fsync it once in a while. They did that so they could look good in benchmarks, and it's why they recommended so strongly that your memory completely fit in RAM or else things would fall apart (pro-tip, any system that recommends that has a poorly designed storage engine). They also screwed up the consistent side of things as well.
- geoPointInSpace 10y agoI'm prototyping in meteor using MongoDB and Compute Engine. I have two VM instances in google cloud platform. One is a web app and the other is a MongoDB instance. They are in the same network. The connection I use is their internal IP. Can other people eaves drop between my two instances?
- tinix 10y agoY'all know other storage engines exist, right? I searched the comments for "percona" and found nothing... Figures. Meanwhile, https://github.com/percona/percona-server-mongodb/pull/17 https://github.com/percona/percona-server-mongodb/pull/17
- wvenable 10y agoI wonder how much data they are storing and in what pattern that they actually need a NoSQL database. I'm curious why someone would make that choice.
- deleted 10y ago[deleted]
- spullara 10y agoIt literally returns wrong answers for queries. I can't believe anyone this thread is defending it.
- twunde 10y agoThe real problem with Mongo is that it's so enjoyable to start a project with that it's easy to look for ways to continue using it even when Mongo's problems start surfacing. I'll never forget how many problems my team ended up facing with Mongo. Missing inserts, slow queries with only a few hundred records, document size limits. All while Mongo was paraded as web scale in talks.
- 1879325235 10y ago> 1) Have migrations (except they're going to be some scary ad hoc nodejs script that loop through your document store and modify fields on the fly). Unless you have never written migrations in SQL before you would know that they are even scary and a big bunch of sql. What is even worse is that, despite nobody here acknowledging it, must SQL databases are used with an ORM. Now in your migration script you have to hack SQL and language-based ORM commands together into a big pile of shit. > 2) Have schemas (except they'll be implicit and undocumented) Actually they can be just as explicit and documented. Mongo even has schemas now I think. But you can see the schema in the class definitions of your code. Once again, most people use an ORM and do this anyway. > 3) Constraints (except they'll be hidden inside your app logic, and violating them will cause data corruption). Errr... SQL constraints are very limited and literally every application has additional constrains embedded in the code. The PostgresSQL propaganda on this forum is such bullshit. None of you have any idea what you are talking about even though you are supposed to be hackers.
- nevi-me 10y agoI've just commented with pretty much the same as you're saying. Some people sound like they haven't even tried Mongo. It's great that PG has JSON support now, but a few years ago people were having together stuff on hstore, storing their geo objects as binary blobs that you can't read without extensions etc. SQL databases have been playing catch up with JSON, MSSQL being the worst. Even with JSON support, it's still a pain looking at some SQL queries that have to be written. If postgres does it for you, don't go bashing everything else that tries to be an alternative I feel.
- Renner1 10y ago> If postgres does it for you, don't go bashing everything else that tries to be an alternative I feel. People bash on MongoDB because they care. People just want to raise awareness that there are two types of MongoDB users: Those who understand it is a deeply broken system, and those who haven't used it thoroughly. The truly careless and anti-social approach would be for people to leave no negative comments about MongoDB and let others fall into the same trap.
- deleted 10y ago[deleted]
- bbcbasic 10y agoAhhh the Trough of Dissolutionment! [1] https://setandbma.wordpress.com/2012/05/28/technology-adoption-shift/ https://setandbma.wordpress.com/2012/05/28/technology-adopti...
- throoooowaway 10y agoBut is your database webscalwebscale? MongoDB is a web scale database.
- bbcbasic 10y agoYou've seen THAT YouTube video then!
- throoooowaway 10y agoBut is your database webscalwebscale? MongoDB is a web scale database.
- 1879325237 10y agoThis is exactly what I thought. You are a DBA. You are not a programmer. Whenever I read these threads I think "either these people are all DBAs or they have an extreme emotional attachment to a particular database". There it is - this is a thread full of DBAs rebelling against a database that takes away their power and responsibility.
- bbcbasic 10y agoMongo requires no administration then?
- gaius 10y ago"Power" in an organisation is the formal authority necessary to fulfill your reponsibilities. The choice of a particular database technology doesn't make the responsibility for the security, availability etc of the data "go away". As I say you can outsource the mundane tasks like "doing backups" but the buck still stops with someone. If you don't know who that is, it might be you...
- chris_wot 10y agoI'm both, and I'm as equally at home writing C++ code as I am writing T-SQL or PL/SQL. Frankly, getting to grips with relations and digging into functions made me a better programmer.
- daveguy 10y agoProgrammers should have more than a passing understanding of database administration. They should understand regularization / constraints / ACID / SQL / etc. This was core curriculum in my CS undergraduate. More importantly they should understand why they are important and when they are needed. Most of the commenters here are probably programmers who understand these things -- or programmers who have gotten burned by database stew and learned the hard way. It seems like you would to well to take a deep breath, drop the antagonistic view of traditional databases (just another tool), and educate yourself on their use and implementation.
- Jweb_Guru 10y agoThe people who work on databases are (in my estimation) usually up there with those who work on operating systems, compilers, and video games. Most programmers are simply not dealing with very interesting constraints in terms of latency, extensibility, throughput, storage, concurrency, or other challenging requirements, but database people are. Don't be irritated when they ask for better tools.
- aavotins 10y agoMongoDB reminds me of an old saying that if you have a problem and you use a regex to solve it, you end up with two problems. I have personally used MongoDB in production two times for fairly busy and loaded projects, and both times I ended up to be the person that encouraged migrating away from MongoDB to a SQL based storage solution. Even at my current job there's still evidence that MongoDB was used for our product, but eventually got migrated to PostgreSQL. Most of the times I've thought that I chose the wrong tool for the right job, which may be true, but still leaves a lot of thought about the correct application. Right now I have a MongoDB anxiety - as soon as I start thinking about maybe using it(with an emphasis on maybe), I remember all the troubles I went through and just forget it. It is certainly not a bad product, but it's a niche product in my opinion. Maybe I just haven't found the niche.
- MoOmer 10y agoI literally brought up the regex joke in a meeting yesterday. A data warehouse was built on top of Mongo, and I get to help clean up the mess.
- clentaminator 10y agoAn interesting read into the development of a project that started using MongoDB and switched to PostgreSQL after eight months in production: http://www.sarahmei.com/blog/2013/11/11/why-you-should-never-use-mongodb/ http://www.sarahmei.com/blog/2013/11/11/why-you-should-never...
- danbmil99 10y agoOh, the fud of it. The behavior is well documented here https://jira.mongodb.org/browse/SERVER-14766 https://jira.mongodb.org/browse/SERVER-14766 and in the linked issues. Seasoned users of mongodb know to structure their queries to avoid depending on a cursor if the collection may be concurrently updated by another process. The usual pattern is to re-query the db in cases where your cursor may have gone stale. This tends to be habit due to the 10-minute cursor timeout default. MongoDB may not be perfect, but like any tool, if you know its limitations it can be extremely useful, and it certainly is way more approachable for programmers who do not have the luxury of learning all the voodoo and lore that surrounds SQL-based relational DB's. Look for some rational discussion at the bottom of this mongo hatefest!
- lars_francke 10y agoI wouldn't call a JIRA ticket good documentation. While I agree that it's good to know the limitations of the tools you chose those limitations should be clearly spelled out in the documentation. I don't think most programmers have the luxury of learning all the voodoo and lore that surrounds MongoDB from JIRA tickets and blog posts.
- danbmil99 10y ago> I don't think most programmers have the luxury of learning all the voodoo and lore that surrounds MongoDB from JIRA tickets and blog posts. That's how I learned everything I know about most FOSS products I have encountered - through the code pages and social media surrounding the project. Pretty much everything about the mongodb hate derives from their marketing and sales. The truth is, they've obviously stumbled onto something the market wants, otherwise they would never have become so successful. For me, as a long-time programmer with no database experience, the mental mapping of JSON constructs as both data and query language was far easier for me to absorb than the relational model, which didn't fit the paradigms that I was used to. At my present gig, we've used Mongo DB for two years, scaling up to quite a large production setup. Like any technology it has strengths and weaknesses, but it has not been the utter failure that readers of Hacker News would be led to expect. We adopted it knowing quite a bit about its history, and it has turned out to be an excellent choice that has held up over time. Periodically we've considered switching to postgres, and we may do so for part of our stack. But for the core jobs of data collection and batch processing data with fluid schema, I'm pretty sure we will stick with mongodb for the duration. It's just a tool, folks.
- alkonaut 10y agoSo it's a bit weak in the design department, offers a bit less rigid semantics than one might hope, and from the start it's a technology that was almost a reaction to the rigid and enterprise-y of old. Mongo reminds me a wee bit of JS...
- opless 10y agoBut it's web scale! </sarcasm>