3 ms·
I'll give the reasons why I hate it, and they have nothing to do with it not being a "good fit" for our model or abstraction problems or anything of that nature
by mediocregopher 14y ago
I'll give the reasons why I hate it, and they have nothing to do with it not being a "good fit" for our model or abstraction problems or anything of that nature. The problem we have with it is that it's an absolute pain to maintain in a large production.
The setup alone is just bizarre, for each replica set you have to have 3 (and only 3!) config server instances associated with your set, along with possibly an "arbiter" instance depending on what circumstances you're operating under. I could do a whole rant (and have, in the past, to anyone who would listen) on the arbiter/voting system mongo uses. We've been left in a situation in the past where one of our secondaries randomly died (more on these lovely "random occurences" later), leaving the cluster with a master and an arbiter left over. Mongo decided that since there was an even number of votes (and why is my production-critical database voting on things?!) it couldn't promote a master (even though the master never actually died), dropped the master down to being read-only, leaving us completely boned. They have since fixed this issue (I think), but it definitely garnered a lot of ill-will from me, and should never EVER have even been a problem.
"Random occurrences" have always plagued us with mongo. Whether it's been random segfaults (less common with more recent versions, but oh dear god were they frequent pre-2.0), secondaries being promoted to master with no clear reason as to what sparked this (which has left us flailing twice now, when a master decided to step down while one secondary was in RECOVERING mode and the other was AWOL), config servers getting out of sync (one time one of the config server sets decided it was going to start hosting the config for an entirely different replica set, luckily no production mongos instances had to reload the config before we noticed), mongos instances getting out of sync/crashing (less common now than before, thankfully), and I'm sure much more that I'm not thinking of now.
To make all of the above about 10x worse the mongo logs are terrible. There is no distinction between INFO messages and ERROR messages in the logs, so everything has to be treated as an error of some kind. And there are many "errors" that we should "just ignore". This makes debugging pretty much impossible. Another nice feature is when you restart a mongos instance (and possibly a mongod instance too, although I could be wrong on that) all the logs from the old process get clobbered by the new one (so backup those logs folks!). It's extremely difficult to track down why slow queries are happening unless you can catch them in the act, and even then it's not trivial.
Mongo has many great qualities. It's one of the few (if the only, that I know of) that can do many of the things that it does, and for that it's very useful. But for it to have been marketed as a production-ready database was a bit disingenuous I think. It scales OK. Not well, just OK. You can make it scale if you try real hard and tiptoe around it so you don't wake it up and make it cry. I think someday in the future all of these issues will be fixed, but at the moment I don't recommend anyone use mongo for anything they plan on a significant number (200k+ users at any given moment) of users being dependent on.
- jasonmccay 14y agoInteresting feedback ... I agree on the log messages that look terrifying yet, when googled, are like "oh, that's normal, don't worry about that." I also agree that the 2.x versions of MongoDB have been a great step forward for stability and performance. On the logging, not sure what you mean by the old logs get clobbered. Is this, perhaps, some house-cleaning job that you have on your server? When we stop/start processes, the logs are appended and keep going. On the issues with primary/secondary and state changes, do you host your databases on Amazon? We have often found that when we have this issue, it was pinpointed back to some temporary intra-zone networking glitch with AWS ... either with their name resolution or a blip in one zone being unable to see another zone in the window of time that MongoDB has set for the response. These often (if not all the time) go unreported by AWS until you get them to dig a bit and report back. We effectively utilize priority to keep the member we want as primary and if a state change does occur, it moves back once the full set is healthy again.
- mediocregopher 14y ago>On the logging, not sure what you mean by the old logs get clobbered. For us the logs are not opened in append-mode, just write. We don't do any house cleaning except for logrotate, but that only runs nightly and isn't the cause of what we're seeing. Like I said, it may not be for the normal mongod instance, but it's definitely true for mongos. I'll double-check on the version number of what we're running > do you host your databases on Amazon? Nope, all owned servers, over a local network on a switch we own as well. The switching happens on some clusters more then others, so it's possibly a usage related thing.
- ericcholis 14y agoAll very constructive points, thanks for this. On the point of slow queries, nailing a slow query is a bit of a mixed bag at the moment. I've found that Dex (http://blog.mongolab.com/2012/06/introducing-dex-the-index-bot/ http://blog.mongolab.com/2012/06/introducing-dex-the-index-b...) and NewRelic help diagnose trouble queries pretty well.
- jasonmccay 14y ago