5 ms·
One thing I always find interesting about these kinds of problems is that most DBs don't describe how they're implemented. It's easy to use the wrong tool, and
by arohner 13y ago
One thing I always find interesting about these kinds of problems is that most DBs don't describe how they're implemented. It's easy to use the wrong tool, and then once you learn how e.g. Mongo is implemented, it's obvious, "oh, that's why things are slow".
I'd love to see http://eagain.net/articles/git-for-computer-scientists/ http://eagain.net/articles/git-for-computer-scientists/, but for every DB technology.
- trimbo 13y agoYou nailed it. I think the issue is that people don't know what questions to ask when gathering the requirements. I'd like to know more about how this part of the article came to be: "This choice was made early on and it was supposed to be a temporary one." HOW was that choice made. What requirements were out there. I think too many people choose Mongo because they believe it's "schemaless"[1] and faster for development, but don't look at the requirements for their actual use case. [1] - There's always a schema. Either it's informally defined by your code or represented formally somewhere else.
- relistan 13y agoIt was one of those, "we need to do something now and this will work" solutions. We had a really talented consultant working for us, writing some of the early code. He was familiar with Mongo and wanted to go that route. Early on I said we should use Cassandra for this, but it took us quite some time to get to the point of being able to migrate. A testament to his code and the "this will be temporary" foreknowledge is the fact that we could swap data stores in a pretty massive sense without much in the way of outward-facing changes to our processing engine. I think the hard thing for us up front was that trying to explain Cassandra data models to someone (and this guy is really really good) and then hand the rest of the work over to them to implement now, on a contracting rate is not a trivial problem. And we needed to hurry for both deadlines and burn rate.
- trimbo 13y agoCool, thanks for the info!
- mtkd 13y agoA pattern I've seen on a lot of the negative Mongo articles has been people using it for things they probably shouldn't. I've yet to use it for any load, and am struggling to triangulate from all the articles I read on whether it does/doesn't scale efficiently. All the issues I have hit so far have been self-inflicted, it is still one of the best new technologies I've used in years - but is has taken a while to stop thinking in SQL equivalents and start thinking natively.
- otterley 13y agoFor what kind of application or workload is MongoDB superior to the alternatives?
- Xorlev 13y agoThe kind that lives on a single VPS and requires a rich query layer to bootstrap quickly.
- relistan 13y agoHopefully this article wasn't seen as bashing MongoDB. I think I mention it was pretty reliable for us even under load. We were just solving a problem with it for which it wasn't the best solution.
- arohner 13y ago> whether [Mongo] does/doesn't scale efficiently It doesn't. Three words: "global write lock". Writes block reads, reads block writes. Implications: if you run a query in production that doesn't hit an index, all traffic stops. The notablescan setting is a very, very good idea. This also means all queries must have an index, so Mongo ends up with more indexes than say, postgres would. It's impossible to configure a clustered mongo environment to not lose data: http://aphyr.com/posts/284-call-me-maybe-mongodb http://aphyr.com/posts/284-call-me-maybe-mongodb Sharding configuration is baroque, and limited.
- francesca 13y agoThe global lock was removed in 2.2 https://blog.serverdensity.com/goodbye-global-lock-mongodb-2-0-vs-2-2/ https://blog.serverdensity.com/goodbye-global-lock-mongodb-2... Now locking is on the database level.
- teraflop 13y agoAgreed. Seeing a nice, clean user-facing API tells you practically nothing about the performance characteristics and failure modes you can expect to see. The best database engine I've encountered in this respect is SQLite. It has plenty of information about what design tradeoffs it makes, and why, e.g.: http://sqlite.org/lockingv3.html http://sqlite.org/lockingv3.html http://sqlite.org/fileformat2.html http://sqlite.org/fileformat2.html http://sqlite.org/atomiccommit.html http://sqlite.org/atomiccommit.html
- mh- 13y ago> I'd love to see [git-for-computer-scientists], but for every DB technology. this is a fantastic idea. if someone gets this going I'll enthusiastically contribute.
- hcarvalhoalves 13y ago+1 It would be useful to treat databases as those big data structures, knowing the best/average/worst case, cpu vs. memory trade-offs for search, etc.