5 ms·
He's right that MongoDB could use improvements like string interning so you don't need to worry about field names. But overall, I think this article is very mis
by functional_test 13y ago
He's right that MongoDB could use improvements like string interning so you don't need to worry about field names. But overall, I think this article is very misleading.
If you use MongoDB in production, you should definitely take he time to learn about the durability options on the database side AND in your driver. By using them appropriately, you can have as little or as much as you like. Data sets larger than 100GB are no problem either -- right now I'm running an instance with a 1.6TB database.
As always, use the right tool for the right job. If you need joins/etc. and don't need unstructured data, Mongo probably isn't a great choice (even with the aggregation framework).
- yummyfajitas 13y agoFor what use cases is Mongo the right tool?
- mattdeboard 13y agoIt's "a" right tool in any case where distributed storage of unstructured data in JSON format is wanted, where database-level locks won't be an issue of concern, and availability is the primary, overriding concern.
- jbooth 13y agoI'd submit that database-level locks make any claims of availability or distributed storage a little overblown. If a single query can blow you out of the water, you're really not highly available. Although I don't have a lot of experience doing big mongodb personally, so maybe I'm missing something.
- sunils34 13y agoThat's exactly right imo. Running MongoDB in production, you end up concerned over the performance of each query (as you should be). MongoDB's profiler makes this easier to investigate. If you hit a db level lock limit, you're probably running a sub-optimal or unindexed query.
- mattdeboard 13y agoThat doesn't have anything to do with availability, at least not in the CAP theorem sense as I understand it. What I think you're talking about (being "blown out of the water" is pretty vague, though) is partition tolerance: high-latency requests that are practically indistinguishable from network partition events. I'm not sure what MongoDB returns (or how its clients react) when there are no available connections because of a lock whose duration exceeds the configured timeout. I'm pretty confident, though, that this sort of thing is covered by basic driver config.
- yummyfajitas 13y agoWhat does mongo offer over (possibly sharded) postgres for this use case? Postgres won't hit you with db-level locks and gives you master/slave replication for availability. You can also get great performance if you put the WAL on a ramdisk, which I think is roughly equivalent to how mongodb handles writes. I'm really not trying to be argumentative here, I'm just trying to understand what mongodb is for.
- mattdeboard 13y ago> What does mongo offer over (possibly sharded) postgres for this use case? Doesn't really matter for the point I'm making. It's a solution for a given set of constraints. Not the solution, or the very best tippy-top solution in all the kingdom, just a solution. Point being I can't think of a use case where this is true, but if you read the article, the author does include what he says is the only reasonable use case for using MongoDB.
- defen 13y agoIIRC, until fairly recently (well after Mongo had launched), "master/slave replication for availability" in Postgres was a bitch to set up, requiring 3rd-party tools + manual failover if the master died. It was a lot easier to get going with Mongo, which is really what matters if you're a 2 person startup just trying to validate an idea.
- mattdeboard 13y agoStrongly disagree here. MongoDB (10gen, at the time) had absolutely insane, irresponsible defaults set in all its drivers until, like, 1.8 (very recent). This is anathema to "we're just trying to validate an idea" especially when "our idea took off and now 6 months later we actually do need to scale." They've fixed it like I said but that whole "we're just using it to validate an idea" thing is a total con. "Nothing so permanent like a temporary [solution]."
- jbooth 13y agoFrom what I've seen, the data model has a lot of utility as long as you don't need super high concurrent performance. Basically, the same area as where rails is the right tool - we want easy features and rapid development, will worry about scaling later.
- yummyfajitas 13y agoWhat does mongodb offer above and beyond using postgres or redis for this use case?
- jbooth 13y agoI've barely used it, but the json document thing with a lot of random convenience functions in the query language seem to lend themselves well to rapid development. For postgres you'd be mapping to a relational schema, and for redis you'd be storing the json yourself as a blob, without any server-side manipulation capabilities (or using redis maps/sets/etc, which are awesome, but aren't as general as json). I haven't been doing very much web dev the last few years though so it's possible that my first impressions are wrong. I'm just repeating what I've been told, basically.
- deleted 13y ago[deleted]
- Sanddancer 13y agoPostgres has had a native json type and the latest version improves upon the functions given to manipulate json data. So data that may not map well to a relational schema can just be put in the json object type and handled accordingly. Also, because it's attached to an SQL engine, you can use things like views on your json data if it makes sense for the type of data you're querying. There has been a considerable amount of work put into postgres over the past few years for getting it to handle your data regardless of what it looks like. The developers seem to have a very good grasp on the fact that not all data is alike, and giving tools that will work well, and together with, all your data leads to a lot fewer headaches in the long run.
- functional_test 13y agoI use it to store a lot of historical time series data that doesn't change once written (at least, not often). I can easily achieve the write performance necessary to record the data streams live. Since it's all append-only, I don't need to worry about fragmentation. With replication, it's possible to access the data with very high throughput which is useful when the data is being accessed by a cluster, for example. I also use it as a metadata "scratch space" for highly available applications (things where failures are not acceptable and must run for days at a time). Again, with replication and automatic fail overs, I've been able to maintain 100% uptime outside of maintenance windows. Obviously that can't last, but so far it's been >2 years with no major problems. EDIT: I should point out that although the size of the metadata objects can be highly variable, since I usually had a small number of them relative to the time series, fragmentation was still not an issue.
- JulianMorrison 13y agoYou have a smallish number of documents where some particular field of fixed size gets overwritten a lot, the old values are uninteresting, and it wouldn't really be a tragedy if your data got trashed. For example, it's the player's score. You want a fixed-size, rolling backlog of time series data such as logs.
- twic 13y agoIs it better than a relational database for that?
- yummyfajitas 13y agoPostgres update performance is pretty bad. When running a big data migration, it's generally faster to copy the old table to a new temporary table and rename the temp table to the old table than it is to run an update.
- JulianMorrison 13y agoMost RDBMSs will, if you rewrite a field, write a fresh row and tombstone the old one, and clear it down in the next compaction. This is what MVCC means in practise: that the old version doesn't disappear while the new is being written. MongoDB by contrast will simply mmap that block of file, overwrite the contents, and fsync. Yes, this has obvious downsides.