5 ms·
Unless he's planning to build sharding on postgres too, I think he's missing the point.
by nsanch 14y ago
Unless he's planning to build sharding on postgres too, I think he's missing the point.
- jerrysievert 14y agowhile sharding is an important aspect of mongodb, i don't consider it the most important feature.
- jshen 14y agothat doesn't mean that others don't find it to be an extremely important feature.
- nsanch 14y agoI don't know if it's the _most_ important feature, but I wouldn't build a serious site on top of anything that didn't have some sort of built-in sharding story. With postgres you have to roll your own. If you want to bridge the gap from postgres to mongo, I think that's where you have to start.
- jshen 14y ago"but I wouldn't build a serious site on top of anything that didn't have some sort of built-in sharding story." There are many serious sites that don't need sharding.
- nsanch 14y agoFair, my statement was overly broad. Sites that are read-only or store blob data in something like S3 can often avoid sharding for quite a while and rely on machines to just get bigger over time. That said, if your site grows in some way you didn't originally anticipate and you get to a point where you need to shard, but can only do so by changing data stores, then it's sad.
- jeffdavis 14y ago"That said, if your site grows in some way you didn't originally anticipate and you get to a point where you need to shard, but can only do so by changing data stores, then it's sad." I think you're being too absolute. For instance, Instagram used sharding in postgres, and they didn't have to throw anything away or dedicate any huge engineering team to solve it.
- leothekim 14y agoThey had to put engineering effort into it. With Mongo, you don't.
- moe 14y agoWith Mongo, you don't. Bullshit. The sharding impl in MongoDB still[1] crumbles pitifully[2] under load. Regardless of sharding MongoDB still halts the world[3] under write-load. Their map/reduce impl is a joke[4][5]. If you had done the slightest research you'd know that every single aspect that you need to scale out Mongo is either broken by design or so immature that you can't rely on it. MongoDB may be fine as long as your working set fits into RAM on a single server. If you plan to go beyond that then you'd better start with a more stable foundation - or brace yourself for some serious pain. [1] https://groups.google.com/group/mongodb-user/browse_thread/thread/131720fe4fc73f4c https://groups.google.com/group/mongodb-user/browse_thread/t... [2] http://highscalability.com/blog/2010/10/15/troubles-with-sharding-what-can-we-learn-from-the-foursquare.html http://highscalability.com/blog/2010/10/15/troubles-with-sha... [3] http://2.bp.blogspot.com/_VHQJkYQ5-dY/TUO3RAn8SNI/AAAAAAAABqs/FJKgl_HgBWA/s1600/mongo-rw-fail.png http://2.bp.blogspot.com/_VHQJkYQ5-dY/TUO3RAn8SNI/AAAAAAAABq... [4] http://stackoverflow.com/a/3951871 http://stackoverflow.com/a/3951871 [5] http://steveeichert.com/2010/03/31/data-analysis-using-mongodb-map-reduce.html/ http://steveeichert.com/2010/03/31/data-analysis-using-mongo...
- nsanch 14y ago[1] looks like an example where the data didn't fit in RAM. Mongo works best when data fits in RAM or if you use SSD's. Yes, it's sub-optimal. [2] is from a year and a half ago. It doesn't belong in a sentence that includes the word "still." I work at foursquare, btw. Those outages happened on my first and second days at the company. I wasn't so keen on mongo then either. We've gotten much better at administering it. Basically all our data is in mongo, and it has its flaws, but I'm still glad we use it. [3] is also from a year and a half ago. Mongo 2.2 will have a per-database write lock, which is at least progress, even though it's obviously not enough. Since 2.0 (or 1.8?) it's also gotten better at yielding during long writes. I have no experience with their mapreduce impl and can't speak to it.
- ibotty 14y ago"With postgres you have to roll your own [sharding]" but it will be trivial if you do not do joins.
- einhverfr 14y agoPostgres-XC is also in beta. Basically it's sharded PostgreSQL with full RI enforcement between shards, and seamless query integration. I assume they are working off the 9.2 codebase (hence the beta being the same time as Pg 9.2) but maybe it's only 9.1 (i.e. no JSON). Postgres-XC is probably the most exciting PostgreSQL-related project out there. It promises full write-extensibility across the cluster without sacrificing consistency.
- cbsmith 14y ago> I don't know if it's the _most_ important feature, but I wouldn't build a serious site on top of anything that didn't have some sort of built-in sharding story. You know people say this, but in practice I find that simple hash bucketing with a redundant pair works surprisingly well, particularly in the cloud. Yes it isn't fancy, but it is trivial to manage and debug, and you can do a lot of optimizations given such a clear cut set of partitioning rules. Your problems have to get really big before a more sophisticated mechanism really pays off in terms of avoiding headaches, and often the more sophisticated mechanisms actually cause more headaches before you get there.
- dickeytk 14y agoIt also seems to be the flakiest feature
- deleted 14y ago[deleted]
- jeffdavis 14y ago"Unless he's planning to build sharding on postgres too, I think he's missing the point." NoSQL seems to have three general claims (I'm not saying whether these are correct or not): (1) ease of administration in some cases; (2) different data model; (3) better performance or availability in some situations. The author is clearly addressing the second, and you are clearly talking about the third.
- nsanch 14y agoGood point.
- bri3d 14y agoShould be pretty trivial for him to shard collections across multiple databases with the same level of intelligence as MongoDB's automated sharding simply using the primary key, since MongoDB doesn't join. Not sure how sharding is the point of MongoDB, though - in most of the universe, sharding is a database architecture/schema-level thing, not a database-server level thing, and for good reason - it's pretty darn hard for a database server to shard effectively without some knowledge of the app layer (and which keys are likely to become hot). Personally, if I really wanted a flexible-schema "document based" database, I'd have implemented this using the FriendFeed K/V + Index model ( http://backchannel.org/blog/friendfeed-schemaless-mysql http://backchannel.org/blog/friendfeed-schemaless-mysql ) plus Postgres's HStore functionality, storing K/V per document in an HStore rather than in one giant K/V table like FriendFeed. That way I wouldn't need to use V8 and JSON parsing to run queries, and the mythic MongoDB-style "sharding" would be just as easy (just distribute the document -> hstore table across shards keyed on ID again).
- nsanch 14y agoI disagree that sharding can ever be a trivial problem if you're going to try to tackle moving data between shards while staying online. I'm not saying it's impossible, just that it's not trivial.
- bri3d 14y agoMongoDB's relatively simple approach involves continuing to use the old shard as an authoritative source (and committing updates to it) while shipping data in the background, then pushing the additional changes across and marking the new shard as "master." Such an approach wouldn't be horribly difficult to implement in SQL using a copy table and write triggers - almost identically to how SoundCloud's Large Hadron Migrator allows writes to occur over a MySQL InnoDB table that's locked for migration (but even simpler because the table schema can't conflict afterwords). The entire problem is admittedly nontrivial, since if the application happens to be writing data to the shard under migration too quickly (or the shard being migrated to dies), the server can end up in a situation where the new shard is never able to catch up and become a master. However, the easy solution (give up and retry later) is Good Enough for most situations (and is pretty much how MongoDB works).
- mark_l_watson 14y agoYou have a good point. I must admit that after many years of using and loving PostgreSQL I have never had to scale it out: that is outside of my experiences. On the other hand, old fashion MongoDB master slave configurations and now replica sets have been easy for me to set up when on just a few occasions I had to use a scaled out MongoDB setup.
- einhverfr 14y agoEasily enough done on Postgres-XC, and you still get full ACID compliance across shards as well!