9 ms·
Migrating from MongoDB to Cassandra
- fiatmoney 13y agoSounds the intended use case for ElasticSearch. "Given some input piece of data {email, phone, Twitter, Facebook}, find other data related to that query and produce a merged document of that data"
- Xorlev 13y agoI will say ElasticSearch features heavily in our infrastructure elsewhere, but for the Person API product, it's purely a primary-key lookup. My coworker wrote a bit about how the search functionality of our offering works here: http://www.fullcontact.com/blog/sherlock_search_engine_that_does/ http://www.fullcontact.com/blog/sherlock_search_engine_that_... That might make more sense why we do PK lookups.
- monkey26 13y agoThis article caught my interest as I've been reading into Cassandra. But some previous research had me thinking that Cassandra works best with under a TB/node. Is SQL still better when you have really large nodes (16-32TB) and only really want to scale out for more storage? I'm currently humming along happily with Postgres, but some of the distributed features, and availability of Cassandra look really nice.
- jbellis 13y agoCassandra 2.0 can handle 5TB per node easily, 10TB with some care. Best to scale out, not up. That said, if someone else has already made the hardware choice for you, you can always run multiple C* nodes on a single machine. I know several production clusters that fit this description.
- krenoten 13y agoIt's much more about desired usage patterns than amount of storage. Cassandra and RDBMS's differ quite a lot in how you replicate, consistency guarantees, performant read patterns, performant write patterns, how you handle recovery, etc... If you intend to bring anything to scale it helps to understand the strengths and weaknesses of the underlying architecture.
- cnlwsu 13y agoWe run at about 1TB a node and it works well (high write load things like metrics and telemetry data). But we also use SQL server where appropriate (i.e. transactional account stuff). I am a fan of using the right tool for the right job providing you have the team to support it.
- brown9-2 13y agoIs DynamoDB never a serious option for people in situations like this and already heavily on AWS?
- Xorlev 13y agoOne of the key detractors for us was the 64KB limit. "The total size of an item, including attribute names and attribute values, cannot exceed 64KB." While we don't often have values over 64KB, it's possible. We didn't want to have to store profiles separate from their metadata.
- rco8786 13y agoread first half of article get spammed leave immediately
- brightsize 13y agoSame here.
- Xorlev 13y agoAuthor here, sorry about that. It's not supposed to display on engineering posts but something must have changed/been broken recently. We know you guys aren't interested in marketing content. Again, sorry about that and thanks for bringing it up.
- deleted 13y ago[deleted]
- MBCook 13y agoI read the article on my iPhone and found the hovering "top" button in the lower right both unnecessary and very obtrusive. iOS has a standard way to jump to the top (tap top of screen), there is no need to interfere with the content.
- jbeja 13y agoOMG, it make jump a little, seriously ;).
- DigitalSea 13y agoWhen I read posts like this all but confirming MongoDB isn't the great product 10Gen make it out to be, I wonder how the heck MongoDB are still even relevant and then I remind myself of the fact that 10Gen have one of the best marketing and sales teams in the game at the moment. While MongoDB has improved greatly over previous versions, I can't help but feel if 10Gen put as much effort into improving their product as they do selling it, Mongo would be a force to be reckoned with! MongoDB is good at some things, but I think most people that try and fail with it fall into one of two camps: 10Gen sold them into it or they bought into the hype without assessing project requirements and ensuring MongoDB was a sensible choice.
- jb007 13y agoMongodb has marketed the database as a general purpose one when in reality it doesn't even come close to one. And the case is the same for all other NoSQL systems. No More general purpose database? The developers at http://www.amisalabs.com http://www.amisalabs.com are tackling today's database problems.
- ddorian43 13y agoyou know they have been tackling without a release for many months now
- CptCodeMonkey 13y agoKnowing some of the FC people first hand, MongoDB did actually serve them fairly well for a substantial amount of time. Until they started hitting max limits, it didn't really make sense to move to c. Going straight to C or something like it would have been almost cargo-cultish ( eg If we build industrial strength, we will get industrial levels of traffic ).
- threeseed 13y agoExactly. Can you imagine all of the Java/C developers telling Ruby developers that they are idiots and don't know anything about programming ? Simply because they choose a technology that is designed for developer productivity at the expense of scalability. Because that what seems to happen for every database discussion.
- jchrisa 13y agoViber, one of the largest over the top messaging apps, recently shared their conversion from Mongo to Couchbase. They ended up requiring less than half of the original servers, and better performance. If you want to see a video of their engineer telling the story, it's available here: http://www.couchbase.com/presentations/couchbase-tlv-2014-couchbase-at-viber http://www.couchbase.com/presentations/couchbase-tlv-2014-co...
- krenoten 13y agoUsually people who get burned by hypedb think twice before making the same mistake again.
- prottmann 13y agoLike always: "Use the right tool for the job". I did not think that this was MongoDB(10gen)s fault, they (viber) choose the wrong database type for their needs.
- pessimizer 13y agoThat's really easy to say, so it's important to show your specific reasoning.
- mcot2 13y agoTokumx would solve all of these issues. 2TB goes a long way in tokumx with lzma compression.
- lynchdt 13y ago"To buy us time, we ‘sharded’ our MongoDB cluster. At the application layer. We had two MongoDB clusters of hi1.4xlarges, sent all new writes to the new cluster, and read from both..." I'm curious about this. Why were you doing the sharding manually in your application layer? Picking a MongoDB shard key - something like the id of the user record - would produce some fairly consistent write-load distribution across clusters. Regardless - it seems like write-load was a problem for you, yet you sent all the write load to the new cluster - why not split it?
- Xorlev 13y agoAs explained, it was a stop-gap solution for data storage only, we did not have a problem with write load on SSDs. We were at the point that MongoDB sharding was just as difficult to deploy as moving to Cassandra, which better fit our goals of availability. MongoDB sharding isn't instant by any means for existing clusters.
- talas9 13y agoYet another shining example of throwing money and time away to work within AWS constraints when bare metal and openstack (1) would have solved it cheaper (2) and arguably faster. 1 (if you insist on cloud provisioning instances, even though it makes little sense if the resources are as strictly dedicated as they are in this case) 2 (VASTLY, over time -- these guys are pissing money away at AWS and I hope their investors know it)
- Xorlev 13y agoWe understand that AWS comes at a premium, however we find the opportunity cost of losing the agility we have on AWS at this stage of our organization more expensive than the delta in cost between moving on to our own hardware and AWS. Our organization is acutely aware of our costs and still strives to minimize them. Our move to Cassandra saved 79% over continuing to run our reserved SSD nodes & backup replica.
- bfrog 13y agoAWS in general is a waste of time and money for most standard web hosting requirements I think