12 ms·
Then where did the growth pains that necessitated Cassandra come from? I'm trying to build up a mental genealogy of the engineering decisions made so that the k
by codewright 14y ago
Then where did the growth pains that necessitated Cassandra come from? I'm trying to build up a mental genealogy of the engineering decisions made so that the knowledge isn't lost to future programmers. Most of you OG peeps have moved onto greener pastures.
- keysersosa 14y agoketralnis will likely remember this better, but by my recollection we switched to cassandra after we hit the limit of what we could do with memcachedb: http://blog.reddit.com/2010/01/why-did-we-take-reddit-down-for-71.html http://blog.reddit.com/2010/01/why-did-we-take-reddit-down-f... We were already using memcachedb as a persistent keystore, "caching" precomputed listings and comment threads. When it started becoming more trouble than it was worth to maintain, we decided to try cassandra as a beefier replacement.
- ketralnis 14y agoIt was capping memcachedb's performance that was the last straw, but it became pretty clear that we were going to similarly cap the performance of any disc that we could afford to put in a postgres write master and that we needed a solution that eliminated that bottleneck, as well as it as a single point of failure. Cassandra (or more generally the dynamo model) did that nicely.
- jedberg 14y agoI'll cover that in the talk, since I'm sure others would like to hear the answer as well. Thanks for the suggestion! If I don't cover it sufficiently, you can hit me up for more info.
- ketralnis 14y ago> Then where did the growth pains that necessitated Cassandra come from? Increased traffic? I don't think I understand the question
- raldi 14y agoMostly porn and ragecomics, wasn't it?
- ketralnis 14y agoAnd advice animals
- codewright 14y agoA little less porn these days, no?
- codewright 14y ago>Increased traffic? I don't think I understand the question :| I was trying to determine what component broke first.
- ketralnis 14y agoThere's always some component "breaking". That's just how performance works, you're always bounded by some resource. An app like reddit has most of its backend tweaked or rewritten continuously as it becomes more important to find and alleviate some new bottleneck. If you do it right, you predict the next bottleneck before it becomes one but there will always be something. Maybe writes take too long. Why? You've capped the performance of the disc in your DB master (or more importantly of the most expensive single set of discs you can afford). Maybe adding app servers proportional to site-wide traffic stopped helping you keep up at some point. Turns out you're network bandwidth bound. Maybe requests are too slow, but apps aren't using much CPU. Why? Because you're network-latency bound. Hey, it turns out that some tight loop was doing single memache GETs when it could have been doing a single multi get. That last one is just a general performance bug, but that's just it. Every bottleneck that keeps you from just adding resources is a performance bug. Assuming your app is not infinitely fast, there's always something.