4 ms·
Just-In-Time Scalability
- cte 18y agoPPT makes me sad :( I want to read about your JIT scalability, but I'm a lazy linux user!
- eries 18y agoGood point. Sorry about that. I just posted a PDF version, let me know if it works for you: http://www.speakeasy.org/~ericries/Just-In-Time%20Scalability_%20Agile%20Methods%20to%20Support%20Massive%20Growth%20Presentation.pdf http://www.speakeasy.org/~ericries/Just-In-Time%20Scalabilit...
- cte 18y agoSweet! Thanks.
- jacobscott 18y agoInteresting slides, I like the rocket ship/driving analogy. A few quibbles/clarifications: From slide 13: "Learning: cross shard joins & transactions aren’t required" Do you mean that there was never a time where you had to do a join or transaction between shards? Would this in turn imply that your data was embarrassingly parallel? Also, just-in-time here seem like it refers to rapid architecture iterations. I was expecting something ec2/appengine-ish, maybe this is just a problem of terminology overloading.
- eries 18y agoI'm not sure who would be embarrassed if our data did turn out to be that parallel :). IMVU has a pretty typical data profile for a social network + IM site. We worried all the time that we would have to build a really fancy distributed transaction engine. We started by just doing all of those transactions on the centralized master database. Since we were incrementally moving load off the master as-needed, we just did those tables last. Turns out, that approach lasted more than three years, much longer than we ever anticipated given the rapid growth we experienced over that period. Of course, we did eventually have to tackle it, which we did by building a "durable message-passing" system. As for terminology, you're completely right. We started using this term a few years ago, so it wasn't quite as overloaded as it is today. I am sometimes surprised how much more time and energy seems to be given over to thinking about the hardware, as opposed to the architectural, aspects of scaling.
- jacobscott 18y agoParallel data stuff makes sense. As for the hardware aspects of scaling: I don't think that people are thinking about hardware much beyond "more boxes". My understanding of the hype surrounding ec2 was that it mostly concerned the rich API/management options for configuring your scaling (in software). If you mean that people are ga-ga over how many of Amazon's machines they could scale to but not how they can add layers of caching and database partitioning -- that's probably accurate. IMHO the key point is that Amazon has a generic framework. The architecture you choose to scale with probably ends up being a lot more ad-hoc, "driven" by your problem domain. I'm fairly confident that if someone comes up with a rocket-ship approach to architecture that generalizes well for web startups, time and energy will be given over to thinking about it.
- furiouslol 18y agoHow did you handle the autoincrement id problem with distributed databases? Did you replace them with hashes?
- eries 18y agoWhen absolutely necessary, we would solve this by keeping global (centralized) ID's in an autoincrement table either on the master or on a vertical shard. It didn't come up that often, because most tables (in our app) are really dedicated to one specific entity (ie a customer's inventory items).
- furiouslol 18y agoUpload it to slideshare.com
- furiouslol 18y agoI really like the solutions that you guys come up with. Instead of complex enterprisey solutions, you came up with simple, smart hacks like adding executable logic as SQL comment.