59 ms·
Twitter says they fixed the Memcache calcification problem. Dormando disagrees.
- alttab 14y agoI enjoyed the jab at twitter's "we are the only website on the planet to have scaling issues" holier than thou attitude. Rails is still trying to get over the character assassination by twitter when they failed to scale it. I know first hand that rails can scale very well. Do bad carpenters blame their tools?
- kevinh 14y agoThere may be a difference between scaling well and scaling 9th on alexa well. Sometimes bad carpenters blame their tools, but sometimes the tools weren't planned for something of that scale.
- timaelliott 14y agoTwitter's attitude is laughable. Yes, you have a decent amount of traffic but you're also only dealing with ~160 uncompressed bytes (plus whatever overhead) per event. The hurdles you've overcame aren't that particularly amazing nor challenging.
- achompas 14y agoYes, you have a decent amount of traffic but you're also only dealing with ~160 uncompressed bytes (plus whatever overhead) per event. The hurdles you've overcame aren't that particularly amazing nor challenging. >42M uniques last month.[0] Are you really going to assert Twitter hasn't dealt with amazing or challenging hurdles in getting this far? [0] http://siteanalytics.compete.com/twitter.com/ http://siteanalytics.compete.com/twitter.com/ EDIT: this ignores that twitter.com is not the only Twitter client--they served 15B (!!) requests/day (!!!) as of a year ago. Not to mention metadata, instrumentation for services, logging, DB backups, and managing configuration of all of those distributed resources. Are we still talking about the ease of 160B? http://www.readwriteweb.com/hack/2011/07/twitter-serves-more-api-calls.php http://www.readwriteweb.com/hack/2011/07/twitter-serves-more...
- timaelliott 14y agoYes. In 2008/2009, another engineer and myself built an ad-platform that received around 500M impressions per day, 5M clicks per day. And it wasn't just recording a tweet or publishing out to followers. We took the user input query, had to do some keyword/relevancy targeting, geofiltering, matching to advertisers and deliver back a large result set of adverts. All within 100ms. Our platform was also apache, mod_php, memcached, mysql and rabbitmq. So definitely not the most optimal of platforms by any means. We had two colos with ~20 servers (dell r410s) at each facility. Twitter just recently announced 400M tweets/day. I'm not trying to brag about my experiences, because looking back now we made numerous amateur mistakes, but just showing that Twitter's "scale" is a joke compared to everyday challenges at any large internet ad network.
- InclinedPlane 14y agoYou are off by nearly 2 orders of magnitude from twitter's scale. They have billions of views per day and each of those views is a stream comprised of hundreds of different sub-streams.
- achompas 14y agoSee my edit. Your ad impressions reached approximately 3% of Twitter's daily request load last year. Note that those requests can serve up to 200 tweets + metadata. This doesn't account for Twitter's budding ad service, which one can assume has some of the same functionality (targeted advertising, information retrieval) as traditional ad networks.
- shadowfiend 14y agoYou understand that 400M tweets a day is the number of tweets posted to their system, right? That speaks not at all to the consumption of those tweets, which is the metric you're using for your ad platform. Additionally, they don't just deal with 160 characters, because again, somehow you're still talking about data being posted, and not data being consumed. Data is consumed off their site via polling APIs, streaming APIs, and a website, all of which are pushing those 400M tweets a day out to plenty of consumers. They may not have as ridiculous a scale as they act like they do. But let's be clear: it is nowhere near as trivial as you make it out to be, either. Armchair quarterbacking is always easy, because you aren't exposed to the complexity that arises when you've spent a few months and years hitting the corner cases of the problem you're commenting on.
- taligent 14y agoDecent amount of traffic ? Sorry but the only thing laughable is that comment.
- InclinedPlane 14y agoWhat does content size matter? The challenge is that every single page of content except for each individual tweet is utterly unique for every user. That defeats the vast majority of straightforward caching implementations. You can't cache fully rendered pages ever because the chance that one random timeline view at a given time will be identical to any other view (even by the same person at a different time) is pretty much as close to zero as possible. Every view is dynamically generated content from up to several hundred or thousand different streams of data and needs to be put in order and have all of the per-user metadata set correctly. Once you start looking into the actual mathematical constraints of the problem of twitter you realize that it's a scaling nightmare. Hundreds of millions of updates per day and tens of thousands of views per second (billions per day). There's only a few people in the world who have the right to look down on stats like that.
- burningout 14y agoAgain, as the parent poster also posted, I think you have never worked on large data. Twitter is like a big mailbox, only that every mail only has 160 bytes. This has been solved 10 years ago.
- achompas 14y agoWait, see my comment below. Twitter received 15B (yes, B) API calls/day last July. How does that compare to your typical email client? I don't want to argue that Twitter is astoundingly hard, but serving ~170K requests/sec can't really be that trivial, even if they're 160 bytes (they're not, since Twitter sends metadata, logs those messages, tracks service metrics, etc. for those messages)
- jasonwatkinspdx 14y agoIf you don't understand that the request distribution matters more than payload size, you aren't even seeing the problems. I encourage you to analyze infrastructure for a twitter style app using inbox duplication. Once you model this against hardware costs you'll learn something about how utterly expensive write amplification is in a hot data set that must be backed by ram due to availability requirements.
- jasonwatkinspdx 14y agoYour assumptions are laughable. Tweets have considerable metadata that pushes them far beyond 160 bytes. Fragmenting this into a secondary object is counterproductive due to the constant factor of 2 requests vs one larger payload. Someone who's been around the block a few times understands that it's difficult to make pronouncements without informed observation. That you are not willing to extend twitter's engineering staff the benefit of the doubt considering your lack of visibility into their measurements speaks loudly.
- InclinedPlane 14y agoAdmittedly I am not well versed in Rails at present. However, the question in regard to twitter isn't so much whether Rails can scale but whether it could scale when twitter needed it too, which was a fair number of years ago. Did Mongrel or Unicorn even exist back in, say, 2007?
- jasonwatkinspdx 14y agoMongrel did.
- InclinedPlane 14y agoWas it mature enough then to have been used in production at twitter's scale?
- alttab 14y agoI had that thought as well. And honestly I'm not entirely sure which version they were running. By now the community has learned a bunch of lessons and we take some of those for granted. Still, the guys over at Twitter could have said "We couldn't scale this technology given our requirements and current knowledge so we went with what we are comfortable with, and thats Java." Instead, they blamed Rails. And now any Rails hater brings up Twitter in a flame war. Even though it was some 4-5 years ago.
- sulam 14y agoTwitter is still the largest Rails site on the web as far as I know, so anyone who says that Twitter hates Rails is either being taken out of context or confused.
- jasonwatkinspdx 14y agoYes. I was on a call with a couple of the twitter devs around this time. They were running on 30 odd instances at joyent. So while they were decent size, they were nothing like they are now. Mongrel was always pretty solid actually. It was originally developed for verisign as I recall, and even in its early versions stuck much more closely to the http specs than other webservers. It caught a lot of flack because of application code blasting out the ruby heap by touching too many objects. Mongrel isn't really to blame for that. With early versions of ActiveRecord it was easy to materialize large result sets without realizing it. A lot of people felt the pain of using associations everywhere without thinking about what would happen when the joins would go north of 10k objects. Not really mongrel's fault.
- SoftwareMaven 14y agoRegardless of the Twitter/Memcache controversy, bad carpenters do blame their tools. Good carpenters succeed in spite of their tools.
- krakensden 14y agoThat response from manjuraj was kind of evil
- floydprice 14y agoHas open source changed? I remember when people didn't fork projects, rebadge them as there own and then promote them over the original. Now don't get me wrong, I think its amazing that Twitter are opening up these enhancements to the community, but it feels like a kick in the teeth to the memcached folks to slap a twitter badge on it, why isn't this a collaboration that benefits the whole community? you know, like open source used to work. I know at 34 I'm a dinosaur in this industry but I do try to keep up with the new way of doing things... This just feels wrong to me.
- zxypoo 14y agoI think open source has changed a bit with the rise of the GitHub era. IMHO, I think @mikeal did a good post regarding this change "Apache considered harmful" http://www.mikealrogers.com/posts/apache-considered-harmful.html http://www.mikealrogers.com/posts/apache-considered-harmful.... Even in the old days, it's easier to fork than work with upstream. I think these days it's just easier share those forks with services like GitHub. It should help spread ideas and improved solutions IMHO so downstream consumers actually benefit. In Twitter's case, they are planning to do what works for them at the moment: "While we initially focused on the challenging goal of making Memcached work extremely well within the Twitter infrastructure, we look forward to sharing our code and ideas with the Memcached community in the long term."
- zxypoo 14y agoJust to clarify things since I helped write the initial blog post about twemcache... in our original post, we had an error regarding the slab calcification problem we mentioned, this problem ONLY applied to our v1.4.4 fork of memcached. After speaking with the upstream maintainers, we learned that recent memcached versions have addressed some of these problems. These are the type of conversations we want to have. At the time we adopted memcached, that's the version we went with and made sure it worked well in our production environment as we scaled as a company. We also open sourced twemproxy [https://github.com/twitter/twemproxy https://github.com/twitter/twemproxy] which is a lightweight proxy for memcached which has worked well for us in combination with twemcache and may work well for others too. We just want to reiterate that twemcache has worked well for our unique environment and any teams evaluating memcached should try all their try all their options, just like any other piece of software you adopt in your stack. One of the reasons of open sourcing our work was to share our ideas with the memcached community to see what worked well for us and help everyone. For example, this is also how we treat our work with our MySQL fork [https://github.com/twitter/mysql https://github.com/twitter/mysql] which we maintain in the open and have signed an OCA with Oracle to help get work pushed upstream so everyone benefits in the long run.
- khangtoh 14y agoApparently another Twitter engineer says it's doing 23M/s "23 million queries per second with zero fucks given" https://twitter.com/timtrueman/status/222793786345013248 https://twitter.com/timtrueman/status/222793786345013248
- forgotusername 14y agotl;dr your memcache branch fragments my consulting clients, you should use my memcache branch instead