18 ms·
Heroku's Ugly Secret: The story of how the cloud-king turned its back on Rails
- bifrost 14y agoI am only going to suggest a small edit -> s/Postgres can’t/Heroku's Postgres can't/ PG can scale up pretty well on a single box, but scaling PG on AWS can be problematic due to the disk io issue, so I suspect they just don't do it. I'd love to be corrected :)
- joevandyk 14y agoIt has to deal with the number of open connections to postgresql, not disk IO. You don't want to have 4,000 connections open at once (unless you were using a connection balancer, but that would only work in transaction mode).
- gleb 14y agoPostgres has a limitation on number of open connections. This is because 1 connection = 1 process. MySQL uses threads, which scales better but has other downsides. The thread-based approach is also possible with Postgres using 3rd party connection pooling apps, e.g. pgbouncer.
- gleb 14y agoBTW, here's a counter-intuitive solution if using pgbouncer is not possible. Simply drop and reestablish connections on every request. In theory this is horrible, since PG connections are so expensive. In practice the cost of establishing a connection is negligible for a Rails app. I do suspect this will make performance "fall-off-the cliff" as you get close to capacity.
- pvh 14y agoEh, I wouldn't recommend going this route. I'd go down the pgbouncer road instead.
- pvh 14y agoData Dep't here. Postgres scales great on Heroku, AWS, and in general. We've got users doing many thousands of query per second, and terabytes of data. Not a problem. The issue with the number of connections is that each connection creates a process on the server. We cap the connections at 500, because at that point you start to see problems with O(n^2) data structures in the Postgres internals that start to make all kinds of mischief. This has been improved over the last few releases, but in general, it's still a good idea to try and keep the number of concurrent connections down. *EDIT: thanks. not a thread. :)
- thehodge 14y agoShame that this seems to have been flagged off the homepage before a reasonable discussion can ensue
- deleted 14y ago[deleted]
- pestaa 14y agoPeople invested too much time/money/energy into Heroku and get defensive when a problem needs discussion? Just guessing.
- thehodge 14y agoHmm seems to be back on the homepage again, not seen that before
- thedufer 14y agoI think this can happen when its been erroneously flagged as having been voting-ringed.
- kmfrk 14y agoThis has a weird habit of happening - or getting noticed at least - when the discussion concerns YC companies.
- omfg 14y agoSomeone from Heroku really needs to weigh in on this.
- VeejayRampay 14y agoGiven how bold the claims are, I'd say it'd be best if Heroku reacted quickly, indeed. Cause right now, it seems that the author is either not understanding how Heroku works or ignoring important specifics/nuances inherent to their "secret-sauce" algorithms. Whether the analysis done in the article is sound or the author is actually pushing an agenda remains to be discussed and the sooner the better. I'm also quite wary of the incentive a 20K monthly bill would give you to try and shake Heroku down for a rebate. By the way, the figure in itself seems very high, but out of context it's impossible for me, the reader, to evaluate if that's actually good money or not. Maybe other solutions (handling everything yourself) would actually be WAY more costly, maybe Heroku actually provides a service that is well-worth the money or maybe the author is right and it's actually swindling on Heroku's part, no way to know.
- timmaah 14y agoThis is not a new revelation. I got them to admit to it 2 years ago. http://tiwatson.com/blog/2011-2-17-heroku-no-longer-using-a-global-request-queue http://tiwatson.com/blog/2011-2-17-heroku-no-longer-using-a-... and specifically: https://groups.google.com/forum/?fromgroups#!msg/heroku/8eOosLC5nrw/Xy2j7GapebIJ https://groups.google.com/forum/?fromgroups#!msg/heroku/8eOo...
- VeejayRampay 14y agoBut then again, those two links don't address the core of the problem: Heroku is used by tons of people around the world. Some of them are paying good money for the service. Given the amount of scrutiny under which they operate, what is the incentive for them to turn an algorithm into a less effective one and still charge the same amount of money in a growing "cloud economy" where companies providing the same kind of service are a dime a dozen (AWS, Linode, Engine Yard, etc)? How does that benefit their business if "calling their BS" is as easy as firing Apache Benchmark, collecting results, drawing a few charts and flat out "prove" that they're lying about the service they provide?? I mean, I doubt Heroku is that stupid, they know how their audience doesn't give them much room for mistakes. So as nice as the story sounds on paper, I'd really like another take on all this, either from other users of Heroku, independent dev ops, researchers, routing algorithms specialists or even Heroku themselves before we all too hastily jump to sensationalist conclusions.
- FireBeyond 14y agoThis should be more prominent. I want to love Heroku, and am sure that I could. But really, throwing in the towel at intelligent routing and replacing it with "random routing" is horrific, if true. It's arguable that the routing mesh and scaling dynamics of Heroku are a large part, if not -the- defining reason for someone to choose Heroku over AWS directly. Is it a "hard" problem? I'm absolutely sure it is. That's one reason customers are throwing money at you to solve it, Heroku.
- timmaah 14y agoEven their own docs were wrong on this for a long time. It bit me in the ass back in 2011 and I got them to clarify and update the documentation just a little. http://tiwatson.com/blog/2011-2-17-heroku-no-longer-using-a-global-request-queue http://tiwatson.com/blog/2011-2-17-heroku-no-longer-using-a-...
- michaelrkn 14y agoThanks for the blog post, by the way. When we were struggling with our own Heroku scaling issues last year (we eventually moved to AWS), I came across it and it was good vindication that somebody else was facing the same issue.
- chc 14y ago> But really, throwing in the towel at intelligent routing and replacing it with "random routing" is horrific, if true. The thing is, their old "intelligent routing" was really just "we will only route one request at a time to a dyno." In other words, what changed is that they now allow dynos to serve multiple requests at a time. When you put it that way, it doesn't sound as horrific, does it?
- arcatek 14y agoThe requests will not be served at the same time, that's the whole point. If a request is routed to a busy dyno, you will have to wait that the previous job finish before being able to start yours.
- simpletouch 14y agoThis is something that I have been struggling with the past long while. Very troublesome when a dyno cycles itself (like they always will at least every 24 hours), as the routing layer continues to send it requests, resulting in router level "Request Timeouts" if it takes too long to restart. Especially difficult to diagnose when the queue and wait time in your logs are 0. What is the point of these in the logs if it never waits or queues?
- michaelrkn 14y agoWe ran into this exact same problem at Impact Dialing. When we hit scale, we optimized the crap out of our app; our New Relic stats looked insanely fast, but Twilio logs told us that we were taking over 15 seconds to respond to many of their callbacks. After spending a few weeks working with Heroku support (and paying for a dedicated support engineer), we moved to raw AWS and our performance problems disappeared. I want to love Heroku, but it doesn't scale for Rails apps.
- WillieBKevin 14y agoWe moved our Twilio app off Heroku for the same reasons. Extensive optimizations and we would still get timeouts on Twilio callbacks. The routing dynamics should be explained better in Heroku's documentation. From an engineering perspective, they're a very important piece of information to understand. We're with https://bluebox.net https://bluebox.net now and are very happy.
- andrewcooke 14y agoi don't think some details of the argument hold. it alleges that you need more dynos to get the same throughput. but that's not true once you have sufficient demand to keep a queue of about sqrt(n) (i think - someone who knows more theory than me can correct me) in size on the dyno (where you have n dynos). because at that point all dynos will be running continuously, and the throughput will be the same with either routing. the average latency will be higher, though (and the spread in latency larger).
- lil_tee 14y agoBut you never want to have a queue on any of your dynos! A queued request means that a user is waiting with no response. If your goal is to have 0 (or less than epsilon) requests queued, it takes far fewer dynos if the requests are routed intelligently If you have 10 dynos and 1000 simultaneous requests, the difference between naive and intelligent might well be reduced, but that's also a scenario in which your end user response times would be horrendously slow and so you'd need more dynos either way
- ratherbefuddled 14y agoI don't think the wording's great. It's not throughput that's important, it's throughput at an acceptable latency. You don't need 50x as many dynos to get the same throughput, you need 50x as many dynos to get the same latency characteristics at that throughput.
- codex_irl 14y agoPersonally - I prefer Linode to Heroku, sure there is more of my time consumed with sys admin, but I like having full control over my platform & setup, rather than having it virtually dictated to me. I'm always open to change but this strategy has served me very well for almost 3 years now.
- kawsper 14y agoI have a Capistrano config that I can slab in my Rails projects. I can then do a: cap deploy:setup cap deploy:cold cap deploy And now my app is running on my server, I then add routing and I am good to go. It is less fancy than Heroku if you want to play with some new technology, you need to install it, and get it configured, and get it to run properly.
- mattj 14y agoSo the issue here is two-fold: - It's very hard to do 'intelligent routing' at scale. - Random routing plays poorly with request times with a really bad tail (median is 50ms, 99th is 3 seconds) The solution here is to figure out why your 99th is 3 seconds. Once you solve that, randomized routing won't hurt you anymore. You hit this exact same problem in a non-preemptive multi-tasking system (like gevent or golang).
- aristus 14y agoI do perf work at Facebook, and over time I've become more and more convinced that the most crucial metric is the width of the latency histogram. Narrowing your latency band --even if it makes the average case worse-- makes so many systems problems better (top of the list: load balancing) it's not even funny.
- jhspaybar 14y agoI can chime in here that I have had similar experiences at another large scale place :). Some requests would take a second or more to complete with the vast majority finishing in under 100MS. A solution was put in place that added about 5 MS to the average request, but also crushed the long tail(it just doesn't even exist anymore) and everything is hugely more stable and responsive.
- genwin 14y agoHow was 5 ms added? Multiple sleep states per request? I imagine the long tail disappears in a similar way that a traffic jam is prevented by lowering the speed limit.
- mixedbit 14y agoI think you misunderstood: they optimized the long running requests and the optimization incurred 5ms performance loss for short requests. It is not that the additional 5ms solved the problem.
- zenazn 14y agoRandomized routing isn't all bad. In fact, if Heroku were to switch from purely random routing to minimum-of-two random routing, they'd perform asymptotically better [1]. [1]: http://www.eecs.harvard.edu/~michaelm/postscripts/mythesis.pdf http://www.eecs.harvard.edu/~michaelm/postscripts/mythesis.p...
- lil_tee 14y agoI'm not familiar with minimum-of-two-random routing, but it does seem like assigning request to dynos in sequence would perform much better than assigning randomly (ie in a scenario with n dyno capacity, request 1 => dyno 1, request 2 => dyno 2, ... request n => dyno n, request n+1 => dyno 1, ..., repeat) That'd be probably significantly better than the case of (request i => dyno picked out of hat) for all i
- icebraining 14y agoThat's essentially round-robin, but it still requires syncing the data of which dyno is next, which is probably what they're trying to avoid.
- pencilcode 14y agoThanks! Very useful.
- jemfinch 14y agoIf Heroku had the data needed to do minimum-of-two random routing, they'd have the data needed to do intelligent routing. The problem is not the algorithm itself: "decrement and reheap" isn't going to be a performance bottleneck. The problem is tracking the number of requests queued on the dyno.
- gojomo 14y agoIf Heroku had the data needed to do minimum-of-two random routing, they'd have the data needed to do intelligent routing. Not strictly true; imagine that they can query the load state of a dyno, but at some non-zero cost. (For example, that it requires contacting the dyno, because the load-balancer itself is distributed and doesn't have a global view.) Then, contacting 2, and picking the better of the 2, remains a possible win compared to contacting more/all. See for example the 'hedged request' strategy, referenced in a sibling thread by nostradaemons from a Jeff Dean Google paper, where 2 redundant requests are issued and the slower-to-respond is discarded (or even actively cancelled, in the 'tiered request' variant).
- squidsoup 14y agoGiven that ElasticBeanstalk has support for rails now, does Heroku still have any advantage over AWS for a new startup?
- kmfrk 14y agoSSL pain can be a major pain to set up. Is the process of setting it up remotely easy compared to Heroku?
- rubyrescue 14y agonot if you use elastic load balancer... it's incredibly easy
- sqs 14y agoSetting up SSL on Elastic Beanstalk was very easy for us. The documentation explained the entire process. It is easier if you get a wildcard SSL cert, so then you can use the same SSL cert for your various deployments under the same domain.
- kmfrk 14y agoThat's very intriguing. Do they support postinstall/post-deploy scripts/hooks like dotCloud to run some framework set-up?
- cameronh90 14y agoI found Elastic Beanstalk had extremely serious latency and performance problems for PHP5, that didn't occur when setting up manual EC2/Load balancer. I intended to investigate it, but never had time.
- sigre 14y agoI spent a lot of time trying to come up with a stable installation of EB for a Rails 3.2 app using RDS and just couldn't get it to a state where I'd ever deploy it as production. Here's where I decided to pull the plug: https://forums.aws.amazon.com/thread.jspa?messageID=410445񤍍 https://forums.aws.amazon.com/thread.jspa?messageID=410445&#... Quite a few people still report a myriad of issues with Rails applications. I really want EB to work well with Rails, but just didn't have confidence in it.
- kawsper 14y agoThis explains why some of my benchmarking tools gave very different and sometimes weird results when figuring out how many dynos our application needed. As this blogpost also states Heroku really need to keep their documentation up to date. I sometimes stumble across something old referring to an old stack, or something contradicting.
- seivan 14y agorails server -p $PORT rake jobs:work Webrick and DJ No procfile? No Unicorn or Puma? No worker process or threads defined
- bignoggins 14y agoRap Genius is employing a classic rap-mogul strategy: start a beef
- parsnips 14y agoNot only that, East Coast vs. West Coast at that...
- hunterhusar 14y agoLOL
- hunterhusar 14y ago-4 for removing an LOL darn
- jussij 14y agoWe've had Battle Rap, Gangsta Rap, you name it Rap. Maybe this is the start of Off the Rails Rap.
- rdl 14y agoWow. I suspect Rap Genius has the dollars now where it's totally feasible for them to go beyond Heroku, but it still might not be the best use of their time. But if they have to do it, they have to do it. OTOH, having a customer have a serious problem like this AND still say "we love your product! We want to remain on your platform", just asking you to fix something, is a pretty ringing endorsement. If you had a marginal product with a problem this severe, people would just silently leave.
- zende 14y agoRap Genius is limited more by time than by money if anything. It would make more sense to throw money at the problem instead of people.
- gleb 14y agoIt doesn't appear that running on Heroku is free for them in terms of time.
- rdl 14y agoThere's also the outage hell. It's been ok for a month or two, but getting killed whenever AWS has a blip in US-East (there's no cross-region redundancy, and minimal resilience with an AZ or Region-wide service has serious problems) isn't great. It probably doesn't hurt RG as much as lower overall performance during normal operations does, though.
- kapilkale 14y agoIt was before. http://success.heroku.com/rapgenius http://success.heroku.com/rapgenius
- lquist 14y agoHow does this compare to EngineYard/AppFog/any other Heroku competitors?
- film42 14y agoIt would be hard to say without a lot of assigned resources (money) and proper testing.
- carbon8 14y agoEngine Yard is more like opinionated configuration management. It allocates and configures EC2 instances that you can log into like normal. The software stack is HAProxy, nginx, unicorn, etc, and customizable through the web interface and/or chef.
- kawsper 14y agoIs there a reason for using both HAproxy and Nginx?
- photomattmills 14y agoHAproxy is a more efficient load balancer for really high scale apps. That said, only about 5% of people would see a difference in load/memory usage. Source: I work for EY.
- felipelalli 14y agocacilda!
- leoh 14y agoUgh. These guys are so cocky.
- htsh 14y agoWhy not hire a devops guy & rack your own hardware? Or get some massive computing units at amazon (just as good but more expensive)? This reminds me of the excellent 5 stages of hosting story shared on here from a while back: http://blog.pinboard.in/2012/01/the_five_stages_of_hosting/ http://blog.pinboard.in/2012/01/the_five_stages_of_hosting/
- c3 14y agowe just switched 1/3rd of our infrastructure off our existing host (engineyard, which uses AWS) onto raw AWS and saved about $2500/month. You can do it too!
- douglasfshearer 14y agoInterested in how you achieved this. Did you change server setup significantly from the default EY stack?
- adanto6840 14y agoI'm curious about this too, especially the last part (similar / identical stack) ?
- Vitaly 14y agoAnd how much more engineering time do you waste on it now?
- bherms 14y agoI've been around a few companies migrating from EY/Heroku -> AWS and the cost effectiveness is always astonishing. In addition you gain full control over your architecture, which is a major plus.
- regularfry 14y agoBecause the whole point is that you shouldn't have to.
- deleted 14y ago
- tim_sw 14y agorandomized routing is not necessarily bad if they look at 2 choices and pick the min. See http://en.wikipedia.org/wiki/2-choice_hashing http://en.wikipedia.org/wiki/2-choice_hashing and http://www.eecs.harvard.edu/~michaelm/postscripts/handbook2001.pdf http://www.eecs.harvard.edu/~michaelm/postscripts/handbook20...
- jemfinch 14y agoThis is a knee-jerk reply. I know, because my knee jerked as well. Think about the problem a little more: if you have the data necessary to pick the min-of-two, then you have the data you need to do intelligent routing.
- jules 14y agoNot necessarily. Heroku claims that a global request queue is hard to scale, and therefore they switched to random load balancing. The comment above shows that a global request queue is not necessary. Lets say that minimum-of-n scales up to 10 dynos. If your application requires 40 dynos, you can have one front load balancer which dispatches the requests to 4 back load balancers, each of which has 10 dynos assigned to them on which they perform min-of-10 routing. This gives you routing that's almost as good as min-of-40 but it scales up nonetheless.
- jacques_chester 14y ago> Heroku claims that a global request queue is hard to scale, and therefore they switched to random load balancing. I wish Heroku would tell us more about what they tried. I can imagine a few cock-a-mimie schemes off the top of my head; it would be good to know whether they thought of those.
- tim_sw 14y agoAssuming by intelligent, you mean minimally loaded, choice of 2 requires less bookkeeping than this. (choice of 2 requires only local (per-node level) info, where as most other intelligent load-balancing requires global info ie. the min, etc.) Taking the case of minimally loaded, you need to keep track of how many active requests each node/replica is serving, as well as globally keeping track of the min. (which past a certain load, will suffer a lot of contention to update) To do choice of 2, all you need is to keep track of active requests per node/replica. Under spiky workloads, there is also a problem with choosing minimally loaded. The counter for numRequests of a node might not update fast enough, so that a bunch of requests will go to that node, quickly saturating its capacity. Choice of 2 doesn't suffer this problem bec of its inherent randomization.
- lquist 14y agoHeroku implements this change in mid-2010, then sells to Salesforce six months later. Hmm...wondering how this impacted revenue numbers as customers had to scale up dynos following the change...
- toast76 14y agoWow. This is explains a lot. We've always been of the opinion that queues were happening on the router, not on the dyno. We consistently see performance problems that, whilst we could tie down to a particular user request (file uploads for example, now moved to S3 direct), we could never figure out why this would result in queuing requests given Heroku's advertised "intelligent routing". We mistakenly thought the occasion slow request couldn't create a queue....although evidence pointed to the contrary. Now that it's apparent that requests are queuing on the dyno (although we have no way to tell from what I can gather) it makes the occasional "slow requests" we have all the more fatal. e.g. data exports, reporting and any other non-paged data request.
- runarb 14y agoIs it so that a dyno can only handle a single user request at a time? Why dos it not use some kind of scheduling system to handle other task while one task is waiting on i/o?
- Cushman 14y agoIt's not exactly so, if you use a server that spawns child processes: http://michaelvanrooijen.com/articles/2011/06/01-more-concurrency-on-a-single-heroku-dyno-with-the-new-celadon-cedar-stack/ http://michaelvanrooijen.com/articles/2011/06/01-more-concur... you can potentially handle 3-4 requests per dyno at a time. That doesn't fix the root problem, though.
- toast76 14y agoInvestigating this approach now. It won't fix the problem, but will certainly reduce the occurrence of blocked dynos. Thx! EDIT: will need to look into our memory perf though, looks like we'll need to do some work to get more than a couple of workers.
- joeya 14y agoI can confirm this. We experimented with Unicorn as a way to get some of the benefits of availability-based routing despite Heroku's random routing. Our medium-sized app (occupying ~230 MB on boot) would quickly exceed Heroku's 512 MB memory limit when forking just 2 unicorn workers, so we had to revert to thin and a greater number of dynos.
- habosa 14y agoWow. Normally when I read "X is screwing Y!!!" posts on Hacker News I generally consider them to be an overreaction or I can't relate. In this case, I think this was a reasonable reaction and I am immediately convinced never to rely on Heroku again. Does anyone have a reasonably easy to follow guide on moving from Heroku to AWS? Let's keep it simple and say I'm just looking to move an app with 2 web Dynos and 1 worker. I realize this is not the type of app that will be hurt by Heroku's new routing scheme but I might as well learn to get out before it's too late.
- michaelrkn 14y agoMy company switched off of Heroku for our high-load app because of these same problems, but I still really like Heroku for apps with smaller loads, or ones in which I'm will to let a very small percentage of requests take a long time.
- mixedbit 14y agoI think this analysis and simulation does not account for one important thing: random routing is stateless and thus easy to be distributed. Routing to the least loaded Dyno needs to be stateful. It is quite easy to implement when you have one centralized router, but for 75 dynos this router would likely become a bottleneck. With many routers, intelligent routing has its own performance cost, the routers need to somehow synchronize state, and the simulation ignores this cost.
- badgar 14y ago> With many routers, intelligent routing has its own performance cost, the routers need to somehow synchronize state, and the simulation ignores this cost. Which is why we pay companies like Heroku to engineer clouds in which to run our applications. Because they're supposed to be better at this than us and spend the time and money building this difficult infrastructure well. That includes a scalable, stateful intelligent routing service.
- habosa 14y agoSomewhat unrelated: Does anyone else think that RapGenius makes a great blogging platform? I'd love a plugin that enabled similar annotations on any blog, even if they're just by the original author and not crowdsourced.
- kmfrk 14y agoIt's been tried before (Apture and a crowd-sourced proof-reading plug-in I can't remember the name of). It needs critical mass to work, but it might very well on a very community-focused platform.
- kmfrk 14y agogooseGrade! I remembered the name. Here's a link: http://www.crunchbase.com/company/goosegrade http://www.crunchbase.com/company/goosegrade.
- kmfrk 14y agoLooks like Cloud 66 couldn't have picked a better day to announce their service: http://news.ycombinator.com/item?id=5213862 http://news.ycombinator.com/item?id=5213862.
- tlrobinson 14y agoQuestion from a non-Ruby-expert: does Thin, which uses Event Machine, help with this at all, or do requests still block on other IO like database calls, etc?
- siong1987 14y agoI don't think that thin will help in this case because Rails is blocking in general. So, you are right because other IOs will still block. You probably need an app that is built on like: https://github.com/raggi/async_sinatra https://github.com/raggi/async_sinatra
- alekseyk 14y agoHow do they know what algorithm Heroku uses for randomization to stimulate the results? The differences in simulations are astonishing, I would not think Heroku's engineers were fine with this approach. 'Let's push this random balancing out.. 1000% increase in resources? Oh well, just update documentation!'
- mononcqc 14y agoI'd guess the problem wouldn't be as bad if each instance could handle more connections/requests than one or two. Allow say, 10 of them, and you will reduce the problem by a lot I believe.
- fatbird 14y agoMaybe this is a dumb question, but wouldn't straightforward Round Robin routing by Heroku restore their "one dyno = one more concurrent request" promise without incurring the scaling liabilities of tracking load across an arbitrarily large number of dynos?
- joevandyk 14y agoNope, requests could still get queued behind a dyno that's busy with a long request.
- fatbird 14y agoSure, but the real issue the article identifies is that, under random routing, they need to keep doubling the number of dynos to halve the odds of bad queueing, which leads to absurd factor of 50 requirements to get back to what they had before. With round robin, the increase should be much more linear.
- mononcqc 14y agoOver many requests, both should average to N requests served to each instance assuming a uniform random distribution. The real issue is being able to figure out instances to avoid when some requests end up being slow. To put it another way, ideal balancing in this case isn't about evenly splitting all requests, but evenly splitting all processing (or waiting) time. If you can guarantee that requests tend to take pretty stable and uniform time, then random or round-robin distribution should give good results. If you can't, some requests will be stuck waiting behind others and their waiting time will accumulate. You'll see worse behaviour when two or more of the bad slow requests get queued one after the other.
- aaron695 14y agoI think you are right. I had to create a quick sim but it does pan out. With round robin it's going to chose the dynos most likely to have the shortest queue given the simple information available. (longest time since it got something) so it's biased to putting stuff on empty queues. Where as random picks randomly, so there's no bias to empty queues, so random should be less efficient.
- zrail 14y agoFor those of you looking to migrate to other, barer hosting solutions like AWS or another VPS provider, I've put together a Capistrano add-on that let's you use Heroku-style buildpacks to deploy with Nginx doing front-end proxy. I use it for half a dozen apps on my VPSs and it works swimmingly well. https://github.com/peterkeen/capistrano-buildpack https://github.com/peterkeen/capistrano-buildpack
- goronbjorn 14y agoAside from the Heroku issue, this is an amazing use of RapGenius for something besides rap lyrics. I didn't have to google anything in the article because of the annotations.
- yeonhoyoon 14y agorap genius is definitely something bigger than rap lyrics. it could potentially be used for a platform for collectively deciphering any text. http://rapgenius.com/Marc-andreessen-why-andreessen-horowitz-is-investing-in-rap-genius-lyrics http://rapgenius.com/Marc-andreessen-why-andreessen-horowitz...
- brusch 14y agoFunny - I thought this was a really interesting article - but I couldn't stand all these annotations. And when I was selecting text (what i do mindlessly when I am reading an article) all the hell broke loose and it tried to load something. Very annoying when I want to concentrate on the technical details. So we'll see once again - everyone's different.
- jules 14y agoWhy not get yourself ONE beefy server? (or two) That should be able to handle your 150 requests per second, simplify your architecture a lot, and buying it would be cheaper than 1 month on Heroku (at $20,000/month).
- sabat 14y agoBecause when that one beefy server goes tits-up, you're out of business. Same with two.
- jules 14y agoWhy? The second is idling until the first fails, then the second takes over. Unless both fail simultaneously of course, for example due to power outage, but then your 40 servers will also fail simultaneously. Not to mention that with just one running server there are a lot fewer failure scenarios.
- ehm_may 14y agoAnd when AWS goes down, your heroku dynos go tits up
- anon640 14y ago"For a Rails app, each dyno is capable of serving one request at a time." Is this a deliberate design choice on Heroku's part, or is this just how Ruby and Rails work? It sounds bizarre that you would need multiple virtual OS instances just to serve multiple requests at the same time. What are the advantages of this over standard server fork()/threaded accept designs?
- kawsper 14y agoIt is how Rails server behaves in itself, but that is also how Heroku tells you to do it. Rails can be served with Unicorn ( http://unicorn.bogomips.org/ http://unicorn.bogomips.org/ ) which is a forking app-server. I do believe there was a trick a while back where you could get Heroku to run a Unicorn process on a dyno to get more requests out of it. The process is described here: http://blog.codeship.io/2012/05/06/Unicorn-on-Heroku.html http://blog.codeship.io/2012/05/06/Unicorn-on-Heroku.html
- anon640 14y agoIsn't this kind of a step backwards? Is Rails really that great that people are willing endure these kinds of limitations just to use it?
- kawsper 14y agoRails server is mostly for development mode. When deployed I think most people use either Unicorn, or throw their applications on JRuby that runs on the JVM with some kind of appserver. JRuby have this advantage of being multithreaded, so you can parallelize within a single process, and don't rely on forking. Stock Ruby with MRI have a GIL, and as far as I know only runs on one core. The limitations of stock Ruby is being worked on, but there is still a long way.
- awj 14y agoRails is commonly run as one or more application servers behind an http server that proxies requests to them. Rails itself doesn't manage threads or forked processes for accepting requests, so the only way it fits into Heroku's dyno model is as an app server per dyno. > What are the advantages of this over standard server fork()/threaded accept designs? It's simple to build and manage in that you don't have to worry about thread safety and can use the already built and tested proxy capabilities of existing web servers to distribute traffic.
- blatyo 14y agoI assumed people were running multiple rails processes on their dynos. http://michaelvanrooijen.com/articles/2011/06/01-more-concurrency-on-a-single-heroku-dyno-with-the-new-celadon-cedar-stack/ http://michaelvanrooijen.com/articles/2011/06/01-more-concur...
- abat 14y agoThe cost of New Relic on Heroku looks really high because each dyno is billed like a full server, which makes it many times more expensive than if you were to manage your own large EC2 servers and just have multiple rails workers. New Relic could be much more appealing if they had a pricing model that was based on usage instead of number of machines.
- googletron 14y agoDoes anyone know how if python applications are affected by this? I know they can handle multiple requests per dyno, I would be interested to know if random routing affects python apps too.
- regularfry 14y agoAs long as they've got a limit on the maximum amount of concurrent requests, they'll be affected. It might well not be as obvious.
- beambot 14y agoFor python / Django you can use a Procfile specifying gunicorn (rather than stock manage.py) with multiple worker threads, eg. web: gunicorn myapp.wsgi -b 0.0.0.0:$PORT -w 5 Then you will have 5 parallel "single-threaded" instances on each dyno rather than just 1. This will partially ameliorate the problem, but probably not 100%. (NOTE: This is speculation since Heroku hasn't weighed in yet)
- jotto 14y agohttp://stackoverflow.com/questions/6370479/heroku-cedar-slower-response-time-than-bamboo http://stackoverflow.com/questions/6370479/heroku-cedar-slow... (from june 2011), the first discovery of cedar stack being slower than the bamboo stack
- pdog 14y agoWhat's the advantage of randomized routing over intelligent routing? Why would this change be made?
- regularfry 14y agoIntelligent routing would presumably need to know and act on a lot of state across their cluster, and if that's got to flow through a single node, you can see how it would present a bottle-neck as the cluster size and requests per second increased. On the other hand, you can do randomised routing without knowing any state at all. You can do it with more than one routing node as well, which makes scaling almost trivial. I presume there are Hard Problems associated with partitioning a Heroku-style cluster for intelligent routing, or that's what they would have done.
- gojomo 14y agoStatelessness. You don't have to remember where you've sent recent requests, which ones are still in process, or how long they've taken.
- deleted 14y ago[deleted]
- Mc_Big_G 14y agoIt's easier/cheaper for them to maintain and you have to pay for more dynos. A lot more.
- reddit_clone 14y agoWin/Win. For them.
- chc 14y agoOK, maybe I'm missing something here, but it seems to me that the OP's real problem is that he's artificially limiting himself to one request per dyno. They now allow a dyno to serve more than one request at a time, and he's presenting that as a bad thing! It seems to me that the answer to Rap Genius' problems is not "rage at Heroku," but rather "gem 'unicorn'".
- michaelrkn 14y agoWe did this, but all it did was buy us a bit of extra time before we ran into the same problem again - a very small percentage (<0.1%) of requests creating a queue that destroyed performance for the rest of them. Also, FWIW, Heroku does not officially support Unicorn, and you have to make sure that you don't run out of memory on your dynos (we tanked our app the first time we tried Unicorn with 4 processes).
- nthj 14y agoI'm inclined to wait until Heroku weighs in to render judgement. Specifically, because their argument depends on this premise: > But elsewhere in their current docs, they make the same old statement loud and clear: > The heroku.com stack only supports single threaded requests. Even if your applicaExplaintion were to fork and support handling multiple requests at once, the routing mesh will never serve more than a single request to a dyno at a time. They pull this from Heroku's documentation on the Bamboo stack [1], but then extrapolate and say it also applies to Heroku's Cedar stack. However, I don't believe this to be true. Recently, I wrote a brief tutorial on implementing Google Apps' openID into your Rails app. The underlying problem with doing so on a free (single-dyno) Heroku app is that while your app makes an authentication request to Google, Google turns around and makes a "oh hey" request to your app. With a single-concurrency system, Google your app times out waiting for Google to get back to you and Google won't get back to you until your app gets back to you so hey deadlock. However, there is a work-around on the Cedar stack: configure the unicorn server to supply 4 or so worker processes for your web server, and the Heroku routing mesh appropriately routes multiple concurrent requests to Unicorn/my app. This immediately fixed my deadlock problem. I have code and more details in a blog post I wrote recently. [2] This seems to be confirmed by Heroku's documentation on dynos [3]: > Multi-threaded or event-driven environments like Java, Unicorn, and Node.js can handle many concurrent requests. Load testing these applications is the only realistic way to determine request throughput. I might be missing something really obvious here, but to summarize: their premise is that Heroku only supports single-threaded requests, which is true on the legacy Bamboo stack but I don't believe to be true on Cedar, which they consider their "canonical" stack and where I have been hosting Rails apps for quite a while. [1] https://devcenter.heroku.com/articles/http-routing-bamboo https://devcenter.heroku.com/articles/http-routing-bamboo [2] http://www.thirdprestige.com/posts/your-website-and-email-accounts-should-be-friends-part-ii http://www.thirdprestige.com/posts/your-website-and-email-ac... [3] https://devcenter.heroku.com/articles/dynos#dynos-and-requests https://devcenter.heroku.com/articles/dynos#dynos-and-reques... [edit: formatting]
- jarcoal 14y agoThis is how I run my apps as well, and they seem to handle more than one request concurrently per dyno, but I'm not smart enough to dispute this post, so I'm just sitting back and watching.
- ratherbefuddled 14y agoIf it's true, I can't see how random routing can be anything but a cynical cash grab. Even a very simple algorithm like round robin would give you a significantly better latency characteristic wouldn't it?
- jasonwatkinspdx 14y agoThe problem is the request arrival rate vs the distribution of service times in your app. New Relic may be giving you an average number you feel happy about, but the 99th percentile numbers are extremely important. If you have a small fraction of requests that take much longer to process, you'll end up with queuing, even with a predictive least loaded balancing policy. This is a very common performance problem in rails apps, because developers often use active record's associations without any sort of limit on row count, not considering that in the future individual users might have 10000 posts/friends/whatever associated object. Fix this and you'll see your end user latency come back in line.
- joeblau 14y agoThanks for the write up. I've been looking for some more reviews on Heorku's platform and this in-depth review definitely illuminates some challenges with the platform.
- benihana 14y agoI really like Rap Genius, but I wish they would tone down the blackness of the background. Reading #CCC text on #000 background makes my eyes bug out.
- jpatokal 14y agoI'm glad you clarified that your first sentence is about their CSS choices. ;)
- marcamillion 14y agoPerspective is a hell of a thing....The way this comment reads, I thought this was going to get racist real quick - but was relieved when I finished reading and did agree with you :)
- sergiotapia 14y agoI think that speaks more on your dormant racism than anyone elses.
- marcamillion 14y agoMy dormant racism how? Vs black people?
- Djehngo 14y agoFor this blog-post where the annotations aren't such a big deal the readability bookmarklet[1] works quite well. [1] http://readability.com/bookmarklets http://readability.com/bookmarklets
- gojomo 14y agoThey want to force the issue with a public spat. Fair enough. But, they also might also be able to self-help quite a bit. RG makes no mention of using more than 1 unicorn worker per dyno. That could help, making a smaller number of dynos behave more like a larger number. I think it was around when Heroku switched to random routing that they also became more officially supportive of dynos handling multiple requests at once. There's still the risk of random pileups behind long-running requests, and as others have noted, it's that long-tail of long-running requests that messes things up. Besides diving into the worst offender requests, perhaps simply segregating those requests to a different Heroku-app would lead to a giant speedup for most users, who rarely do long-running requests. Then, the 90% of requests that never take more than a second would stay in one bank of dynos, never having pathological pile-ups, while the 10% that take 1-6 seconds would go to another bank (by different entry URL hostname). There'd still be awful pile-ups there, but for less-frequent requests, perhaps only used by a subset of users/crawler-bots, who don't mind waiting.
- gojomo 14y agoOn further thought, Heroku users could probably even approximate the benefits from the Mitzenmacher power-of-two-choices insight (mentioned elsewhere in thread), without Heroku's systemic help, by having dynos shed their own excess load. Assume each unicorn can tell how many of its workers are engaged. The 1st thing any worker does – before any other IO/DB/net-intensive work – would be to check if the dyno is 'loaded', defined as all other workers (perhaps just one, for workers=2) on the same dyno already being engaged. If so, the request is redirected to a secondary hostname, getting random assignment to a (usually) different dyno. The result: fewer big pileups unless completely saturated, and performance approaching smart routing but without central state/queueing. There is an overhead cost of the redirects... but that seems to fit the folk wisdom (others have also shared elsewhere in thread) that a hit to average latency is worth it to get rid of the long tails. (Also, perhaps Heroku's routing mesh could intercept a dyno load-shedding response, ameliorating pile-ups without taking the full step back to stateful smart balancing.) Added: On even further thought: perhaps the Heroku routing mesh automatically tries another dyno when one refuses the connection. In such a case, you could set your listening server (g/unicorn or similar) to have a minimal listen-backlog queue, say just 1 (or the number of workers). Then once it's busy, a connect-attempt will fail quickly (rather than queue up), and the balancer will try another random dyno. That's as good as the 1-request-per-dyno-but-intelligent-routing that RapGenius wants... and might be completely within RapGenius's power to implement without any fixes from Heroku.
- zeeg 14y agoIf this is such a problem for you, why are you still on Heroku? It's not a be-all end-all solution. I got started on Heroku for a project, and I also ran into limitations of the platform. I think it can work for some types of projects, but it's really not that expensive to host 15m uniques/month on your own hardware. You can do just about anything on Heroku, but as your organization and company grow it makes sense to do what's right for the product, and not necessarily whats easy anymore. FYI I wrote up several posts about it, though my reasons were different (and my use-case is quite a bit different from a traditional app): * http://justcramer.com/2012/06/02/the-cloud-is-not-for-you/ http://justcramer.com/2012/06/02/the-cloud-is-not-for-you/ * http://justcramer.com/2012/08/30/how-noops-works-for-sentry/ http://justcramer.com/2012/08/30/how-noops-works-for-sentry/
- deleted 14y ago[deleted]
- barmstrong 14y agoWe were very surprised to discover Heroku no longer has a global request queue, and spent a good bit of time debugging performance issues to find this was the culprit. Heroku is a great company, and I imagine there was some technical reason they did it (not an evil plot to make more money). But not having a global request queue (or "intelligent routing") definitely makes their platform less useful. Moving to Unicorn helped a bit in the short term, but is not a complete solution.
- homosaur 14y agoWhile I generally agree with your thoughts, I also wonder what's the reason for continuing to misrepresent their service until you dig 20 layers deep in docs.
- AliEzer 14y agoInteresting article but every time I read something on RapGenius and move my eyes from the screen, I keep seeing white lines, very annoying. White font on black background is bad. Off topic I know, but still.
- rubyrescue 14y agothis is a bit hyperbolic. "Heroku Swindle Factor" just seems rude.
- dblock 14y agoI believe routing is not random, but round robin. I'd like Heroku to confirm. It's still a problem. If you are looking to run Unicorn on Heroku, use the heroku-forward gem (https://github.com/dblock/heroku-forward https://github.com/dblock/heroku-forward). Works well, but application RAM is quickly its own issue, we failed to run that in production as our app takes ~300MB.
- mattbillenstein 14y agoSingle request per server? What year is this? 2003?
- bhauer 14y agoI was thinking the same thing. And 200ms for what exactly? To me the elephant in the room is that 99 server instances are spun up to handle 15M uniques per month.
- lkrubner 14y agoGood lord!!!!! Percentage of the requests served within a certain time (ms) 50% 844 66% 2977 75% 5032 80% 7575 90% 16052 95% 20069 98% 29282 99% 30029 100% 30029 (longest request) Those numbers are amazingly awful. If I ever run ab and see 4 digits I assume I need to optimize my software or server. But 5 digits? Why in the world would a company spend $20,000 a month for service this awful?
- eli 14y agoWell, at X level of concurrency, wouldn't most set ups with load balancers start to spit numbers like that?
- trotsky 14y agono, at x level of concurrency most set ups wont spend 16 seconds or more on 10% of their requests.
- eli 14y agoYou mean because they'll crash before then? Otherwise I don't follow. Surely there's always a limit to how many simultaneous requests can be processed at once.
- trotsky 14y agoHmm, you're right of course. Somewhere our terms got crossed, I wouldn't call (total requests/servers) concurrency, I'd call that request density.
- deleted 14y ago[deleted]
- csense 14y agoHigh cost/risk associated with switching providers, and frog-in-heating-water syndrome.
- 14y ago
- jedahan 14y agoTake a look at deliver if you like heroku push, but want it on a machine you control at a bit lower level: https://github.com/gerhard/deliver https://github.com/gerhard/deliver I got it working on EC2 / Ubuntu real easy, and even added some basic support for SmartOS/illumos for joyent cloud
- eignerchris_ 14y agoThanks for calling this out. As you said, random routing is about as naive as it gets. They need to make upgrades to the routing mesh - expose some internal stats about dyno performance and route accordingly. Even if the stats were rudimentary, anything would be an improvement over random.
- cachvico 14y agoCan't they make the choice of intelligent or random scheduling a per-platform setting? Java, node.js use random. Django + Rails use intelligent.
- rapind 14y agoI'd been using Heroku since forever, but bailed on them for a high traffic app last year (Olympics related) due to poor performance once we hit a certain load (adding dynos made very little difference). We were paying for their (new at the time) critical app support, and I brought up that it appears to be failing at a routing level continuously. And this was with a Sinatra app served by Unicorn (which at the time at least was considered unsupported). We went with a metal cluster setup and everything ran super smooth. I never did figure out what the problem was with Heroku though and this article has been a very illuminating read.
- haddr 14y agoI would like to see why actually Heroku fell back to random routing. It doesn't really make sense. Of course this all routing stuff is really tricky, but on the other hand there is a lot of work done (look at TCP agorithms). When I was studying ZeroMQ routing based stuff for one project, I came across "credit-based flow control" pattern, that could make perfect sense in this kind of situation (Publisher-Subscriber scenario). Why not implementing such thing?
- mleach 14y agoThe balance of a subjective, sensationalist headline with objective statistical simulation was impressive. I'm a huge Heroku fan using Cedar/Java, but can't help but wonder how many optimization options remain for Rails Developers, assuming nothing else changes on Heroku: * Serving static HTML from CDN * Unicorn * Redis caching with multiget requests
- trotsky 14y agoThose charts of "simulated" load balancing strategies don't look at all reasonable at first glance. You certainly don't see such spiky patterns with normal web loads. I think you'd have to have some crazy amount of std. dev in completion time cranked way, way up in your simulation before you saw a bunch of servers stacked at 30 with others at 1. It's not that there is no benefit to better balancing, it's just that I've never seen it have anything close to that impact. It seems like it's only being perceived as a problem here because somebody drank too much of the (old) kool-aid. Some of the other numbers are hard to take at face value as well. 6000ms avg on a specific page? If requests are getting distributed randomly shouldn't all your pages show a similar average time in queue? Sounds more like they're using a hash balancing alg and the static page was hashing on to a hot spot.
- waxjar 14y ago> If requests are getting distributed randomly shouldn't all your pages show a similar average time in queue? A common misconception, called "the law of small numbers". Probability theory tells us this is only true over a large amount of requests, i.e. in the long term (the law of large numebrs). In the short term, results can vary wildly and thus form these kind of queues.
- encoderer 14y agoYou don't think that page was chosen as an example because it should be a very inexpensive page to render? I read that as saying that the performance problem does affect all pages "even on THIS mostly static page." As for your skepticism of the graphs, take a look at the annotated R source they provided. I didn't do a deep dive on it or anything but it looked reasonable to me.
- oellegaard 14y agoThis really sux. I like all their other offerings though - I'm considering running the Cloud Foundry "dyno" part alone and using the heroku services with it.
- benjamincburns 14y agoThis kind of validates an idea I've been flirting with: a Heroku-like service which routes requests via AMQP or similar message broker and actually exposes the routing dynamics to the client apps. From a naive, inexperienced view the idea of having web nodes "pull" requests from a central queue rather than the queue taking uneducated guesses seems to be a no-brainer. I can see this making long-running requests (keep-alive, streaming, etc) a bit more difficult, but not impossible. What am I missing? This seems so glaringly obvious that it must have been done before...
- bdittmer 14y agoI believe Mongrel2 (http://mongrel2.org http://mongrel2.org) is close to what you're talking about. It uses zeromq (http://www.zeromq.org http://www.zeromq.org) to talk to a backend application and the backend application talks back to the web server using the same protocol. Not exactly nodes pulling messages off a queue, but closer to something like that?
- dblock 14y agoThe pull model is very hard to implement because the router behaves like a proxy for a much larger set of dynos (think tens of thousands). When you have 10K clients yielding "i'm available" 10 times a second, you have a nightmare, it's not sustainable. A possible solution for the proxy and the dynos to agree on a protocol where the proxy passes a request to the dyno and the latter can give up with a status code that says "retry with another dyno". This could go on to up to the 30s timeout limit that Heroku has now.
- jacques_chester 14y agoIn Mongrel2, app servers subscribe to a named ZeroMQ queue and, when they're done, they send the response on a different queue. You can actually configure an arbitrary number of different queues if you like, switching on request path and some other stuff I don't just now recall.
- X-Istence 14y agoThat is exactly how an app I develop at $WORK works. Requests come in on a front-end, it gets passed off to a router that has multiple different workers connected. The workers send a request to the router letting the router know that they are ready to start responding to requests. The router hands the worker a request, marks the request as being worked on, and moves on to the next request. It is a basic Least Recently Used queue at that point, and if all workers are busy but worker number 3 which received work last finished first, he gets handed new work instantly. The worker then sends the request back to the router, which sends it back to the appropriate front-end that originally responded to the user. We are using ZeroMQ for our communication. For our use case we can handle around 200 requests a second from a TCP/IP connected client to our router, to a worker and back to a client. That is with 3 backend workers, which are hitting the disk/database. It has worked very well for us.
- MediaSquirrel 14y agoHeroku: The Rap Genius "Success" Story http://success.heroku.com/rapgenius http://success.heroku.com/rapgenius
- evan2 14y agoGreat article. Have you thought about the alternative of building your own auto-scaling architecture with 99.9% uptime? I'd be interested to hear if you plan to move off heroku and, if so, what your plans are.
- wastedbrains 14y agoHeroku for static content is always terrible, I am always surprised at how many people host static sites on Heroku, it is really easy to host of S3 buckets and it is much faster for static pages.
- adminonymous 14y agoI do hope that someone brings this up during Heroku's "Waza" developer conference. It's the perfect opportunity to air it out.
- grandalf 14y agoI'd think that most of the requests being served by rapgenius.com would be highly cacheable (99% are likely just people viewing content that rarely changes). Seems weird that the site would have such a massive load of non-cacheable traffic. Heroku used to offer free and automatic varnish caching, but the cedar stack removed it. Some architectures make it easy to use cloudfront to cache most of the data being served. My guess that refactoring the app to lean on cloudfront would be easier and more cost-effective (and faster) than manually managing custom scaling infrastructure on EC2.
- JuDue 14y agoI'd love to see some better tutorials on how to use AWS Beanstalk to scale Rails apps. There is this one, but it doesn't give me a sense of the scalability or management http://docs.aws.amazon.com/elasticbeanstalk/latest/dg/create_deploy_Ruby_rails.html http://docs.aws.amazon.com/elasticbeanstalk/latest/dg/create... Any recommendations?
- JuDue 14y agoFor example, a small instance is $69/yr for 1.7GB of memory, with additional hourly costs that are quite low This is very economical compared to Heroku, and most startups can survive on that initially if they cache properly. But if there is any level of success, how hard is it to scale compared to the extra cost of Heroku? I'm not convinced it's THAT hard, but would love to see more blog posts about Beanstalk. The AWS doco feels quite mechanical.
- JuDue 14y agoAnd example of what is confusing... $69 to reserve an instance for a year. But that is for "light utilization"?! What does it mean to reserve and instance, but to commit to light usage? And if you are expecting heavy usage, the price goes up to $195. But how can you buy an instance for a year but also commit to your usage level? If it's my instance, why is my utilization anyones business?
- dangrossman 14y agoYou're not committing to a usage level, you're committing to a pricing level. If you reserve a small instance, you get a small instance, no matter what utilization level you choose. It's the same exact resources no matter what you pick. The utilization levels are pricing tiers: Light utilization = lowest upfront cost, highest hourly rate. Medium utilization = medium upfront cost, medium hourly rate. Heavy utilization = highest upfront cost, lowest hourly rate. The names are meant to signify the trade-off you're making. If you run your instance only an hour a day, you will pay the least by choosing "light utilization": the hourly cost is high but you're only going to multiply that by a small number, so the savings in the up-front cost will dominate the total cost. If you run your instance 24 hours a day, then the hourly rate will dominate your total costs, so you'll save money by choosing "heavy utilization" with a higher up-front cost but lower hourly cost. Segmenting the costs makes the pricing table more difficult to read, but it optimizes for everything else: you pay the lowest possible price for guaranteed resources, and Amazon has better knowledge of how much spare capacity it actually needs to handle the reservations.
- JuDue 14y agoI'd love to see some better tutorials on how to use AWS Beanstalk to scale Rails apps. There is this one, but it doesn't give me a sense of the scalability or management http://docs.aws.amazon.com/elasticbeanstalk/latest/dg/create_deploy_Ruby_rails.html http://docs.aws.amazon.com/elasticbeanstalk/latest/dg/create... Any recommendations for good tutorials?
- zeeg 14y agoHere's a very simple gevent hello world app. This is run from inside AWS on an m1.large: https://gist.github.com/dcramer/4950101 https://gist.github.com/dcramer/4950101 For the 50 dyno test, this was the second run, making the assumption that the dynos had to warm up before they could effectively service requests. You'll see that with 49 more dynos, we only managed to get around 400 more requests/second on an app that isnt even close to real world. (By no means is this test scientific, but I think it's telling)
- clouddevops 14y agoPerhaps easy deployments are not worth the performance and blackbox trade-off. An alternative approach is a cloud infrastructure provider with baremetal and virtual servers on L2 broadcast domain, and one that provides a good API and orchestration framework so that you can easily automate your deployments. Here are some things we at NephoScale suggest you consider when choosing an infrastructure provider: http://www.slideshare.net/nephoscale/choosing-the-right-infrastructure-provider http://www.slideshare.net/nephoscale/choosing-the-right-infr...
- deleted 14y ago[deleted]
- juanbyrge 14y agoLOL, heroku is not designed for real apps. It's designed for side projects and consulting projects that don't go anywhere. Anytime you get traffic, move off ASAP!
- hpguy 14y agoCan anyone explain why this random routing is supposedly good for Node.JS and Java? I mean the net effect is busy dynos might serve more requests while idle ones remain idle and that is certainly not good for Node.JS or anything. What am I missing?
- ajsharp 14y agoMy hunch is that Heroku isn't doing this to bleed customers dry. I know more than a few really, really great people who work there, and I don't think they'd stand for that type of corporate bullshittery. If this were the case, I think we'd have heard about it by now. My best guess is that they hit a scaling problem with doing smart load balancing. Smart load balancing, conceptually, requires persistent TCP connections to backend servers. There's some upper limit per LB instance or machine at which maintaining those connections causes serious performance degredations. Maybe that overhead became too great at a certain point, and the solution was to move to a simpler random round-robin load balancing scheme. I'd love to hear a Heroku employee weigh in on this.
- Dylan16807 14y agoThe thing that baffles me is that you could do high-level random balancing onto smaller clusters that do smart balancing. This would solve most of the problem of overloaded servers. An entire cluster would have to clog up with slow requests before there was any performance impact. So why don't they do this?
- alttab 14y agoThey should hire you. I mean that as a compliment.
- ajsharp 14y agoMy off the cuff answer to your question is, because it's probably not quite that simple ;)
- Dylan16807 14y agoThat's why I asked ;)
- nacho2sweet 14y agoWhy is everyone like against rapgenius.com for "forcing the issue with a public spat". They are the customer not getting a service they are paying for. I would be fucking pissed too. Heroku isn't being the darling the service they advertised. They tried to work on it with Heroku. This is useful information to most of you. Are most of you against Yelp?
- gojomo 14y agoThe only person using the words 'force the issue with a public spat' is me, and I judged that as 'fair enough'. I'm not against RapGenius and I'm glad the issue is being discussed. But we haven't seen Heroku's comments, and while some parts of RapGenius's complaint are compelling, I'm not sure their apparent conclusions - that 'intelligent routing' is needed, and its lack is screwing Heroku customers — are right. I strongly suspect some small tweaks, ideally with Heroku's help, but perhaps even without it, can fix most of RapGenius's concerns. Perhaps there was a communication or support failure, which led to the public condemnation, or maybe that's just RG's style, to heighten the drama. (That's an observation without judgement; quarrels can benefit both parties in some attention- and personality-driven contexts.)
- michaelfairley 14y agoThere's another fun issue that falls out of this: any requests sitting in the dyno queue when the app restarts get dropped with a 5xx error. https://github.com/michaelfairley/unicorn-heroku/issues/1#issuecomment-8601906 https://github.com/michaelfairley/unicorn-heroku/issues/1#is...
- joshwa 14y agoFundamentally we're talking about a load balancer. Even the most basic load-balancers can use a least-connections algorithm. Even a round-robin algorithm would be better since that would give each dyno (number-of-dynos * msec per request) to finish a long-running request. Random routing is a viable option where the number of concurrent requests a node can handle is large or unknown, but when the limit is known and in the single-digits, random routing is a recipe for disaster.
- bryanwbh 14y agoThanks for the write-up on this as I am currently reviewing heroku as an option for a proper PAAS for my app.
- nsrivast 14y agoOP is a friend of mine, and when I first heard of his problem I wondered if there might be an analytical solution to quantify the difference between intelligent vs naive routing. I took this problem as an opportunity to teach myself a bit of Queueing Theory[1], which is a fascinating topic! I'm still very much a beginner, so bear with me and I'd love to get any feedback or suggestions for further study. For this example, let's assume our queueing environment is a grocery store checkout line: our customers enter, line up in order, and are checked out by one or more registers. The basic way to think about these problems is to classify them across three parameters: - arrival time: do customers enter the line in a way that is Deterministic (events happen over fixed intervals), randoM (events are distributed exponentially and described by Poisson process), or General (events fall from an arbitrary probability distribution)? - checkout time: same question for customers getting checked out, is that process D or M or G? - N = # of registers So the simplest example would be D/D/1, where - for example - every 3 seconds a customer enters the line and every 1.5 seconds a customer is checked out by a single register. Not very exciting. At a higher level of complexity, M/M/1, we have a single register where customers arrive at rate _L and are checked out at rate _U (in units of # per time interval), where both _L and _U obey Poisson distributions. (You can also model this as an infinite Markov chain where your current node is the # of people in the queue, you transition to a higher node with rate _L and to a lower node with rate _U.) For this system, a customer's average total time spent in the queue is 1/(_U - _L) - 1/_U. The intelligent routing system routes each customer to the next available checkout counter; equivalently, each checkout counter grabs the first person in line as soon as it frees up. So we have a system of type M/G/R, where our checkout time is Generally distributed and we have R>1 servers. Unfortunately, this type of problem is analytically intractable, as of now. There are approximations for waiting times, but they depend on all sorts of thorny higher moments of the general distribution of checkout times. But if instead we assume the checkout times are randomly distributed, we have a M/M/R system. In this system, the total time spent in queue per customer is C(R, _L/_U)/(R _U - _L), where C(a,b) is an involved function called the Erlang C formula [2]. How can we use our framework to analyze the naive routing system? I think the naive system is equivalent to an M/M/1 case with arrival rate _L_dumb = _L/R. The insight here is that in a system where customers are instantaneously and randomly assigned to one of R registers, each register should have the same queue characteristics and wait times as the system as a whole. And each register has an arrival rate of 1/R times the global arrival rate. So our average queue time per customer in the dumb routing system is 1/(_U - _L/R) - 1/_U. In OP's example, we have on average 9000 customers arriving per minute, or _L = 150 customers/second. Our mean checkout time is 306ms, or _U ~= 3. Evaluating for different R values gives the following queue times (in ms): # Registers 51 60 75 100 150 200 500 1000 2000 4000 dumb routing 16,667 1,667 667 333 167 111 37 18 9 4 smart routing 333 33 13 7 3 2 1 0 0 0 which are reasonably close to the simulated values. In fact, we would expect the dumb router to be comparatively even worse for the longer-tailed Weibull distribution they use to model request times, because you make bad outcomes (e.g. where two consecutive requests at 99% request times are routed to the same register) even more costly. This observation seems to agree with some of the comments as well [3]. [1] http://en.wikipedia.org/wiki/Queueing_theory http://en.wikipedia.org/wiki/Queueing_theory [2] http://en.wikipedia.org/wiki/Erlang%27s_C_formula#Erlang_C_formula http://en.wikipedia.org/wiki/Erlang%27s_C_formula#Erlang_C_f... [3] http://news.ycombinator.com/item?id=5216385 http://news.ycombinator.com/item?id=5216385
- teich 14y agoThis is Oren Teich, I run Heroku. I've read through the OP, and all of the comments here. Our job at Heroku is to make you successful and we want every single customer to feel that Heroku is transparent and responsive. Getting to the bottom of this situation and giving you a clear understanding of what we’re going to do to make it right is our top priority. I am committing to the community to provide more information as soon as possible, including a blog post on http://blog.heroku.com http://blog.heroku.com.
- GhotiFish 14y agoI'm looking forward to hearing why Heroku is using such a strange load balancing strategy.
- character0 14y agoWhile I think it is appropriate for Heroku to respond to this thread (and other important social media outlets covering this), linking to a blog without any messaging concerning your efforts might not be the greatest move... This may not be a sink or swim moment for Heroku, but tight management of your PR is key to mitigating damage. Best of luck, Heroku is a helpful product and I want to see you guys bounce back from the ropes on this one.
- csense 14y agoTelling people where to look for a reply when they have one is a great idea, IMHO.
- tjbiddle 14y agoLooking forward to your blog post. Hoping things get cleared up!
- avodonosov 14y agoI hope the solution will not break the possibility for multithreaded apps to receive several requests
- 14y ago
- cwalcott 14y agoManaging load with thin workers isn't very hard...haproxy [http://haproxy.1wt.eu/ http://haproxy.1wt.eu/] makes it pretty easy to setup rather complex load distributions (certainly more complex than random!).
- Sami_Lehtinen 14y agoInteresting. No I didn't read it because the page crashes my mobile browser everytime reliably.
- aren55555 14y agoThis was a great read.
- bad_user 14y agoI noticed problems with Heroku's router too. However, contrary to the author, I'm serving 25,000 real requests per second with only 8 dynos. The app is written in Scala and runs on top of the JVM. And I was dissatisfied that 8 dynos seem like too much for an app that can serve over 10K requests per sec on my localhost.
- jaytaylor 14y agoThis sounds interesting, but kind of suspicious. What app is serving that kind of volume continuously other than fb or goog? You running zynga on Heroku or something?
- bad_user 14y agoIt's an integration with an OpenRTB bidding marketplace that's sending our way that traffic. And 25K is not the whole story. In a lot of ways it's similar to high frequency trading. Not only do you need to decide in real time if you want to respond with a bid or not, but the total response time should be under 200ms, preferably under 100ms, otherwise they start bitching about latency and they could drop you off from the exchange. And the funny thing is 25K is actually nothing, compared to the bigger marketplace that we are pursuing and that will probably send us 800K per second at peak.
- jmount 14y agoTurns out to be a great queueing problem. Please check out my analysis of much simplified version of the random routing algorithm that fails with near certainty: http://www.win-vector.com/blog/2013/02/randomized-algorithms-can-have-bad-deterministic-consequences/ http://www.win-vector.com/blog/2013/02/randomized-algorithms...
- gtirloni 14y agoVMs running full frameworks that are single-threaded. Why does that feel like wasting resources or bloating the architecture?
- lucian303 14y agoMost issues with cloud providers are not ones of technology but ones of trust. Trust always take precedence.
- aneth4 14y agoWould love to see some more perspectives on this. We also spend a lot of resources on heroku. I'm not sure if this change by heroku is worse than the intermediary popups on all the links on this blog.
- dschiptsov 14y agoWhy on Earth any sane engineer would think that adding layers of "virtualized" crap in front of your application will be of any benefit?) The only advantage of virtualization is on the developing stage and it is an ability to add quickly more slow and crappy resources you not own.) Production is an entirely different realm, and the less layers of crap is in between of your TCP request and DB storage - the better. As for load balancing - it is Cisco level problem.) Last question: why each web site must be represented as a hierarchy of some objects, instead of thinking in terms of what it is - a list of static files and some cached content generation on demand?)
- dangrahn 14y agoI was in contact with Heroku support a couple of weeks ago since we experienced some timeout on our production app. Got a detailed explanation how the routing on heroku works by a Heroku engineer, and thought I could share: "I am a bit confused by what you mean by an "available" dyno. Requests get queued at the application level, rather than at the router level. Basically, as soon as a request comes in, it gets fired off randomly to any one of your web dynos. Say your request that takes 2 seconds to be handled by the dyno was dispatched to a dyno that was running a long running request. Eventually, after 29 seconds, it completed serving the response, and started working on the new, faster 2 second request. Now, at this point it had already been waiting in the queue for 29 seconds, so after 1 second, it'll get dropped, and after another 1 second, the dyno will be done processing it, but the router is no longer waiting for the response as it has already returned an H12. That's how a fast request can be dropped. Now, the one long 29 second request could also be a series of not-that-long-but-still-long requests. Say you had 8 requests dispatched to that dyno at the same time, and they all took 4 seconds to process. The last one would have been waiting for 28 second, and so would be dropped before completion and result in an H12."
- EGreg 14y agoAnd this kind of thing is why I prefer to have our own VPS. Linode is great, but we are slowly switching over to AWS and automating all the scaling up/down.
- jhuckestein 14y agoWatch out, this affects small rails applications with few dynos as well. If you hit the wall with one dyno and add another one, you won't get twice the throughput even though you pay twice the price. I've always had suspicions about this on some smaller apps but never really looked into it. You can configure New Relic to measure round-trip response times on the client side. At peak loads those would be unreasonably high. Much higher than can be explained by huge latencies even.
- Uchikoma 14y ago~$20,000 sounds like a lot of money, they would need at least $100M a year in revenue to justify this number. This will be a major challenge if they want to grow profitable after the $15M VC money runs out. I'd assume they'd get the same for $5000 in rented servers which would free up enough money - outside of the valley - to have an DevOps and another developer.
- Terretta 14y agoNot following. Why do you say they need $100 million a year in revenue to justify $240K (00.25% of revenue) in hosting expenses? For an online biz, 10% - 50% isn't uncommon for profitable businesses. Many "virtual" companies (online only) do fine at 80%.
- Uchikoma 14y agoOtherwise you do not have enough traffic to justify $240k.
- zen_boy 14y agoHow do you need $100M a year in revenue to cover $20k/mo or $240k/year?
- filvdg 14y agoYou can Model Queues and calculate the service level when you do inteligent routing The Erlang C formula expresses the probability that an arriving customer will need to queue (as opposed to immediately being served). http://owenduffy.net/traffic/erlangc.htm http://owenduffy.net/traffic/erlangc.htm
- krutulis 14y agoI can't help but wonder if this kind of surreptitious change to the platform might in any way be connected to Byron Sebastian's sudden resignation last September from Salesforce. Is that nutty of me? http://gigaom.com/2012/09/05/heroku-loses-a-star-as-ceo-and-salesforce-evp-sebastian-resigns/ http://gigaom.com/2012/09/05/heroku-loses-a-star-as-ceo-and-...
- jonnycat 14y agoThis might be the case "out of the box", but it's very simple to go multithreaded on the Cedar stack and avoid this issue (provided that your app is threadsafe, of course). You can do this pretty easily with a Procfile and thin: bundle exec thin -p $PORT -e $RACK_ENV --threaded start And then config.threadsafe! in the app Regarding Rails app threadsafety, there are some gotchas around class-level configuration and certain gems, but by and large these issues are easily manageable if you watch out for them during app development.
- zen_boy 14y agoAre the some resources on the general topic what it means to build multi-threaded Rails app instead of a traditional one?
- knodi 14y agoThis kind of routing wouldn't be a problem if they didn't charge $35 a dyno. It's such a high cost for a dyno.
- Giszmo 14y agoI searched for variance and apparently nobody mentioned this before: They talk of Mean request time: 306ms Median request time: 46ms Which indicates a very high variance, so don't take for granted that an x50 increase of performance would result from intelligent routing. The problem is that the fast tasks suffer from being queued after the slow tasks, so each such fast task takes an extra latency. If the variance is lower, the random routing will be favorable at some point as the delay of getting the task from the router queue to the dyno is not zero neither. In the case of no variance, "intelligent routing" would always add that delay as soon as all dynos are at their limit. Before that, the router would simply keep a list of idle dynos and send work there without delay. Sure if you never hit 100% load, intelligent routing is cheap and comes at no delay. Imagine 40ms jobs getting all dynos to 100% load. Now the dynos would be idle for the duration of the ping that it takes to report being idle. let that be 4ms. That is 10% less throughput than with items queuing up on the dyno. The router being the bottleneck would therefore justify to make it stateless and give the dynos a chance to use these last 10% of processing power as well, ultimately increasing the throughput by 10%. Sure, a serious project would not run its servers at 120% load hoping to eventually get back to 100% within time, so all this being said I would always favor intelligent routing to get responsive servers, add dynos in rush hours and only opt for dyno-queuing for stuff that may come with a delay (scientific number crunching, …)
- philipDS 14y agoOff-topic: RapGenius should really open source their "Explain tooltips" with the inline explanation window. Awesome :)
- vineet 14y agoIt seems that the delay is in the variance in the length of the different jobs. Having slow jobs is generally not a good idea, and I can imagine that they are happening for uncommon tasks. When you are running a 100+ servers it seems like a simple answer would be to think about these uncommon tasks differently. Options would be for prioritizing them differently, showing different UI indicators, and also wanting them happening on a separate set of machines. Doing these would mean that an intelligent routing mechanism would not have as much use. Am I wrong here? I do believe that Heroku should document such problems of theirs more clearly, so that we know what challenges that we are facing as we develop applications, but in this particular case, it seems that they do have the right plumbing, and that they just need to be used differently.
- stevewilhelm 14y agoHeroku Support Request #76070 To whom it may concern, We are long time users of Heroku and are big fans of the service. Heroku allows us to focus on application development. We recently read an article on HN entitled 'Heroku's Ugly Secret' http://s831.us/11IIoMF http://s831.us/11IIoMF We have noticed similar behavior, namely increasing dynos does not provide performance increases we would expect. We continue to see wildly different performance responses across different requests that New Relic metrics and internal instrumentation can not explain. We would like the following: 1. A response from Heroku regarding the analysis done in the article, and 2. Heroku-supplied persistant logs that include information how long requests are queued for processing by the dynos Thanks in advance for any insight you can provide into this situation and keep up the good work.
- stevewilhelm 14y agoHeroku's response: Hi Steve, I've been reading through all the concerns from customers, and I want every single customer to feel that Heroku is transparent and responsive. Our job at Heroku is to make you successful. Getting to the bottom of this situation and giving you a clear and transparent understanding of what we’re going to do to make it right is our top priority. I am committing to the community to provide more information as soon as possible, including a blog post on http://blog.heroku.com http://blog.heroku.com. Oren Teich Heroku GM
- pointful 14y agoJust adding a top-level post to point out something buried in one of the threads here that is an important point on what is happening here: The "queue at the dyno level" is coming from the Rails stack -- it's not something that Heroku is doing to/for the dynos. Thin and Unicorn (and others, I imagine) will queue requests as socket connections on their listener. Both default to 1024 backlog requests. If you lower that number, Heroku will (according to the implications in the documentation on H21 errors) try multiple other dynos first before giving up. See https://devcenter.heroku.com/articles/error-codes#h21-backend-connection-refused https://devcenter.heroku.com/articles/error-codes#h21-backen... For a single-threaded process to be willing to backlog a thousand requests is problematic when combined with random load balancing. Dropping this number down significantly will lead to more sane load-balancing behavior by the overall stack, as long as there are other dynos available to take up the slack. Also, the time the request spends on the dyno, including the time in the dyno's own backlog, is available in the heroku router log. It's the "service" time that you'll see as something like "... wait=0ms connect=1ms service=383ms ...". Definitely wish New Relic was graphing that somewhere...
- izietto 14y agoFrom the Heroku docs: [...] Request distribution The routing mesh uses a random selection algorithm for HTTP request load balancing across web processes. [...] If the algorithm is random, the load balancing simply doesn't happen, am I wrong? https://devcenter.heroku.com/articles/http-routing#request-distribution https://devcenter.heroku.com/articles/http-routing#request-d...
- frankc 14y agoI know nothing about Heroku's architecture than what I just read in this post, but couldn't you alleviate this problem greatly by having the dyno's implement work stealing? Obviously the they would have to know about each other then, but perhaps that is easier to do that global intelligent routing.
- lil_tee 14y agoWe have updated our post to incorporate some popular suggestions and reactions into our simulations: http://rapgenius.com/1504221 http://rapgenius.com/1504221
- pm90 14y agoDid anyone else find the headline a bit confusing? I got the feeling that they had abandoned RoR for another framework and almost skipped the article itself
- justinhj 14y agoSeems like message here is that if you use an off the shelf solution you need to work around its limitations. In this case random load balancing may sound dumb but it's actually quite a reasonable way to spread load. The customers real problem is the single threaded server bottleneck compounded by the sporadic slow requests. Seems like they have outgrown Heroku and a more custom solution is required. Either that or rebuild the server in whole or in part with a more concurrent one.
- damian2000 14y agoA bit of context: http://success.heroku.com/rapgenius http://success.heroku.com/rapgenius
- craigkerstiens 14y agoThe initial response from GM of Heroku - https://blog.heroku.com/archives/2013/2/15/bamboo_routing_performance/ https://blog.heroku.com/archives/2013/2/15/bamboo_routing_pe...