4 ms·
Not sure if this is called for. Heroku has a performance issue and their documentation had a mistake. They accepted everything, apologized and are working to re
by danielpal 14y ago
Not sure if this is called for. Heroku has a performance issue and their documentation had a mistake. They accepted everything, apologized and are working to resolve it. What am I missing here?
- markokocic 14y ago> What am I missing here? A refund to affected customers? The fact that they apologized and accepted the blame does not change the fact that they knowingly degraded performance of their oldest and most loyal customers and forced them to pay much higher costs for years.
- tomlemon 14y agoI appreciate Heroku's apology, but it: 1) Understates the problem – It affects not only Bamboo, but all thin (and other non-concurrent web servers) on Cedar. And since thin is the default on Cedar, the problem affects all apps on Cedar by default 2) Understates how long they knew about it – I notified Heroku about the problem INCLUDING the simulation results 3 days before the blog post came out and yet the apology claims they didn't know about the problem until seeing the blog post.
- ChuckMcM 14y agoThere are two ways to look at this, and depending on your point of view you might be upset.The root of the dispute is how they scale and how that affects latency. According to these write-ups, Heroku scales performance by doing dynamic scheduling on an array of identical servers (called 'dynos'). The documentation talks about a feature named "Intelligent Routing" which only sends work to a dyno which is available to do work. That is a pretty ideal setup because in practice it means that you get linear scaling by adding server instances, and since costs are based on total server instances you get both linear scale increase with linear cost increase. However, there is a very classical problem, first noted by Gene Amdahl, about the cost of figuring out how best to parallelize a stream of requests, vs the rate at which you could satisfy those requests. It became known as "Amdahl's Law"[1]. It limits the practical scalability of a lot of systems. So at some point, Heroku got big enough, that the cost of figuring out which server instance wasn't busy, was taking "too long". (that cost is the (1-P) part in the Amdahl equation) so they decided to reduce the cost of making the choice by replacing a "data driven choice" with a "statistics driven choice". This too is actually a pretty well known way of doing things (Google and Blekko use it to send search queries to a bunch of waiting backends) But unlike the 'idealized' case which has every server instance handling at most one queue, the value becomes a probability that the server instance is either not-busy, or that the current transaction will finish quickly. This works well for systems where the cost of every transaction is nearly identical, you just add servers until the 90th percentile of requests hits your target, but poorly for systems where each request has a variable amount of work it might do. I spent a number of years studying these sorts of systems while solving scalability issues at Network Appliance. A file system built out of a distributed set of nodes providing access to a single file system image needs to know a-priori the cost of each transaction flowing through it in order to optimize scheduling. Similarly RAID subsystems need to know which disk I/Os are going to land in the cache or on the disk, and if they land on the disk will they result in a seek or not. You end up with a directed graph of weighted probabilities being shoved through a channel of fixed bandwidth. Its all amazingly fun until someone says "I have to get data back from the disk in no less than 10mS every time" (databases would say stuff like that) and then you start trading dollars for milliseconds as they say. So Heroku changed their algorithm, didn't tell anyone, and the systemic behavior changed in a very user visible way for large users (in this case Rap Genius). The folks at Rap Genius were pissed off that they made this change without informing them, and they, Rap Genius, looked bad to their customers because of it. Nobody in operations wants to say "Uh, I don't know why your experience with our service is currently sucking." I can see why Rap Genius is mad, and I can see that Heroku might not have fully thought through the ramifications of their algorithm change. [1] http://en.wikipedia.org/wiki/Amdahl%27s_law http://en.wikipedia.org/wiki/Amdahl%27s_law
- chc 14y agoI think the explanation for the change is a bit simpler than that. Someone can correct me if I've misunderstood, but AFAIK it isn't that "Intelligent Routing" was too expensive, but that it depended on simplifying assumptions that stopped being true. Originally, Heroku was only for Ruby, and it depended on the assumption that a server could only handle one request at a time. All the talk of Intelligent Routing seems to date from this time, so it appears that Intelligent Routing just meant they never routed more than one request at a time to a server. But then Heroku wanted to add support for things like Java and Node.js, which can support multiple requests per server. This meant the simplifying assumption of "1 dyno = 1 request" baked into the old routing algorithm was no longer valid, so they had to switch to something else or they'd be crippling pretty much everything but Ruby.
- ChuckMcM 14y agoThat would make sense, if Heroku didn't know which were ruby requests and which weren't. But it seems like they did. If only as a set of VIPs (virtual IPs) landing on the router being tagged as 'for ruby' or 'not ruby' which could pick the appropriate routing algorithm. Understand that keeping millisecond accurate state on 10 machines is doable, on 100 machines its hard, and on a few thousand machines? It really starts to break down. One way I've seen that done is that on ingress a request is wrapped in a message for server X which is taken out of the 'free' pool, and then when server X returns the answer back through the router it gets added back into the free pool. But the next order effect is that list insertion / removal has different sorts of behaviors, if you shift frees into the end of the list and pop them from the front (a round robin approach) you get good distribution but sometimes send things 'far' away when they could be served locally. If you push/pop things from the front you get some really hot servers and some really cold servers. Early on Google played some games which were designed to maximize the use of available network backbone bandwidth (its always oversubscribed from the server to the 'net'). Like any of the more interesting problems it starts off easy and then gets harder and harder.
- ChuckMcM 14y agoIn a later post, Rap Genius quotes Adam Wiggins the CTO of Heroku with this bit: " There are a lot of reasons, but the two big ones are 1) the "intelligent routing" doesn't scale, since it relies on distribution locking which effectively destroys parallelism, and 2) it's incompatible with the evented and realtime apps which are increasingly common on the modern web. "Intelligent routing" sounds good, but in the end it wasn't good for our customers." So, in short they ran into Amdahl's law and changed the way they do things.
- leepowers 14y agoTransparency is definitely called for. Especially for a use case that's getting a lot of attention. Heroku's lack of transparency is what caused this problem in the first place.