4 ms·
Super interesting post. Following blog links, the timeline in https://slack.engineering/all-hands-on-deck-91d6986c3ee https://slack.engineering/all-hands-on-dec
by nik_0_0 6y ago
Super interesting post. Following blog links, the timeline in https://slack.engineering/all-hands-on-deck-91d6986c3ee https://slack.engineering/all-hands-on-deck-91d6986c3ee also offers a look at the play by play.
However, as far as I can read it, they have somewhat different views on the root cause?
"Soon, it became clear we had stale HAProxy configuration files, as a result of linting errors preventing re-rendering of the configuration."
vs.
"The program which synced the host list generated by consul template with the HAProxy server state had a bug. It always attempted to find a slot for new webapp instances before it freed slots taken up by old webapp instances that were no longer running. This program began to fail and exit early because it was unable to find any empty slots, meaning that the running HAProxy instances weren’t getting their state updated. As the day passed and the webapp autoscaling group scaled up and down, the list of backends in the HAProxy state became more and more stale."
Maybe a combination of the two?
- ketzo 6y agoHonestly it’s a bit tough for me to parse, but the way I’m reading it, 1. Stale configs led to an overabundance of web apps, and then 2. Old instances of the web app couldn’t be removed because of the consul-template bug. so, yes, a combination (in sequence) of the two. Hard for me to be sure because I’m by no means knowledgeable on this stuff.
- aeyes 6y agoEven easier way to understand what happened: - slots full - to update slots with a new host you need an empty slot - hosts went away but updating config was impossible -> errors because config referenced non-existing hosts
- folkhack 6y agoAgree but one more: - monitoring was broke so we didn't learn about it until it was too late
- a2tech 6y agoReal root cause was a poorly written home built tool to manage haproxy configs. The tool did not handle the slots being full and crashed. Respawn, rinse and repeat. Haproxy config got stale and when their automated tools removed servers it started with the machines haproxy actually knew about, and then the service died.
- sradman 6y ago> they have somewhat different views on the root cause? It sounds like an issue with naming the failure pattern rather than understanding it. The root cause was equivalent to a memory leak in their custom auto scaling process; machine instances were not being freed (an “instance leak”). The fixed resource limit was self-induced by a hard-coded ratio between the number of proxy servers to web servers. Historically, the fixed ratio never reached a point where the “instance leak” caused failures but on one specific “Terrible, Horrible, No-Good, Very Bad Day” it failed badly.
- user5994461 6y agoThe way they are doing things. HAProxy is configured with a fixed amount of slots. This effectively acts as a maximum limit, so should be enough for the running instances + newer instances coming up anytime due to auto scaling. They have a tool listening to applications starting and shutting down. It's adjusting the configuration live while running to remove shut down instances (free a slot) and put in newer instances (find a free slot and reconfigure). From the explanation on that day, there were more instances than usual due to high load. It seems the tool was looking for a free slot at some point and found none and crashed. I'd say, it's an issue with capacity planning because they didn't plan enough slots for their infra on high load and an issue with the tool because it shouldn't fail silently when out of slots.
- bigiain 6y ago"640K ought to be enough for anybody!" Eventually, seemingly sane assumptions become anachronistic laughing points. (Even if they're apocryphal...)
- amedvednikov 6y agoThis is not a sane assumption, and he never said it, it's a myth.
- bigiain 6y agoI disagree, specifically depending on when that assumption might have been made. My first computer came with 1KB of ram, expandable to 16KB. I have no doubt the designers of that and most of their peers at the time made similar assumptions. My second computer has interesting "bank switching" to circumvent the 16 bit database limitations of it's 8bit cpu that could openly address 64KB, and managed to mostly usefully have 128KB of ram in it. I suspect it's designers would have also happily made an assumption about 640KB being "enough for anyone". (Also, maybe you should look up the definition of "apocryphal? I know he never said it, and strongly alluded to that, and didn't attribute it to Bill for that reason...)