7 ms·
#4 I work at Facebook. I worked at Twitter. I worked at CloudFlare. The answer is nothing other than #4. #1 has the right premise but the wrong conclusion. So
by bdd 7y ago
#4
I work at Facebook. I worked at Twitter. I worked at CloudFlare. The answer is nothing other than #4.
#1 has the right premise but the wrong conclusion. Software complexity will continue escalating until it drops by either commoditization or redefining problems. Companies at the scale of FAANG(+T) continually accumulate tech debt in pockets and they eventually become the biggest threats to availability. Not the new shiny things. The sinusoidal pattern of exposure will continue.
- marenkay 7y agoOnly one thing to add: Tech debt is accrued in amounts where every VC fund would get wet pants if tech debt was worth dollars paid out.
- kwizzt 7y agoHow does the fact you worked at those companies relate to #4? Edit: I misread the parent and my question doesn't make a lot of sense. Please ignore it :)
- bdd 7y ago> How does the fact you worked at those companies relate to #4? For Facebook I worked on the incident, previous Wednesday. 9.5 hours of pain... And for my past employers, I still have friends there texting the root causes with facepalm emojis.
- liberte82 7y agoDo tell
- deleted 7y ago[deleted]
- deleted 7y ago[deleted]
- deleted 7y ago[deleted]
- captn3m0 7y agoCan you clarify what redefining problems would mean (with an eg)?
- GuiA 7y agoThink of computer vision tasks. Until modern deep learning approaches came around, it was built on brittle, explicitly defined pipelines that could break entirely if something minor about the input data changed. Then the great deep learning wave of 201X happened, replacing dozens/hundreds of carefully defined steps with a more flexible, generalizable approach. The new approach still has limitations and failure cases, but it operates at a scale and efficiency the previous approaches could not even dream of.
- MegaButts 7y agoThat's not redefining the problem, so much as applying a new technology to solve the same problem. Usually using the flashy new technology decreases reliability due to immature tooling, lack of testing, and just general lack of knowledge of the new approach. Also deep learning, while incredibly powerful and useful, is not the magic cure-all to all of computer vision's problems and I have personally seen upper management's misguided belief in this ruin a company (by which I mean they can no longer retain senior staff, they have never once hit a deadline, every single one of their metrics is not where they want it to be, and a bunch of other stuff I can't say without breaking anonymity).
- loblollyboy 7y agoWow, such certainty. "I worked at these companies that went down, so I knew and still know everything, I can even rule out the possibility of other organizations which I did not work for screwing with them, because I know everything about those too. And yes, these FANGs are complex organizations, but not so complex that a former employee like myself wouldn't know what the cause of an outage is or at least isn't (Hell, I'd fix it, but I have to finish explaining why complexity is not the reason why they are crashing, and the series of ten dollar words that constitute my explanation aren't exactly the quickest to type."
- jkaplowitz 7y agoFormer employees and current employees talk via unofficial online and offline backchannels at many companies.
- loblollyboy 7y agoOk, so maybe I overreacted
- bdd 7y agogeez, tough crowd. do you wanna ten dollar hug?
- loblollyboy 7y agoI was just polishing my bit. Not in a bad mood today so much as a bored mood. You seem like you know what you are talking about (yes, I was bored enough to stalk you, too)
- fossuser 7y agoYep, this also matches what I've heard through the grapevine. Pushing bad regex to production, chaos monkey code causing cascading network failure, etc. They're just different accidents for different reasons. Maybe it's summer and people are taking more vacation?
- degenerate 7y agoI actually like the summer vacation hypothesis. Makes the most sense to me - backup devs handling some things they are not used to.
- Avamander 7y agoThese outages mean that software only gets more ~fool~ summer employee proof.
- Balgair 7y agoSo, a reverse Eternal-September? It'll get better once everyone is back from August vacations?
- uber-employee 7y agoNo, because it’ll only get better until next summer.
- kenhwang 7y agoI'm more partial to the summer interns hypothesis.
- bobthepanda 7y agoRule one of having interns and retaining your sanity is that interns get their own branch to muck around in.
- vorticalbox 7y ago
- hexrcs 7y agoLife in tech is like a Quentin Tarantino movie.
- _jal 7y ago...except everyone is sitting at desks typing, there's no blood or surf rock or chases or self-indulgent soliloquies, and the cursing is much less creative?
- jessaustin 7y agoMaybe you're doing it wrong?
- robohoe 7y agocursing is much less creative? I beg to differ.
- aaroninsf 7y agoI will observe, without asserting that it is actually the case, that successful executions of #3 should be indistinguishable from #4. (And this is maybe a consequence of #1).
- mastratton3 7y agolol yes, whats the quote on "Don't assume bad intention when incompetence is to blame"? After seeing how people write code in the real world, I'm actually surprised there aren't more outages.
- newsbinator 7y agoHanlon's Razor: https://en.wikipedia.org/wiki/Hanlon%27s_razor https://en.wikipedia.org/wiki/Hanlon%27s_razor "Never attribute to malice that which is adequately explained by stupidity."
- euske 7y agoI always thought that this cause should also include "greed". But then, greed is kinda one step closer to malice, and I'm not sure if there's a line.
- rossdavidh 7y agoAh, but that's a lot of big corps being more stupid in the last month than last year? If it's two or three more, that's normal variation. We're now at something more like 7 or 8 more. The industry didn't get that much stupider in the last year.
- jethro_tell 7y agoWell we have an entire profession of SRE/Systems Eng roles out there that are mostly based on limiting impact for bad code. Some of the places I've worked with the worst code/stacks had the best safety nets. I spent a while shaking my head wondering how this shit ran without an outage for so long until I realized that there was a lot of code and process involved in keeping the dumpster fire in the dumpster.
- devin 7y agoWhich do you prefer? Some of the best stacks and code I’ve worked in wound up with stability issues that were a long series of changes that weren’t simple to rework. By contrast, I’ve worked in messy code, complex stacks, that gave great feedback. In the end, the answer is I want both, but I actually sort of prefer “messy” with well thought out safety nets to beautiful code and elegant design with none.
- gcbw2 7y agosince all of them happen in high profile business hours, i'd guess either #1 or #5. For #4 to be the actual cause, outages out of business hours would be more prevalent and longer.
- icebraining 7y agoOf course it went down during business hours, that's when people are deploying stuff. It's known that services are more stable during the weekends too.
- peterburkimsher 7y agoThe Archive.org outage of 26th of June was outside PST business hours. https://twitter.com/internetarchive/status/1143604539695616000 https://twitter.com/internetarchive/status/11436045396956160... https://twitter.com/internetarchive/status/1143378990826004480 https://twitter.com/internetarchive/status/11433789908260044...
- Diederich 7y agoI've also worked at a couple of the companies involved. This is the correct analysis on every level.
- idlewords 7y agoFAANG(+T)(-N)(+M)
- 18pfsmt 7y agoI think we 'bumped heads' at Middlebury in '94, and I think you are in store for an "ideological reckoning" w/in 3 years. Pinboard is a great product, so thanks for that. I am surpised you don't have your own Mastodon instance (or do you?).
- wybiral 7y agoI've still never seen this much downtime on these systems so it's weird to happen all at once. It's possible that they're related without requiring any conspiracy theories or anything. Maybe these companies are just getting too big or too sloppy to maintain the same standard of uptime (compared to the past few years)? Or maybe there's some underlying issue that they're all rushing to fix which justifies the breaking prod changes within the same timeframe. But it was weird when a it happened to two or three of them. Now we're going on something like 5 massive failures from some of the biggest services online within a little over a week...
- iamtheworstdev 7y agoFaangt = Facebook amazon Apple Netflix Google tesla?
- gsich 7y agoGmafia
- arrty88 7y agoAdd slack to the list Edit: and stripe
- foobarbecue 7y agoTwitter, not tesla
- gjs278 7y agoit sounds like you’re the common factor between the outages. where else have you worked, maybe we can predict the next failure
- cmroanirgo 7y agoYes, but it always seems to come down to a very small change with far reaching consequences. For this ongoing twitter outage, it's due to an "internal configuration change"... and yet the change has wide reaching consequences. It seems that something is being lost over time. In the old days of running on bare metal, yes servers failed for various reasons, then we added resiliency techniques whose sole purpose was to alleviate downtime. Now we're at highly complex distributed systems that have failed to keep the resiliency up there. But the fact that all the mega-corps have had these issues seems to indicate a systemic problem rather than unconnected ones. Perhaps a connection is the management techniques or HR hiring practices? Perhaps it's due to high turnover causing the issue? (Not that I know, of course, just throwing it out there). That is, are the people well looked after and know the systems that are being maintained? Even yourself who's 'been around the traps' with high profile companies: you have moved around a lot... Were you unhappy with those companies that caused you to move on? We've seen multiple stories here on HN about how those people in the 'maintenance' role get overlooked for promotions, etc. Is this why you move around? So, perhaps the problem is systemic and it's due to management who've got the wrong set of metrics in their spreadsheets, and aren't measuring maintenance properly?
- bdd 7y agoI don't see the value in lamenting the old days of a few machines where you could actually name them as Middle Earth characters, install individually, log in to one single machine to debug a site issue. The problems were smaller and individual server capacity in respect to demand was in meaningful fractions. Now the demand is so high and set of functions these big companies need to offer are so large, it's unrealistic to expect solutions that doesn't require distributed computing. It comes with "necessary evils", like but not limited to configuration management--i.e. ability to push configuration, near real time, without redeploying and restarting--, and service discovery--i.e. turning logical service names to a set of actual network and transport layer addresses, optionally with RPC protocol specifics. I refer to them as necessary evils because the logical system image of these are in fact single points of failures. Isn't it paradoxical? Not really. We then work on making these systems more resilient to the very nature of distributed systems, machine errors. Then again, we're intentionally building very powerful tools that can also enable us to take everything down with very little effort because they're all mighty powerful. Like the SPoF line above, isn't it paradoxical? Not really :) We then work on making these more resilient to human errors. We work on better developer/operator experience. Think about automated canarying of configuration, availability aware service discovery systems, simulating impact before committing these real time changes, etc. It's a lot of work and absolutely not a "solved problem" in a way single solution will work for any scale operation. We may be great at building sharp tools but we still suck at ergonomics. When I was at Twitter, a common knee-jerk comment at HN was "WTF? Why do they need 3000 engineers. I wrote a Twitter clone over the weekend". A sizable chunk of that many people work on tooling. It's hard. You're pondering if hiring practices and turnover might be related? The answer is an absolute yes. On the other hand, these are the realities of life in large tech companies. Hiring practices change over years because there's a limited supply of of candidates experienced in such large reliability operations and industry doesn't mint many of them either. We hire people from all backgrounds and work hard on turning them to SREs or PEs. It's great for the much needed diversity (race, gender, background, everything) and I'm certain the results will be terrific but we need many more years of progress to declare success and pose in front of a mission accomplished banner on an aircraft carrier ;) You are also wisely questioning if turnover might be contributing to these outages and prolonged recovery times. Without a single doubt, again the answer is yes but it's not the root cause. Similar to how hiring changes as company grows, tactics for handling turnover has to change too. It's not like people leave the company, but within the same company they move on and work on something else. The onus is on everyone, not just managers, directors, VPs to make sure we're building things where ownership transfer us 1) possible 2) relatively easy. This in mind, veterans in these companies approach code reviews differently. If you have tooling to remove the duty of nitpicking about frigging coding style, and applying lints, then humans can indeed give actually important feedback on complexity of operations, self describing nature of code, or even committing things along with changes to operations manual living in the same repo. I think you're spot on with your questions but what I'm trying to say with this many words and examples is, nothing alone is the sole perpetrator of outages. A lot of issues come together and brew over time. Good news, we're getting better. Why did I move around? Change is what makes life bearable. Joining Twitter was among the best decisions in my career. Learned a lot, made lifelong friends. They started leaving because they were yearning a change Twitter couldn't offer. I wasn't any different. Facebook was a new challenge, I met people I'd love to work with and decided give it a try. I truly enjoy life there even though I'm working on higher stress stuff. Facebook is a great place to work but I'm sure I can't convince even %1 of HN user base, so please save your keyboards' remaining butterfly switch lifetime, don't reply to tell me how much my employer sucks :) I really hope you do enjoy your startup jobs (I guess?) as much as I do my big company one.
- GrumpyNl 7y agoTurned out to be number #1 The outage was due to an internal configuration change, which we're now fixing. Some people may be able to access Twitter again and we're working to make sure Twitter is available to everyone as quickly as possible.