12 ms·
Cloudflare outage caused by bad software deploy
- tomcam 7y agoI run a service placing bids the last few seconds on eBay. Every time this happens I lose measurable business (we place thousands of bids per day). While it doesn’t affect scheduled bids, they can’t place bids and are likely to move to a competitor. These recent outages have been costly. Does anyone know a more reliable provider?
- sbr464 7y agoFrom experience, using two services (active/active) is really the only way to avoid downtime. DNS can be trickier, but there are providers that can fallback automatically or split requests (dnsmadeeasy etc)
- deforciant 7y agoMaybe use cloudflare only for the landing page and docs but serve your bidding app frontend and backend directly? Since users are already there it will only affect first load :)
- tomcam 7y agoThat is exactly how we use Cloudflare, but I appreciate the guidance. I am new to devops.
- PhasmaFelis 7y agoHaha, people pay you to bid-snipe on eBay for them? It just baffles me that manual/third-party bid-sniping is still a thing. eBay has had automatic bidding for more than twenty years. You'll pay the same whether you put in the winning bid a week in advance or 5 seconds. But people see that "you lost this auction" notice and they're irrationally convinced that it would have gone differently if they'd bid at the last minute, somehow.
- beering 7y agoIf other bidders are irrational, then bid-sniping can work. It doesn't give others the opportunity to contemplate, "I've been out-bid, do I actually want this item more than I originally thought?"
- neilv 7y agoAnd it's well-known since early eBay days that many bidders are irrational, including but not limited to competitive impulse to "win". Plus you sometimes have shill bidders. Sniping approximates sealed bids, with the highest-bidder the second-highest sealed bid amount or a small increment above it. (Unfortunately for eBay, that would tend to decrease their cuts, unless the appeal of the sealed bid format brings in sufficiently more bidder activity.) Another advantage of the software/service is that it automates. If you want to buy a Foo, you can look at the search lists, find a few Foos (possibly of varying value in the details), say how much you'll pay for each one, and let the software attempt to buy each one by its auction end until it's bought one, then it stops. If eBay implemented this itself, it might be too much headache in customer support, but third-parties could provide it to power users. (I don't buy enough on eBay anymore to bother with anything other than conventional manual bids, but I see the appeal of automation.)
- steve_adams_86 7y agoIf everyone is using a bidder like this, isn't it essentially like a blind auction?
- skibble 7y agoThis absolutely does work. If you get outbid normally on eBay, say at least half an hour or so before the auction ends, then people get irrational and you quickly end up in a bidding war. This way, as a bidder, you set the maximum price you are willing to pay outside of a heated 'damn, I missed out' mentality and people don't have time to respond. Also this helps when placing bids on auctions that end when you're asleep and you want to effect the above. I've tried to do it manually in the past but naturally life intervenes and you're somewhere with no signal or in the middle of something. Having tried it both ways sniping is definitely better than regular bidding.
- dfga 7y agohttps://opinionator.blogs.nytimes.com/2013/03/30/those-irritating-verbs-as-nouns/ https://opinionator.blogs.nytimes.com/2013/03/30/those-irrit...
- eganist 7y agoKinda wonder at this point what findings exist on their Availability SOC 2, assuming they've gotten one. The repeated outages plus the constant malicious advertising by scammy ad providers through cloudflare are slowly turning me off to the service as a potential enterprise customer. Unfortunate too since plenty of superlatively qualified people build great things there (hat tip to Nick Sullivan), but it seems like the build-fast culture may now be impeding the availability requirements of their clients. This is also a great example of a case where SLAs are meaningless without rigorous enforcement provisions negotiated in by enterprise clients. Cloudflare advertises 100% uptime (https://www.cloudflare.com/business-sla/ https://www.cloudflare.com/business-sla/) but every time they fall over, they're down for what, an hour at a time? Just this one issue would've blown anyone else's 99.99% SLA out of the water -- https://www.cloudflarestatus.com/incidents/tx4pgxs6zxdr https://www.cloudflarestatus.com/incidents/tx4pgxs6zxdr I love the service, but if I'm to consider consuming the service, they'd do well to have the equivalent of a long term servicing branch as its own isolated environment, one where changes are only merged in once they've proven to be hyper-stable.
- klodolph 7y agoAs an engineer, I get pissed whenever I see 100% uptime, or eleven-nines, nine-nines, or other impossible targets. Like, how am I supposed to design a system with numbers like that?
- Lorin 7y agoDeploy once, never update, and deploy a missile defense to prevent backhoes from digging up fiber?
- klodolph 7y agoAh yes, missile defenses, like the MIM-104 Patriot: https://en.wikipedia.org/wiki/MIM-104_Patriot#Failure_at_Dhahran https://en.wikipedia.org/wiki/MIM-104_Patriot#Failure_at_Dha...
- JaimeThompson 7y ago
- souterrain 7y agoCloudflare should write a guide to doing post-event communication. Or perhaps they shouldn’t, as this seems to be a potential differentiator. This is direct and doesn’t attempt to avoid blame. Well done.
- Pfhreak 7y agoAvoiding blame is different than acknowledging responsibility. A post mortem should be very conscious about blame - never target the engineer who deployed the change, for example. Take responsibility for the machine that allowed the unsafe change to be deployed. (Where machine could be tooling or process, as appropriate.)
- rodgerd 7y agoKarma for shitting on Verizon, maybe.
- neom 7y agoI'm always beyond impressed with how responsive and transparent CF is with incidence and post mortem communication. Given who the CEO and COO are, I suppose this shouldn't be surprising, never the less as a customer it builds a great deal of trust. Kudos.
- grey-area 7y agoYes, they do really well on this - open, transparent, posting information quickly as soon as they were fairly sure what the problem was. I always really enjoy their writing, both incident reports and writeups of new features. The only thing I think they could have managed better was their status page, which claimed they were up (every service was green) when they were not.
- deleted 7y ago[deleted]
- lanrh1836 7y agoA quick look at their Glassdoor reviews paints a very different story if the reviews are to be believed...
- PopeDotNinja 7y agoI don't understand what you are saying.
- gist 7y agoNothing like having what should be a world class company falling prey to the same type of screw-ups that plaque 'the local guy maintaining some wordpress site on a shared server'. Separately there is nothing that says that a company like Cloudflare has to air their dirty laundry (as the saying goes). The vast majority of 'customers' really don't care why something happened at all or the reason. All they know is that they don't have service. Pretend that the local electric company had a power outage (and it wasn't caused by some obvious weather event). Does it really matter if they tell people that 'some hardware we deployed failed and we are making sure it never happens again'. I know tech thinks they are great for these types of post-mortems but the truth is only tech people really care to hear them. (And guess what all it probably means is that that issue won't happen again...)
- chachachoney 7y ago>> Pretend that the local electric company had a power outage (and it wasn't caused by some obvious weather event). >> Does it really matter if they tell people that 'some hardware we deployed failed and we are making sure it never happens again'. It does, and most people would be relieved to hear about an equipment failure rather than a malicious employee, poor security, etc...
- ngold 7y agoI fail to see the point of your post? Are you arguing that less transparency and information is a good thing? I doubt any of the people you are talking about "not caring" read this site to begin with.
- GranPC 7y ago> I know tech thinks they are great for these types of post-mortems but the truth is only tech people really care to hear them. Well, Cloudflare is in luck; most of their customers are "tech people"!
- gist 7y ago100 not true. All you have to do is pull a list of the daily additions and deletions and you will see that they have many customers that are not 'tech' people. Further you are assuming all the customers of theirs that are tech people even read and keep up with blog posts like this.
- lgats 7y agoAt 1402 UTC we understood what was happening and decided to issue a ‘global kill’ on the WAF Managed Rulesets, which instantly dropped CPU back to normal and restored traffic. That occurred at 1409 UTC. So for about 50 minutes, those who relied on the WAF were open to attack?
- foota 7y agoIsn't open to DDoS better than can't be reached?
- detaro 7y agoDepends on the relative costs of the two options?
- hunter2_ 7y agoCan you give an example of where the cost of possibly-denied could ever be higher than definitely-denied? First Cloudflare literally denied service, then as a hotfix there was a higher-than-normal potential for denying service, and eventually the normal potential for denying service was restored. I'm trying to comprehend how the second phase could ever be worse than the first phase. Now, if you're talking about elevating the potential for compromised confidentiality and/or integrity rather than merely availability, I'd agree, but generally [D]DoS refers to availability. Leaning on a WAF to plug gaping vulnerabilities that can be discovered and exploited during the period of time before the WAF was restored means you have much bigger problems than uptime.
- detaro 7y ago> Leaning on a WAF to plug gaping vulnerabilities that can be discovered and exploited during the period of time before the WAF was restored means you have much bigger problems than uptime. It's also, roughly speaking, the selling point of products called "WAF". (and yes, relying on them is not great)
- javagram 7y ago
- cfors 7y agoI wonder what happened with that poor regex expression. My thoughts are immediately shifting to one of my favorite articles of all time "Regular Expression Matching can be Simple and Fast..." [0] [0] https://swtch.com/~rsc/regexp/regexp1.html https://swtch.com/~rsc/regexp/regexp1.html
- djhworld 7y agoA good war story there, at least the problem was relatively simple and quick to identify as the root cause, rather than something deeper. Would be interested to see what the gnarly regex was that was bombing their CPUs so hard!
- BentFranklin 7y agoKind of funny that it was a regexp.
- davidw 7y agoI'm reminded of: "You have a problem, and you decide to use a regexp to solve it. Now you have two problems" Although of course I'm just kidding and I'm sure that a good regexp probably is the right solution for what they're doing in that instance: they have a lot of bright people.
- sequoia 7y agoWhat sort of regular expression pitfalls can cause this sort of CPU utilization? I know they're possible but I am curious about specific examples of something similar to what caused Cloudflare's issue here.
- novas0x2a 7y agoSome regex languages allow backtracking, and backtracking is usually the thing that causes regexes to blow up in resource cost: https://www.regular-expressions.info/catastrophic.html https://www.regular-expressions.info/catastrophic.html
- edwintorok 7y agoYou probably want a regex engine that runs in linear time: * Google's RE2 https://github.com/google/re2/wiki/WhyRE2 https://github.com/google/re2/wiki/WhyRE2 * https://github.com/laurikari/tre/ https://github.com/laurikari/tre/ There is a good series of articles about the problem: https://swtch.com/~rsc/regexp/regexp3.html https://swtch.com/~rsc/regexp/regexp3.html I would strongly recommend deploying such a regular expression matcher to avoid problems like this. There are examples in the above article that you can use to test anything in your production deployment that accepts regular expressions to see how well it copes.
- novas0x2a 7y agoMight have been a misdirect reply (although useful), but yeah, agree, linear-time regex engines are generally a much better idea.
- roro159 7y agoDoS with regex is a thing: https://www.owasp.org/index.php/Regular_expression_Denial_of_Service_-_ReDoS https://www.owasp.org/index.php/Regular_expression_Denial_of... StackOverflow had a similar case a while back: https://stackstatus.net/post/147710624694/outage-postmortem-july-20-2016 https://stackstatus.net/post/147710624694/outage-postmortem-...
- peterwwillis 7y agoHow to implement a multi-CDN strategy (streamroot.io): https://news.ycombinator.com/item?id=18399523 https://news.ycombinator.com/item?id=18399523 Etsy implementing multiple CDN (7 years ago, the CDNcontrol project looks abandoned): https://speakerdeck.com/ickymettle/integrating-multiple-cdn-providers-our-experience-at-etsy https://speakerdeck.com/ickymettle/integrating-multiple-cdn-... https://dyn.com/blog/speaking-with-etsy-about-multi-cdns-and-dns/ https://dyn.com/blog/speaking-with-etsy-about-multi-cdns-and... Basically: you can try to keep a low TTL DNS, but it'll be more DNS traffic, and 5-10% of traffic takes forever to cut over because nobody respects TTL. Worst case you have just as much down time as before, best case most of your traffic is recovered in a few minutes.
- outworlder 7y agoIt may be useful to note, for whoever is reading this, that low DNS TTL only ever makes sense for anything that you can do a cutover either automatically or on short notice, not for all records. Otherwise, you are now at mercy of outages on your DNS providers. Just leaving it out there so one doesn't get the idea that "low TTL == Always Good"
- rob-olmos 7y agoFor the size and importance of Cloudflare some insights to a couple questions would be nice: 1. Why are WAF rules not progressively deployed since there's already a system to do so? 2. Maybe there should also be a testing environment that receives a mirror of production traffic before deployments reach real users? (I understand the WAF change was not set to take action, but a separate environment would be less likely to affect production)
- JakeTheAndroid 7y agoCloudflare does have test colos that use a subset of real network traffic for testing. It's actually the primary testing methodology, and the employees are usually some of the first people forced through test updates. This release wasn't meant to go out, and the fact it did means it would have bypassed the test environments either way.
- nodesocket 7y agoIt is interesting NGINX returned 502 nearly instantly under very heavy CPU load. I would have expected requests to just hang or timeout.
- docapotamus 7y agoI would imagine it's tiered. The Nginx servers at the front returning the 502 probably aren't the boxes running the code
- jgrahamc 7y agoyes
- grey-area 7y agoI really want to know the regexp and corresponding input(s) which killed the internet now :) Was it just aaaaaaaaaaaah? https://swtch.com/~rsc/regexp/regexp1.html https://swtch.com/~rsc/regexp/regexp1.html
- almost_usual 7y agoI'm assuming it's something pretty embarrassing if it's not in the post mortem.
- hunter2_ 7y agoThe first sentence here is "This is a short placeholder blog and will be replaced with a full post-mortem...". I'd bet big money that they do include it.
- almost_usual 7y agoLook forward to the full post-mortem
- snarf21 7y agoIt reminds me of an old joke. I decided to solve a software problem with regular expressions and now I have two problems.
- gbrayut 7y agoProbably something mundane like ^[\s\u200c]+|[\s\u200c]+$ That's the one that took down Stack Overflow a few years ago https://stackstatus.net/post/147710624694/outage-postmortem-july-20-2016 https://stackstatus.net/post/147710624694/outage-postmortem-...
- UI_at_80x24 7y agoCan anybody suggest a Systems Engineer-centric forum/site? (Not Windows 'help I can't print' level, more DataCenter grade.) HN does have some great content/replies that touch on these topics, but I'd like something more.
- gridspy 7y agoPerhaps these QA sites are interesting? https://superuser.com/ https://superuser.com/ https://serverfault.com/ https://serverfault.com/ But yes, the content like this on HN is fascinating and I would also like more.
- aleem 7y agoIf a single regex can take down the Internet for a half hour, that's definitely not good -- for a class of errors that can be easily prevented, tested, etc. The timing is unfortunate too, after calling out Verizon for lack of due process and negligence. I'm sure they have an undo or rollback for deployments but probably worth investing into further. They also need to resolve the catch-22 where people could not login and disable CloudFlare proxy ("orange cloud") since cloudflare.com itself was down.
- pfundstein 7y ago> The timing is unfortunate too, after calling out Verizon for lack of due process and negligence. Nonetheless, Verizon could take a leaf out of their responsiveness and transparency book.
- lima 7y agoYeah. They criticized Verizon for being unresponsive. Mistakes happen.
- _wmd 7y agoYou'd think after leaking private data for literally months less than 3 years ago (and only noticing because Google had to point it out to them) that they'd, y'know, have at least some kind of QA environment fed with sample traffic by now. Really hard to believe they're still getting caught testing in prod
- rainyMammoth 7y agoFor working in that field, the arrogance of CloudFlare is still unbelievable to me. After their huge Cloudbleed issue with the addition of this one, they continue to call out everyone through their blog posts. And everyone seems fine with it because they are a hype company.
- issati 7y agoI don't use CloudFlare nor have any interest in them, but I don't see the arrogance. The issues CloudFlare have are things everyone takes seriously and are working very hard on. Deployment and memory safety are hard problems that happens to the best of the best. It happens Google, Amazon and Facebook. If anything the idea that this would damaging, because it is more public, is arrogant. If CloudFlare would be saying that everything is fine you might have a point, but they aren't. Just like the other companies mentioned they seem to be improving their routines, programming and infrastructure to try and mitigate these problems. What they are criticising however are things like not adopting new protocols or not taking things that affects everyone seriously. This isn't something that would happen if people were trying. And the response from some of the industry is "we know what we are doing", and shortly after the same thing happens again and again and again. So I don't really see CloudFlare being that arrogant, if anything it's the "you are not better than us" from some parts of the industry that is. The day I see CloudFlare not trying I would be happy calling them arrogant. But if anything I would caution that they are too successful by trying more than most.
- bdibs 7y agoSeems like a terrible idea to deploy changes to such a vital piece of their software GLOBALLY without some sort of rollout procedure.
- quickthrower2 7y agoProbably for the kind of work they are doing avoid regex? Or at least the very complicated modern regex (simple autonoma that you can compile in advance might be ok)
- txcwpalpha 7y agoIf you're trying to do pattern matching, is there actually a widely used alternative to regex? The more I can avoid using regex for mission-critical things, the happier I will be, but I'm really not aware of anything better for this type of application.
- quickthrower2 7y agoIve tried parser combinators. They are nice but a bit more labour than writing out a regex and I’m not sure how performance compares
- txcwpalpha 7y agoPeculiar that eastdakota (Cloudflare's CEO) doesn't seem to be tweeting at the Cloudflare team responsible for this, telling them they should be ashamed and are guilty of malpractice. When it was Verizon that took down the internet he felt it was appropriate to do that to the Verizon teams, after all. edit: right after posting this comment, he did tweet the following: https://twitter.com/eastdakota/status/1146196836035620864 https://twitter.com/eastdakota/status/1146196836035620864 > I'd say both we and Verizon deserve to be ashamed. As well as this: https://twitter.com/eastdakota/status/1146170209780113408 https://twitter.com/eastdakota/status/1146170209780113408 > Our team should be and is ashamed. And we deserve criticism. ... I still don't think that publicly shaming anyone is a good leadership style nor is it a good way to motivate people to perform better in the future, but kudos for the self-awareness, at least.
- threezero 7y agoCloudflare was responsive and reasonable. Verizon was unreachable and deflected responsibility when they finally made a statement. And public shaming does often motivate companies to be more responsive to their customers.
- txcwpalpha 7y agoAFAIK Cloudflare isn't in any way a "customer" of Verizon. Verizon doesn't owe Cloudflare any kind of response or devotion of resources. Verizon owes it's actual customers a resolution to their problem, which they gave. I'm not saying Verizon is perfect nor absolved of fault, but Cloudflare was/is not owed any kind of explanation or assistance by VZ, and it's absurd of CF to still be whining about that fact (as they are doing in some other tweets today). If CF wants some kind of SLA with VZ, they should engage them in a business relationship, not try to publicly shame them.
- thegagne 7y agoI’d say what they really need is a representative governing body over major network carriers to establish proper standards and levy fines for those that do not comply. Kind of similar to a homes association saying “hey that trash on your lawn affects your neighbor, clean it up!” It’s true that they are not a customer but at that level what they do affects each other, and it’s better to resolve things civilly and privately instead of publicly on twitter.
- suchow 7y agoIs there a usage error in the first sentence or has English lost the blog / blog post distinction?
- tom_ 7y agoSome people distinguish between the two, some don't.
- deleted 7y ago[deleted]
- rubyn00bie 7y agoThis line kills me: > We were seeing an unprecedented CPU exhaustion event, which was novel for us as we had not experienced global CPU exhaustion before. I'd imagine it was quite novel for most anyone affected /s
- ksara 7y ago>"It doesn't cost a provider like Verizon anything to have such limits in place. And there's no good reason, other than sloppiness or laziness, that they wouldn't have such limits in place."[1] [1] https://blog.cloudflare.com/how-verizon-and-a-bgp-optimizer-knocked-large-parts-of-the-internet-offline-today/ https://blog.cloudflare.com/how-verizon-and-a-bgp-optimizer-...
- Operyl 7y agoThe difference between Verizon and Cloudflare in this case is that Cloudflare generally _does_ fix their mistakes when they screw up (and generally don't make the same type again).. whereas Verizon has screwed up internet routing more times than I'd like to think about. No company is perfect, but I'd say this comparison is pretty apples to oranges.
- ti_ranger 7y ago> We make software deployments constantly across the network and have automated systems to run test suites and a procedure for deploying progressively to prevent incidents. Good. > Unfortunately, these WAF rules were deployed globally in one go and caused today’s outage. Wow. This seems like a very immature operational stance. Any deployment of any kind should be subject to minimum deployment safety, that they claim they have. > At 1402 UTC we understood what was happening and decided to issue a ‘global kill’ on the WAF Managed Rulesets, which instantly dropped CPU back to normal and restored traffic. That occurred at 1409 UTC. Many large companies would have had automatic roll-back of this kind of change in less time than it took CloudFlare to (apparently) have humans decide to roll-back, and possibly before a single (usually not global) deployment had actually completed on all hosts/instances. However, what is more concerning is that it seems you shouldn't rely on CloudFlare's "WAF Managed Rulesets" at all, since they seem to be willing to turn it off instead of correctly rolling back a bad deployment, which they only did > 43 minutes later: > We then went on to review the offending pull request, roll back the specific rules, test the change to ensure that we were 100% certain that we had the correct fix, and re-enabled the WAF Managed Rulesets at 1452 UTC. How were they not able to trivially roll back to the previous deployment?
- londons_explore 7y agoSo many employees deploying so many changes at a time it wasn't clear which one was the cause...?
- pornel 7y agoDeployments are scheduled and managed by the SRE team.
- mschuster91 7y agoWhich is why the entire (mostly in Agile environments) model of "deploy to prod as soon as you can" is absolute nuts. If you're dev at a hipster app maybe a dozen people use to holler "yo" at each other, by all means go for it. If you're operating one of the biggest and most important chonks of Internet infra... maaaaaybe stick to established practices such as stage testing, release schedules and incremental rollouts?
- BuddhaSource 7y agoDo builds go through stage role out? For service like Cloudflare.
- tolgahanuzun 7y agoIt is very difficult to explain this to customers who don't understand technology. 30 minutes is a very big time. :/
- mrzasa 7y agoShameless plug: understanding regex engine implementation can help with avoiding performance pitfalls: https://medium.com/textmaster-engineering/performance-of-regular-expressions-81371f569698 https://medium.com/textmaster-engineering/performance-of-reg...
- pearapps 7y agoNo way!?!?!??!?!?!?!??!?!??!?!?!