4 ms·
They have now released an initial statement [1]: For about 30 minutes today, visitors to Cloudflare sites received 502 errors caused by a massive spike in CPU
by TomAnthony 7y ago
They have now released an initial statement [1]:
For about 30 minutes today, visitors to Cloudflare sites received 502 errors caused by a massive spike in CPU utilization on our network. This CPU spike was caused by a bad software deploy that was rolled back. Once rolled back the service returned to normal operation and all domains using Cloudflare returned to normal traffic levels.
This was not an attack (as some have speculated) and we are incredibly sorry that this incident occurred. Internal teams are meeting as I write performing a full post-mortem to understand how this occurred and how we prevent this from ever occurring again.
[1] https://blog.cloudflare.com/cloudflare-outage/ https://blog.cloudflare.com/cloudflare-outage/
- ilkkao 7y agoInteresting to hear later if the process was followed. Hard to believe they deploy changes to every production location at once.
- jgrahamc 7y agoSorry, I would have posted this myself but was too busy.
- technonerd 7y agoAhh the "we test in prod" method
- jgrahamc 7y agoYeah, except that wasn't meant to be the way things work.
- js2 7y ago> Unfortunately, one of these rules contained a regular expression that caused CPU to spike to 100% on our machines worldwide. Sounds like backtracking. If so, I'll bet there's a conversation happening about switching to re2. edit: hmmm, https://github.com/cloudflare/lua-re2 https://github.com/cloudflare/lua-re2 I'm also curious why this rollout wasn't staged.