20 ms·
Fastly Outage
- alixaxel 5y agoIndeed, part of GitHub (.io) too.
- schappim 5y agoParts of Shopify
- an0n4u 5y agonumpy docs, too. i think it's cloudflare related as well. at least, I keep seeing some cloudflare errors interpolated with the 503 varnish error.
- lapink 5y agoPytorch and Python docs, all down. No stackoverflow. I guess this is a forced bank holiday for developers around the world.
- AkshitGarg 5y agoWell they thought that using a CDN over a CDN would be a good idea
- somishere 5y agoWe've got Cloudflare sitting in front of our Firebase/GCP instance (which I've just found out is Fastly-cached :/). Getting 503s at the origin but we're up on our URL with an always online notice thanks to CF. Double dip isn't all that bad.
- dilawar 5y agoI think reddit in India is down as well.
- threeseed 5y agoShopify's CDN is down. Which is causing $15+ million in lost product sales for every hour of outage. Not to mention the loss of any new customers.
- dspillett 5y agoStackOverflow and all the StackExchange family of sites are down. I suspect the lost productivity from that will be more costly over the whole economy than potential lost sales via Shopify. People can go back to shopify so those transactions not definitely blocked for ever, any time "lost" due to reference resources being unavailable can't so easily be claimed back.
- threeseed 5y agoI don't think you understand how ecommerce works. A very significant amount of people won't go back. It's why the most effective marketing campaign by far is retargeting those people to convince them to come back. Unfortunately that's not possible in this case since you can't track the users as the site is unusable.
- twic 5y agoSome sites are on Cloudflare, right? Looks like we have a natural experiment to test this belief!
- gezfrg321 5y ago> A very significant amount of people won't go back So they didn't need what they were about to purchase and saved their money. Doesn't sound like a net loss to me.
- dspillett 5y ago> I don't think you understand how ecommerce works ... people won't go back I was talking about the economy in general, not specific e-commerce sites. People that actually need what they were looking for but don't go back will buy it elsewhere. The money still flows, just somewhere else. And if they don't need the item(s), they'll perhaps use the money for something more useful.
- optiomal_isgood 5y agoFWIW, Fastly ~8 hours ago (3am UTC) reported another incident: https://status.fastly.com/incidents/1glxxb8sf2zv https://status.fastly.com/incidents/1glxxb8sf2zv and deployed a fix—either the fix made it worse or wasn't sufficient to mitigate the problem.
- sjaak 5y agoPerhaps Fastly is simply taking their commitment to reducing CO2 seriously? Three hurrays for the climate!
- jl6 5y agoWorrying that this is impacting so many dev toolchains and services, which will hinder the ability to respond to the issue.
- tus88 5y agoReddit, Paypal, Guardian, ElasticSearch...crap.
- sleepyshift 5y agoLooks like this has taken out Reddit at least.
- algo_cheese 5y agoAnd a large part of GitLab
- spyke112 5y agoIs it also hitting Github? I'm not getting any css when loading Github.
- nevi-me 5y agoLooks like it is. If you're still able to see much of the UI, don't force-reload the page as it'll invalidate the CSS in the cache. I did that moments ago, and I regret it.
- Jamie9912 5y agoYep, seems like: Reddit BBC News Twitch.tv Twitter emoji cdn? are all down 503 service error
- another-dave 5y agoAh didn't cop that Twitter emoji issue was related! Thought an ad-blocker was stepping up its filters aggressively :) Stack Overflow, The Guardian, Gov.uk too as some other biggish names getting hit.
- strogonoff 5y agoVarious bits of GitHub on the Web (committing edits, editing releases) were broken for the same reason. Failure modes of JS-heavy GUIs are interesting.
- deleted 5y ago[deleted]
- mlnj 5y agoStackOverflow too.
- lpmitchell 5y agoThis seems to be impacting a number of huge sites, including the UK government website[0]. [0] https://www.gov.uk/ https://www.gov.uk/ https://m.media-amazon.com/ https://m.media-amazon.com/ https://pages.github.com/ https://pages.github.com/ https://www.paypal.com/ https://www.paypal.com/ https://stackoverflow.com/ https://stackoverflow.com/ https://nytimes.com/ https://nytimes.com/ Edit: Fastly's incident report status page: https://status.fastly.com/incidents/vpk0ssybt3bj https://status.fastly.com/incidents/vpk0ssybt3bj
- c-fe 5y agoalso https://www.reddit.com https://www.reddit.com (at least in Netherlands) edit: 12:05 up again for me, no images or custom fonts loading though ... and down again 1 minute later edit: 13:01 reliably up again for me
- deleted 5y ago[deleted]
- secondcoming 5y agoDown in UK
- tacticalmook 5y agoDown in US. Also Imgur, which is closely related
- floating_panda 5y agoDown in india
- silviot 5y agoSame here in Germany: imgur and reddit are down, plus a bunch of other sites.
- tarsinge 5y ago
- johnstonnorth 5y agorubygems.org affected too
- csmattryder 5y agoHere's the status page incident for this. https://status.fastly.com/incidents/vpk0ssybt3bj https://status.fastly.com/incidents/vpk0ssybt3bj
- algo_cheese 5y ago> We're currently investigating potential impact to performance with our CDN services. Guys, you are offline with a 503 error, this is a little more than "potential impact to performance".
- superzamp 5y agoLowkey status reports are the norm now :) "some users may experience degraded service" => site completely down for all locations
- sdflhasjd 5y agoI fully expect that if I find a "major outage" on Slack's status page that it could only mean the outbreak of nuclear war.
- TheDauthi 5y ago"Some users may experience brief service disruption."
- tommica 5y ago"By the account of them and us being completely vaporized"
- jconnop 5y agoI was going to link the appropriate XKCD where organised attackers are panicing as they realise they're dealing with a sysadmin muttering about uptime.. .. but of course XKCD is down too. e: https://xkcd.com/705/ https://xkcd.com/705/
- creamyhorror 5y agobasically the internet is down reddit, stackoverflow, github, paypal, pypi, twitter, twitch, NYT, CNN, BBC, the Guardian... edit: wow, even Amazon.com relies on Fastly for some of its edge caches!
- 3np 5y agodebian's main apt repo mirror affected as well
- secondcoming 5y agoBBC is still up at least in the UK
- iso1631 5y agoNot here (although won't be long) dig bbc.co.uk bbc.co.uk. 193 IN A 151.101.64.81 bbc.co.uk. 193 IN A 151.101.128.81 bbc.co.uk. 193 IN A 151.101.192.81 bbc.co.uk. 193 IN A 151.101.0.81
- deleted 5y ago[deleted]
- easytiger 5y agoit's down
- fredoralive 5y agoSeems to be mixed for me, BBC News and Sport works but stuff like Weather, iPlayer (video streaming) and Sounds (audio streaming) have died. I guess the BBC is big enough that different bits of the site run off different solutions (perhaps news and sport are still in spirit running off "news.bbc.co.uk" instead of the main servers?).
- iso1631 5y agohttps://www.washingtonpost.com/technology/2020/04/06/your-internet-is-working-thank-these-cold-war-era-pioneers-who-designed-it-handle-almost-anything/ https://www.washingtonpost.com/technology/2020/04/06/your-in... “This basic architecture is 50 years old, and everyone is online,” Cerf noted in a video interview over Google Hangouts, with a mix of triumph and wonder in his voice. “And the thing is not collapsing.” The Internet, born as a Pentagon project during the chillier years of the Cold War, has taken such a central role in 21st Century civilian society, culture and business that few pause any longer to appreciate its wonders — except perhaps, as in the past few weeks, when it becomes even more central to our lives.
- nindalf 5y agoTaken out xkcd as well.
- iso1631 5y agoIsn't there an xkcd comic about CDN failures?
- deleted 5y ago[deleted]
- SSLy 5y agohttps://xkcd.com/2347/ https://xkcd.com/2347/
- loulouxiv 5y agoxkcd is down too :(
- azureel 5y agoMaybe this one, titled "The Cloud". https://xkcd.com/908/ https://xkcd.com/908/
- clydethefrog 5y agoMight be xkcd.com/503.
- jthetzel 5y agoWe have no way to know. https://xkcd.com/908/ https://xkcd.com/908/
- Archit3ch 5y agoToday's comic is titled "Product Launch", so the joke still works if you assume it's about a disastrous launch. ;)
- schappim 5y agoSMH.com.au
- pts_ 5y agoAre these sites on the same cloud or CDN?
- innocenat 5y agoThey are all on Fastly CDN...?
- deleted 5y ago[deleted]
- deleted 5y ago[deleted]
- mcintyre1994 5y agoDo they have an official status page? Googling gets https://docs.fastly.com/en/guides/fastlys-network-status https://docs.fastly.com/en/guides/fastlys-network-status which is 503 Edit: Elsewhere in the comments: https://status.fastly.com/incidents/vpk0ssybt3bj https://status.fastly.com/incidents/vpk0ssybt3bj
- deleted 5y ago[deleted]
- asicsp 5y agoRelated thread: https://news.ycombinator.com/item?id=27432397 https://news.ycombinator.com/item?id=27432397
- ronyfadel 5y agoTen Percent Happier is down, and now my day is ruined.
- selykg 5y agoWhen viewing a meditation session you can see a download button in the upper right (at least on iOS). I always have a small stash of my favorites saved locally in case of internet outage or I’m caught in a situation where I don’t have internet but need a few minutes. On top of that I’ve been really trying to rely less on an app. So I throw a lightly guided or unguided session in every couple days at least where I focus on going solo so I don’t need an app and just need a timer.
- alexannic 5y agocnn.com is down as well.
- deleted 5y ago[deleted]
- Nilef 5y agoIronically, even this Outage page is out for me
- misnome 5y agopypi.org, but not https://status.python.org/ https://status.python.org/ - I'm impressed that they actually hosted the status page differently!
- JulianWasTaken 5y agoThat's fairly standard practice. Fastly itself has its status page up as well: https://status.fastly.com/ https://status.fastly.com/
- navanchauhan 5y agoNo wonder, The Verge and NYT are down too.
- magicturtle 5y agoreddit down aswell
- tfar 5y agohttps://flutter.dev/ https://flutter.dev/ and https://fastlane.tools/ https://fastlane.tools/ as well.
- barosl 5y agoI didn't know so many sites were depending on Fastly. Stack Overflow, GitHub, reddit, .... Even pip is unavailable. My development workflow is completely janked up. It is a bit scary that we are putting too many eggs in one basket.
- liveoneggs 5y agofastly gives free service to things like pip. It's actually very nice.
- JulianWasTaken 5y agoBit pedantic, but it's PyPI that Fastly gives services to, not pip (and PyPI that's down, not pip). The two are only loosely related – pip is a piece of software.
- joshenders 5y agoBlame site operators that are single homing and not loadbalancing CDNs
- notyourday 5y agoFor sites of any complexity with any dynamic content having CDN redundancy is akin to being multi-cloud -- it is not worth the effort. A lot of dynamic sites use Fastly for its programmatic edge control and a near immediate ( ~1s-4s, typically around 2 ) global cache invalidation for any tagged objects with a single call to the tag. That feature alone simplifies backend logic significantly. To make this feature portable to CDNs that do not support it and provide only regular cache invalidation requires a complicated workflow setup which significantly increases the cache bust time, which in turn removes all the advantages of the treat dynamic content as static and cache bust on write approach.
- joshenders 5y ago>> For sites of any complexity with any dynamic content having CDN redundancy is akin to being multi-cloud — it is not worth the effort. I proposed and lead our multi-CDN project at Pinterest for both static and dynamic content and I can tell you, many many times over, it has been well worth the effort. Everybody should do this if not only for contract negotiating leverage. Cache invalidation is fast enough on all CDNs now for most use cases (yes, including Akamai). But realistically, most sites (Pinterest included) are not using clever cache invalidation for dynamic content because it’s not worth the integration effort (and it’s very difficult to abstract for large 1k+ engineering teams). Most customers are just using DSAs for the L4/L5 benefits (both security and perf). In that case, it’s not complicated to implement multi-cdn.
- mrzool 5y agoWhy is this a link to the Fastly homepage, where absolutely no information is provided? This is the page that should be linked: https://status.fastly.com https://status.fastly.com
- scolvin 5y agoBecause even their homepage is down intermittently/for some people.
- lucasverra 5y agoit is starting to show several Degraded Performance tags
- jmvoodoo 5y agoOddly their homepage rendering an error was a more accurate description of the problem than "investigating potential impact to performance with our CDN"
- taurath 5y agoStuff is down across the web, but the most it says is “degraded performance” and in my area it’s all green even though the sites are still down.
- samhh 5y agoAll looks orange now, but "degraded performance" is a cheeky way to describe "everything is on fire".
- kevincox 5y ago> Fastly’s network has built-in redundancies and automatic failover routing to ensure optimal performance and uptime. But when a network issue does arise, we think our customers deserve clear, transparent communication so they can maintain trust in our service and our team. What a joke!
- mrzool 5y ago
- theginger 5y agoGitHub? I had some issues, checked the service status page said no issues, but images were returning a 503. Maybe they host their service status page elsewhere including using fastly.
- theginger 5y agoGitHub now showing partial outage (images on status pages are fixed)
- fagnerbrack 5y agohttps://dashboard.stripe.com/ https://dashboard.stripe.com/ is down https://github.com/ https://github.com/ is defaced
- rvz 5y agoBasically everything is broken. "Centralising Everything" huh
- Haydos585x2 5y agoSuch a huge number of sites. It seems like it's mostly US based sites and Australians are okay. Sending good vibes to whatever poor person is on support right now.
- nineteen999 5y agoI'm in Australia and there are heaps of sites down for me.
- lysp 5y agoAs per report above - most (or all?) of Asia/Pac servers are down. This incident affects: North America (Ashburn (BWI), Ashburn (DCA), Ashburn (IAD)), Europe (Amsterdam (AMS)), and Asia/Pacific (Hong Kong (HKG), Tokyo (TYO), Singapore (QPG)).
- iso1631 5y agoAffects far more than that
- Haydos585x2 5y agoAh, I meant more sites like ABC, 9NOW, SBS, AFL, Foxtel etc rather than accessing US sites from AU.
- paimoe 5y agoIn Perth, reddit is down. So is Blackboard files for uni
- iso1631 5y agohttps://easydns.com/blog/2020/07/20/turns-out-half-the-internet-has-a-single-point-of-failure-called-cloudflare/ https://easydns.com/blog/2020/07/20/turns-out-half-the-inter... The whole idea of the internet was a distributed network impervious to most attacks. The reality is that a single failure can knock out 90% of the services people use.
- fagnerbrack 5y agoThe internet still works, only the websites are returning the wrong response
- abluecloud 5y agoyeah, the internet is working perfectly. if you want to view 503 errors.
- Doxin 5y agoBelieve it or not but "the internet" and "the world wide web" are not synonyms.
- berkes 5y agoTrue. But the vast majority of use goes via "WWW". For example email - the other big "internet-user" is technically not part of the WWW, but most (? I don't have any stats, just a guess) of our mailclients run on the WWW, nonetheless.
- fmajid 5y agoBitTorrent was half of all Internet traffic for a while, though it has decreased with the rise of legal and convenient streaming services.
- berkes 5y agoMost of which (unfortunately) run on the WWW. I'm not sure what the native clients for Netflix and Spotify actually run, but I use their WWW clients mostly. Making most of my internet bits&bytes go over the WWW.
- csomar 5y agoSo I'm wondering where in the "hundreds of servers around the world" did they exactly go wrong. This happened with Cloudflare before too. I think we are a little too dependent on these services.
- fagnerbrack 5y agoIn Software Engineering we call it "coupling" /s
- 0xbkt 5y agoIt is a meaningless premise when you actually have SPoFs baked deep inside the system.
- patentatt 5y agoI’d love to see a breakdown of what single point of failure causes these worldwide network outages. They even brag about redundancy in their marketing materials. I hope we see a post mortem on this
- jfny 5y agoYeah seriously. Time to rebuilt the architecture from the ground up.
- raylus 5y agogithub.com is pretty broken
- heavydust 5y agoreddit.com is affected too
- ClearAndPresent 5y agoWhat conclusions can we draw about concentrating web content in a few CDNs?
- threeseed 5y agoIn HTML/CSS you should be able to specify a fallback source if the first returns a non-200. Or that companies need to have better DNS strategies.
- richardwhiuk 5y agoWeb Browsers should probably retry a different server in DNS if they get a 503 - but they don't.
- allyant 5y ago> In HTML/CSS you should be able to specify a fallback source if the first returns a non-200. Except if the HTML/CSS is hosted on that CDN?
- ilaksh 5y agoContent-centric networking had been a central research topic for many years. And many potentially useful systems have been proposed and implemented. At some point some of them will start to become popular.
- nickelpro 5y agoDNS didn't fail, and there's nothing you can do in HTML/CS/JS if your CDN fails to serve those things
- itsbits 5y agowe had that experience when cloudfare was down for sometime lastyear. We now setup a minor own static server as a backup, if at all this happens again. Althgh we hadn't so far had to use it.
- npteljes 5y agoThat sometimes they fail but the world goes on.
- hypnoscripto 5y agoLooks like fastly.com uses fastly…
- NewLogic 5y agoEven amazon.com styling is borked for me
- deleted 5y ago[deleted]
- mschuster91 5y agoAh yes, the wonders of centralized internet infrastructure. Let's use a handful of providers for everything, they said. It will be cheaper, they said. It will be easier to manage, they said. And it was cheaper, until downtimes began to affect more and more sites when central SPOFs got hit. And I wonder how much of that need for these centralized SPOFs actually comes from the sheer absurd amount of bloat, ads, code and assets that sites these days "have" to deliver to the customer. I 'member times when pages had 100kb total size, loaded in an instant and were perfectly usable.
- perino 5y agoAnything hosted on Firebase seems to be down
- deleted 5y ago[deleted]
- vfclists 5y agoWhat happens when there is excessive centralization. I thought that one of the principles behind the Internet is to be able to reroute around failures, but neither these service providers nor their clients ever seem to learn. I guess in their mind that only applies to packet routing not services. SMH
- 8K832d7tNmiQ 5y agoThat explains why I couldn't access reddit
- dkarp 5y agoBefore the "Error 503 Service Unavailable" messages appeared, there were a few minutes where the error was a single line: connection failure Not sure if that provides anyone here with more insight into what might have caused this!
- SileNce5k 5y agoIt was `connection failure` for me.
- stordoff 5y agoI got that, then a 'Fastly unknown domain' error (on Reddit), then the 503s on multiple sites (I also had an API I use return a 502 then a 500 error, but I don't know what the full response was as it was just a quickly thrown together script I was using). Edit: and now "I/O error" on Reddit.
- q3k 5y agoI also saw a glimpse of 'I/O error'. That sounds fun.
- ramraj07 5y agoFor a moment I thought all of Western internet was cut off from India. Says how siloed my browsing habits are!
- graphman 5y agoFirebase Dynamic Links is affected too. Checking the IP looks like they are using Fastly which is quite surprising.
- simonbarker87 5y agohow will their devs fix it if stackoverflow has gone down?!
- unfunco 5y agoAmazon being down surely points to something other than Fastly being the cause?
- austinjp 5y agoI just had a look at amazon.co.uk and most assets fail to load, the browser debug console is full of 503 errors. Picking one at random, it's fastly: $ nslookup images-eu.ssl-images-amazon.com Server: 127.0.0.53 Address: 127.0.0.53#53 Non-authoritative answer: images-eu.ssl-images-amazon.com canonical name = m.media-amazon.com. m.media-amazon.com canonical name = media.amazon.map.fastly.net. Name: media.amazon.map.fastly.net Address: 199.232.177.16 Name: media.amazon.map.fastly.net Address: 2a04:4e42:1d::272
- mpitt 5y agoAmazon.com uses Fastly https://www.streamingmediablog.com/2020/05/fastly-amazon-homepage.html https://www.streamingmediablog.com/2020/05/fastly-amazon-hom...
- jfny 5y ago[deleted]
- richardwhiuk 5y agoThey will use S3, but they need a CDN in front. Surprised they don't use CloudFront - maybe that's what they've failed over to.
- macintux 5y agoApparently they switched from CloudFront after determining Fastly was faster for this use case. CloudFront is focused on large streaming services, not small HTTP resources.
- raylus 5y agoWhew, DevOps fire alarms are going off!
- lysp 5y agoThis incident affects: Europe (Amsterdam (AMS), Dublin (DUB), Frankfurt (FRA), Frankfurt (HHN), London (LCY)), North America (Ashburn (BWI), Ashburn (DCA), Ashburn (IAD), Ashburn (WDC), Atlanta (FTY), Atlanta (PDK), Boston (BOS), Chicago (ORD), Dallas (DAL), Los Angeles (LAX)), and Asia/Pacific (Hong Kong (HKG), Tokyo (HND), Tokyo (TYO), Singapore (QPG)).
- kiwijamo 5y agoAffecting Auckland (AKL) which is not on the list so I can only assume it's affecting more locations than they're letting on.
- Banana699 5y ago+= North Africa (Egypt, Cairo) Stackoverflow.com, reddit, qoura down. (and probably more, those are the ones I tested)
- tendencydriven 5y agoTheir status page is now saying every location has degraded performance.
- pimterry 5y agoThis is one of the things that excites me about IPFS: in a world of decentralized data storage, yes self-hosting and control over your data is nice and all, but serious resilience to most random infrastructure outages is a much bigger deal. It's still early days, but I'm hopeful that it can provide a real solution to today's CDN centralization.
- deleted 5y ago[deleted]
- jfny 5y agoI'm pretty sure you can serve hundreds if not thousands of users from a single Raspberry Pi
- pimterry 5y agoI mean, yes, absolutely, and that works to start with, but I'm willing to bet the overall uptime and performance of a raspberry pi in your living room is quite a bit worse that Fastly's :-).
- jokoon 5y agoAgree, but currently, ipfs would serve as a fallback, since it's about files. Decentralized/distributed generally has slower network performance. Unless most nodes are high performance, I guess? Personally I think a distributed database system, where entries are being made redundant in something like a blockchain+dht, would be a good start? Decentralizing the internet works if it financially makes sense for platforms to build such tools.
- pimterry 5y ago> Agree, but currently, ipfs would serve as a fallback, since it's about files. Isn't a CDN fundamentally all about files too? > Decentralized/distributed generally has slower network performance. Unless most nodes are high performance, I guess? There is definitely more work to do here before this is really useful, but it's well within the realm of things that IPFS should be able to do at reasonable performance for production sites in future. Good performance still requires a serious CDN node network similar to traditional CDNs today (to seed your content for day to day use) but with IPFS if that CDN goes down then existing users on your site can _also_ serve the site to other nearby users directly, or other CDNs can serve your site too, etc etc. Your DNS wouldn't be linked to any specific CDN in any way, just to the hash of the content itself, so anybody could serve it. > Decentralizing the internet works if it financially makes sense for platforms to build such tools. There's a platform company called Fleek who already do this today: https://fleek.co/hosting/ https://fleek.co/hosting/ (no affiliation, and I've never even used the product, just looks cool). Seems to be designed as a Netlify competitor: push code with git and it builds it into static content and then deploys to IPFS. The benefits don't exist today of course, because no browsers natively support IPFS, so most users can only access the content via an IPFS gateway, which means you're back to fully centralized server infrastructure again... If we can get IPFS support into browsers though then fully decentralized CDN infrastructure for the web is totally possible.
- devops000 5y agoHeroku is down https://dashboard.heroku.com/ https://dashboard.heroku.com/
- raphaelj 5y agoOnly their main website though. My Heroku apps work pretty well.
- monkeydust 5y agoPretty bad www.gov.uk is down as more services move to digital.
- zelphirkalt 5y agoI don't think moving to digital is the issue here. The issue is relying on third parties, which can have an issue at any moment, taking down whoever relies on them with them. A government should not rely on CDNs like that. In fact government websites should not have any traffic going over third parties. When I want to use/view a government website, I should not be subjected to sharing any data with unwanted third parties and the government should not be affected, when some private company makes mistakes or has outages. It is an unacceptable situation. They can set up their own state-owned CDN, using the same underlying technology. Compared to where they spend all that tax money, some servers and some engineers would be a very cheap investment, in relation to the independence achieved.
- allyant 5y agoThey seem to have migrated across to Cloudfront - working now.
- ur-whale 5y agoWow, talk about a brutal SPOF, most of the things I had planned to work with today are broken: reddit, github, stack overflow.
- plasma 5y agoI briefly saw an output error about "domain not found" when hitting fastly.com, wonder if some list of domains has hit a limit/flushed/etc.
- dkarp 5y agoI get this now on reddit: Fastly error: unknown domain: www.reddit.com.
- raphaelj 5y agoCouldn't be happier I moved https://noisycamp.com https://noisycamp.com to BunnyCDN.com.
- aero-glide2 5y agoisitdownrightnow.com is down
- roachpepe 5y agoThanks for the best laughs in a while friend - that's pure irony right there!
- austinjp 5y agoYeah so it's been mentioned in the comments already, but to everyone in Fastly right now: I feel for you. Something like this must be insanely stressful, and not just during the outage. There will be (should be) a massive post-mortem. People will be losing sleep over this for days, weeks, months. :( Edit: There seems to be a major empathy outage in this thread. Disgusted but not surprised, unfortunately.
- mothsonasloth 5y agoCall me old fashioned but the latest trend of showing "empathy" for a serious incident, then proceeding to dance around the aftermath of it, whilst people give themselves a pat on back in a retro/post-mortem, isn't the way to do it. People need to be blamed, and responsibility for actions taken (without covering asses)
- q3k 5y agoThe point isn't to dance around the incident, but to not blame people. You can blame systems, design, engineering culture, processes, but don't blame people. Even if someone accidentally pressed the 'destroy prod' button, that's not the fault of that person, it's the fault of that button existing and being accessible in the first place. I have no empathy for Fastly-the-company. I hate the fact that the Internet is centralized around CDNs. I wish this idea of 'but we _must_ run a CDN for our 1QPM blog!' would die in a fire. But I can still empathize with the Fastly engineers handling this shitstorm right now.
- tyrex2017 5y agoI disagree. People implemented those systems, so if you are correct that it is the systems fault, then it is also a persons fault. People must be held accountable to have good incentives to reduce such outtages in the future. I do agree though that we should always be compassionate and realistic with other humans.
- taurath 5y ago
- Metacelsus 5y agoI first noticed that xkcd was down. Then I went to post about it on reddit . . . also down! Good thing HN is up.
- ur-whale 5y agoLooks like an SRE team rolled out buggy software.
- ur-whale 5y agoLooks like HN is working ;-)
- timvisee 5y agoThis seems to be a bigger issue. BGP failure?
- stefan_ 5y agoIf they can serve me a garbage Varnish error (shoutout to "software that actually runs your business that none of your devs work on") it's not BGP.
- devops000 5y agoHacker News is the only one UP!
- mothershesha 5y agoGot the same here (Australia)
- montag 5y agoIt's funny, I searched Twitter for "Ebay down" and the top result was an Ebay tweet with some not coincidentally broken Twitter emoji SVGs (as another person mentioned)...
- taurath 5y agoI’ve noticed lots of social media content is tied to this - Reddit and Twitter images and some videos, for one.
- rich_sasha 5y agowww.python.org down as well, with the shortest of messages: 'connection failure'. Probably related?
- rich_sasha 5y ago...and now back up, with reddit et al still down. Hmm.
- lopatin 5y agoTheir status page keeps claiming that my region, Chicago (ORD), is either Degraded Performance, or Operational. But clearly it's down. Is fuzzing metrics like this how they hit their SLA targets?
- dragosbulugean 5y agoAll Webflow sites?
- dragosbulugean 5y agoAnd all Webflow sites it seems...
- fsnowdin 5y agojust had my own site down because of this. glad to see it wasn't my fault lol but good luck to the Fastly people on fixing the issue.
- optiomal_isgood 5y agoAmazon.com was completely broken here (Europe) and they're back, I was observing from where the assets were loaded from and they switched from EU to NA as a failover. Homework well done.
- abluecloud 5y agoStill getting broken assets from the UK.
- optiomal_isgood 5y agoYou're right, I should've said *partially* back. At least the CSSs now load, but a few products images are still gone. However it was completely broken here before (literally loading just the main HTML).
- 00deadbeef 5y agoI was surprised to learn Amazon don't use their own CDN
- optiomal_isgood 5y agoThey used to use AWS CloudFront and switched to Fastly, someone shared this in another comment: [https://www.streamingmediablog.com/2020/05/fastly-amazon-homepage.html](title https://www.streamingmediablog.com/2020/05/fastly-amazon-hom...: CDN Fastly Wins Content Delivery Business For Amazon.com and IMDB Websites) Quoting: > "But with small object delivery, like images loading fast on Amazon’s home page, it’s the opposite. Customers will pay for a better level of performance and in this case, Fastly clearly outperformed Amazon’s own CDN CloudFront. This isn’t too surprising since CloudFront’s strength isn’t web performance, or even live streaming, but rather on-demand delivery of video and downloads."
- dastbe 5y agoAmazon (like a lot of others) use several CDNs for redundancy. You can see from dig that it resolves to combinations of cloudfront, akamai, and (presumably, based on your reported experience) fastly. dig +short www.amazon.com tp.47cf2c8c9-frontier.amazon.com. d3ag4hukkh62yn.cloudfront.net. 65.8.70.16 dig +short www.amazon.co.uk tp.bfbdc3ca1-frontier.amazon.co.uk. dmv2chczz9u6u.cloudfront.net. 13.224.0.89 dig +short www.amazon.in tp.c95e7e602-frontier.amazon.in. d1elgm1ww0d6wo.cloudfront.net. 13.224.9.30 dig +short www.amazon.co.jp tp.4d5ad1d2b-frontier.amazon.co.jp. www.amazon.co.jp.edgekey.net. e15312.a.akamaiedge.net. 104.71.134.162
- easytiger 5y agoI will NEVER understand why people put so much trust in single provider solutions for anything critical.
- sergiomattei 5y agoYikes, seems like a massive outage. EDIT: Hexdocs is down, elixir-lang.org is down
- atymic 5y agoThis has got to be even bigger than when cloudflare went offline, in terms of big companies affected. Clearly they have way more F500 customers than CF. Good luck to the on call engineers!
- yxhuvud 5y agoThe funny part is that it isn't uncommon for sites to depend on both cloudflare and fastly in one way or another, due to buying services from saas companies that also depend on them.
- gansai 5y agowouldn't websites have alternate CDN's managing their traffic, why should they have a single point of failure ? I was assuming there are couple of services like Fastly and companies might have architected keeping in mind the alternatives too, I guess.
- raimondious 5y agoNormally you configure your a record to point at the cdn as the cdn is the thing that gives you multiple points of failure (caches all over the world). Hard to have a fallback to that. Running multiple cdns would be extremely expensive. Cdn caches are kept useful by traffic running through them, so hard to have a backup for that too.
- ImpactStrafe 5y agoBecause interacting and switching between cdns can be very complicated and/or costly It should be planned for, especially by major tech organizations like reddit, or Amazon, etc. But I won't fault news organizations, who already don't have boatloads of money for not having fail over cdns
- tommoor 5y agoHands up if you're also here after being woken up by downtime alerts on the west coast
- permb 5y agoMade my alpine linux docker builds fail as well (varnish) - but shouldn’t it use a mirror when the primary download site is gone? fetch http://dl-cdn.alpinelinux.org/alpine/v3.12/main/x86_64/APKINDEX.tar.gz http://dl-cdn.alpinelinux.org/alpine/v3.12/main/x86_64/APKIN... fetch http://dl-cdn.alpinelinux.org/alpine/v3.12/community/x86_64/APKINDEX.tar.gz http://dl-cdn.alpinelinux.org/alpine/v3.12/community/x86_64/... ERROR: http://dl-cdn.alpinelinux.org/alpine/v3.12/main http://dl-cdn.alpinelinux.org/alpine/v3.12/main: temporary error (try again later)
- DoreenMichele 5y agoI'm having intermittent Reddit issues, as one more data point. I'm grateful for HN. I rebooted my computer. I thought it was my device and then saw this on my phone while rebooting.
- clawphantom 5y agoTwitch isn’t working and not responding and also the web dashboard
- devops000 5y agoBTC/USD is down too.
- creamyhorror 5y agoPerfect time for the crypto whales to dump massively and cause an absolute panic.
- MrGilbert 5y agoInterestingly, https://www.fastly.com/ https://www.fastly.com/ works for me, whereas https://fastly.com/ https://fastly.com/ doesn't.
- gislifb 5y agoFunnily enough, it's the opposite for me...
- clawphantom 5y agoTwitch isn’t responding and also the web dashboard
- monkeydust 5y agoJust occurring to me how CDNs are a major point of failure now for the internet
- choult 5y agoMy money is on an expired internal certificate or CA.
- hugh-avherald 5y agoFastly has scheduled maintenance to retire some TLS certs next week.
- oneeyedpigeon 5y agoGood marketing for Fastly! I had no idea so much of the internet relied on it...
- Omnious58 5y agoI was wondering why my Tidal app just stopped mid song and won't connect, after much googling and absolutely no help or even notifications from Tidal explaining there's an issue it seems this outage is the culprit. Bugger.
- evouga 5y agoSince Fastly’s own website is currently down: What is fastly? Why are a huge number of web sites dependent on them? They are some kind of web host for companies that don’t want to run their own servers/data centers?
- ImpactStrafe 5y agoFastly is a Content Distribution Network (CDN). Basically the closer the server serving the webpage is to the end user the faster it is for the end user to see and interact with. But running servers all over the world 1) isn't efficient 2) costs a lot of money. So a few companies (fastly, cloud flare, akamai) figured, hey, why don't we build a bunch of small data centers all over the world and then provide a distributed way to serve web traffic from it. It originally was brought about for services like Netflix, but has expanded greatly. You still host your servers, but a copy of the webpage/media is given to the CDN to serve to customers.
- evouga 5y agoThanks. That makes sense. Wouldn’t you build in a failsafe that bypasses Fastly and sends traffic to your own servers in the case of this kind of outage? Or outages are so rare that it’s not worth the trouble?
- ceejayoz 5y agoMany sites do this; Amazon's failed over to their own servers for images for me, it appears. It typically just takes some human intervention, I suspect.
- npteljes 5y agoThat's the fallback, but the original stack is not designed with the volume of traffic in mind. So it gets overwhelmed very quickly and makes the website practically unavailable.
- ImpactStrafe 5y agoThe number of serious CDN outages in the world are incredibly rare. In fact, you can probably remember most of them if you were given dates. Plus, going around the CDN can be very complex (depending on the type of content), very expensive (all of a sudden you have a massive data out network traffic that didn't exist previously), and not guaranteed to work (DNS updates can take longer to get to everyone than the actual CDN outage lasts). There are places where it is worth it and useful, but for a lot of the sites listed it's not useful.
- jchandra 5y agohttps://www.greenhouse.io/ https://www.greenhouse.io/ down as well.
- vincentmarle 5y agoWell I know where to go next time if I were to be a Russian hacker
- ilaksh 5y agoLet's make all of the main internet sites dependent upon one central private service. Great idea guys.
- ysavir 5y agoTangential question, but with services like these, is there a known way to handle failure gracefully? Some way to automatically bypass these services if they are known to be down?
- richardwhiuk 5y agoHave two different CDN partners, own your own DNS, and then withdraw one of the CDNs if they are down. Suspect that's what Amazon have done.
- efficax 5y agoYou have to have two separate cdns and use DNS to fail over. The problem is that means paying for a CDN that just sits dormant for the 99.999% of the time that your primary is down. Alternatively you could use DNS to fail over to the content you host, instead of another CDN. But in many cases that would be the same as an outage since the CDN exists to reduce the impact of all those requests on your infra
- deleted 5y ago[deleted]
- ikjus1 5y agoomg! How can i code from now!! T^T
- ZoomStop 5y agoThe outage has already been added to the Fastly Wikipedia page
- abhiminator 5y agoHoly smokes these Wikipedia writers are quick! I'm sometimes impressed by how fast a page on a super recent happening gets populated with all of the currently known details.
- ddtaylor 5y agoSomeone must have 51% attack the Pied Piper blockchain!
- willvarfar 5y agohttps://www.bbc.com/news/technology-57399628 https://www.bbc.com/news/technology-57399628 is rendering and reporting on the story, but BBC itself was down at the start of the outage, with the same 503 varnish error message. Presumably the BBC has some kind of fallback in place. The journalists ought interview their own techies :)
- deleted 5y ago[deleted]
- snookdebook 5y agoI gave it about 10 tries, and it seems a very small percentage of transactions do go through. A decent number of tries is rejected right at the Varnish front door: < HTTP/2 503 < server: Varnish < retry-after: 0 < date: Tue, 08 Jun 2021 10:11:41 GMT < x-varnish: 271470009 < via: 1.1 varnish < fastly-debug-path: (D cache-bma1666-BMA 1623147101) < fastly-debug-ttl: (M cache-bma1666-BMA - - -) < content-length: 450 < Service Unavailable Guru Mediation: Details: cache-bma1666-BMA 1623147101 271470009 Many more reach some backend system that just dumps "connection failure": < HTTP/2 502 < content-type: text/plain; charset=utf-8 < content-length: 18 < connection failure And a tiny few do get through: < HTTP/2 200 < content-type: text/html; charset=UTF-8 < cache-control: max-age=0, must-revalidate < date: Tue, 08 Jun 2021 10:11:43 GMT < via: 1.1 varnish < vary: accept-encoding < set-cookie: ...snip... < server: snooserv < content-length: 275036 < <!doctype html><html>...snip...
- MyOnePiece 5y agoQuick question if the cdns are down why cant traffic be routed to the web servers the central web servers the company owns ? I thought cdns had fallback configured ?
- deleted 5y ago[deleted]
- angled 5y agoNone of the ES/NQ/RTY/YM futures contracts took kindly to the outage! This could have had a much wider financial impact. Most seem to have recovered now.
- _kyran 5y agoThings seem to have come back online in Australia, although not sure if that's just sites switching over their DNS?
- john37386 5y agoIt should be resolve soon. From fastly status page: The issue has been identified and a fix is being implemented. Posted 1 minute ago. Jun 08, 2021 - 10:44 UTC
- abluecloud 5y agoWonder if all the caches will have been wiped, causing knock on issues
- john37386 5y agoYou might be right. Here is another update from fastly: The issue has been identified and a fix has been applied. Customers may experience increased origin load as global services return. Let's see
- grumple 5y agoPhew! That time to find the issue is always the stressful part. < 1 hour is pretty good for weird stuff, and fortunately the east coast of the US is barely online this early (sorry Europe!).
- taosx 5y agoI̶n̶ ̶r̶o̶m̶a̶n̶i̶a̶ ̶e̶v̶e̶r̶y̶t̶h̶i̶n̶g̶ ̶s̶e̶e̶m̶s̶ ̶b̶a̶c̶k̶ ̶t̶o̶ ̶n̶o̶r̶m̶a̶l̶.̶.̶.̶?̶ Edit: nope, just worked for 2-3 requests (10 secs)
- abhiminator 5y agoLooks like they're currently applying a fix. https://status.fastly.com/incidents/vpk0ssybt3bj https://status.fastly.com/incidents/vpk0ssybt3bj
- deleted 5y ago[deleted]
- _kyran 5y agoThose of you that work in DevOps, SRE or are CTOs. What kind of things do you put in place to manage these kind of centralised issues that are beyond your control?
- timthorn 5y agoThese issues are in your control - not for the centralised service but your use of them. You can build appropriate redundancy for the components/providers in your stack and the budget you have.
- TheRealDunkirk 5y agoEvery other comment about what's down in this thread -- as if we needed dozens of site-by-site accountings of this outage in the first place -- is a bitch about reddit. Why is reddit so important to this crowd? The specific topics I used to read the site for (half a dozen years ago) have all been overrun by "bucket people," there is literally never an answer to any question I find a google link to there, and the site's design is actively user-hostile. Seriously: what's keeping that place afloat? Porn, I suppose.
- sergiomattei 5y agoOf course, the Enlightened Folk of this site can no longer use their leisure time on lowly activities such as the "Reddit". Teach me your ways, master! /s Jokes aside, people can do whatever they please. Reddit has a bunch of niche communities around many hobbies and fun things. No need to be bitter about it.
- TheRealDunkirk 5y agoYou have put your finger on it. I AM bitter about it. It used to be really cool, and really nice to use, before the Taylor/Pao dustup, and the redesign.
- afroboy 5y agoold.reddit still a thing and there is a plenty of educational subreddits with really nice community around them, it's just like the internet just pick the things that suits you.
- modshatereality 5y agoreddit taught me to never trust a mod, so it does have some purpose still. i think without glaringly bad examples of how (not) to run a community based site, we would be doomed to repeat it's mistakes.
- loriverkutya 5y agoThe issue has been identified and a fix is being implemented. Posted 3 minutes ago. Jun 08, 2021 - 10:44 UTC
- classicflavour 5y agoMy work's website is down too and the regular sites I use to escape work borderm
- jfny 5y agoDo companies really not run test suites / do manual testing before deploying to production?
- LightG 5y ago"The internet will just route around a local / centralised problem ... like water around an object" Obligatory LOL ...
- heavydust 5y agothe problem has been fixed
- luke2m 5y agoWhen this happens to cloudflare, it will be even more impactful.
- gansai 5y agoFastly is back now. (The issue has been identified and a fix is being implemented.)
- deleted 5y ago[deleted]
- zwirbl 5y agoSpotify is also hit, though it still works without images
- k_ 5y agoUpdate: The issue has been identified and a fix is being implemented. Posted Jun 08, 2021 - 10:44 UTC Seems like this is being resolved; curious to see the details afterwards (from https://status.fastly.com/incidents/vpk0ssybt3bj https://status.fastly.com/incidents/vpk0ssybt3bj)
- optiomal_isgood 5y agoReddit, Stack Overflow, Spotify, all back for me. Good job Fastly engineers!
- jujodi 5y agoWould be fascinating if Fastly is not be able to use GitHub, Travis, Terraform, pip, etc. to deploy their fix
- nraval1729 5y agoInteresting thought. I had not thought about this before. If there is a cyclic dependency (not saying there is at the moment) how would things play out? Do you just ssh into your own servers to deploy the fix?
- john37386 5y agoIt's probably a DDoS attack.
- anotheryou 5y agoLooks fixed: https://downdetector.com/ https://downdetector.com/
- pattyj 5y agoIt would be interesting to see estimations on the man-hour cost of this outage.
- zonezone23 5y agoDown https://yarracity.vic.gov.au https://yarracity.vic.gov.au
- timetosleep 5y agoSeems to be back online
- JosephK 5y agoExtremely long call, but what are the chances this turns out connected to the raids on organised crime using the An0m app that started today?
- cwen 5y agoA real-world Chaos experiment!
- toong 5y agoIt is time to remove that "100% uptime guarantee" claim from the website :grimacing:
- fullstackwife 5y agoNo mention of outage on https://status.cloud.google.com/ https://status.cloud.google.com/, and I wonder why, because apparently this is a GCP problem.
- colesantiago 5y agoLooks like Fastly did not work as advertised, very misleading.
- artembugara 5y agoSeems like another single point of failure. What is a solution to not be affected by such an outage?
- diveanon 5y agoTime to develop CDN for CDNs. It seems like a pattern that CDN have overly centralized the web and lead to issues like this. Maybe its time to build a CDN that distributes your static assets to multiple CDNs and has a set of fallback states for service outtages.
- JCWasmx86 5y ago>The issue has been identified and a fix has been applied. Customers may experience increased origin load as global services return. Is fixed
- Dobbs 5y agoI got a push notification from the CNN app telling me a bunch of the internet was down due to a cloud provider. I clicked the link only for the app to open to a 503. In hindsight not surprising, but quite amusing.
- kypro 5y agoSome people are claiming online that this is a cyber attack. I contract for the UK Gov and I'm hearing reports that traffic is going through the roof right now. Anyone know if there is any legitimacy to this?
- fr2null 5y agoThe fastly monitoring/status page says: "Customers may experience increased origin load as global services return". Which sounds like the increased traffic is to be expected. [1] status.fastly.com
- cph-w 5y agoI did not realise fastly adoption was so wide-spread. Can anyone more enlightened tell my why or have some resource on which use-cases fastly is superior to other CDNs such as CloudFlare?
- grumple 5y agoGood job Fastly for getting the issue identified and resolved so quickly. < 1 hour to identify, <13 minutes to fix (assuming status is accurate).
- cdev 5y agoit seems to be up now
- colesantiago 5y agoAlso, why has this been allowed to happen? Billions of dollars lost because of this one company? I don't understand this.
- tinaranw123 5y agoI HAVE FLUTTER EXAMINATION TOMORROW BUT STACKOVERFLOW IS DOWN. SHIT :)
- tinaranw123 5y agoI HAVE FLUTTER PROJECT FOR MY EXAM TOM AND STACKOVERFLOW IS DOWN. SHIT :) I'M FCKED
- reuben_scratton 5y agoI'm sure it's just a coincidence that today is Patch Tuesday. :-|
- rottc0dd 5y agogithub is back online. SSO too.
- alexchamberlain 5y agoStupid question: why didn't sites "just" fail over to their actual servers to handle the traffic, albeit slowly? I guess they won't be sized to handle the load in a lot of cases, and Fastly was responding, so DNS fail over didn't work?
- abluecloud 5y agoyeah. the dns was up. the problem was the servers weren't able to proxy the traffic. Also, as you say, you'll probably end up bringing down the upstream servers if you just fail open (and not even sure that'd be a possibility with fastly in it's "down" state that we saw).
- altacc 5y agoProbably a different answer for each site. I'm not a DNS expert but I think you're right on both counts. Having failover also requires a duplicate CDN architecture at the fallback location, which is an increase of costs in time, money & maintenance for relatively little benefit. Often there's a fair amount of background integration with a CDN, and each function slightly differently, so it's not simply plug & play.
- osxman 5y agoSTOP Free VPS servers. Now! Explanation: they provide an easy platform for attackers.
- hestefisk 5y agoThe Guardian summarised this as well: https://www.theguardian.com/technology/2021/jun/08/massive-internet-outage-hits-websites-including-amazon-govuk-and-guardian-fastly https://www.theguardian.com/technology/2021/jun/08/massive-i...
- vlan121 5y agoDamn, I thought I cloud blame myself or the provider..
- vlan121 5y agoDamn, I thought I cloud blame myself or the internet provider...
- omk 5y agoThis outage made me realize that github is served over a single IP address (A record) for my point of origin (India). Stackoverflow has 4 A record listing, but all of these belong to fastly. The internet is designed for redundancy. Wonder why these companies don't have a fail over network. Makes me wonder if cost is factor considering their already massive infra. But a single point of failure ... <confused>.
- kayfox 5y agoGithub's DNS likely will serve up a different IP for github when there is an outage. I can't talk about the details but GitHub and the rest of Microsoft use a global load balancing system that works through DNS.
- omk 5y agoWould be interesting to know what these fail over patterns are. As DNS takes a while to propagate, I thought DNS records already indicate fail over addresses.
- kayfox 5y agoI think only MX records indicate any priority for each additional record returned, for A records theres no indication of which records have priority over others and the usual behavior of authoritative DNS servers is to rotate the order in which records for the same thing are returned, so effectively returning more than one record for the same question results in a distribution of requests to the IPs returned rather than any sort of failover behavior. In the case of the software Microsoft uses, it monitors endpoints for the websites in question and then changes which IP(s) are returned based on the availability of those endpoints, the geographic region and other factors.
- bombcar 5y agoSome reliability systems change the routing for the IPs instead of updating the DNS as BGP can propagate faster than DNS caching. Priority for A records would a nice feature.
- i386 5y agoAnyone want to talk about half the internet going out because one provider couldn’t keep their service up instead of SO jokes and feels for the engineers? the entire internet is like a stack of cards from the protocol to the economic model.
- fareesh 5y agoHow does one design a system that has a redundancy for when the CDN goes down? Paying for more than one CDN is probably too expensive isn't it?
- modshatereality 5y agoThis post is suspiciously ranked much lower than it should be (1216 points, 9 hours ago), lower than posts with < 100 points.
- marmot777 5y agoI think the honorable thing would be for them to have a statement easily findable. So many companies sweep this sort of things under the rug if it’s only customer data that’s been breached. If they can’t sweep they have a high priced PR agency do the communicating. I do not trust companies who handle things this way.