25 ms·
Tell HN: AWS appears to be down again
Anyone else seeing this?
- Zelphyr 5y agoI think it's time to face the fact that we all have too many of our eggs in the AWS basket.
- imhoguy 5y agoI guess it is all about log4shell patching in rush.
- ceejayoz 5y agoWe're having troubles in us-west-2. Discourse is reporting trouble, too. https://twitter.com/DiscourseStatus/status/1471140369899290628 https://twitter.com/DiscourseStatus/status/14711403698992906...
- supermathie 5y agous-west-1 also seems offline, but us-east-1 (ironically) seems fine
- nic_wilson 5y agoWe are seeing issues with requests to Auth0, which I believe is hosted on AWS and has historically gone down when AWS has had issues
- romanhotsiy 5y agoWe see issues with Auth0 too. Other AWS services we use seem to be working fine so far (us-east-1)
- heartbreak 5y agoAWS is reporting an issue in us-west-2 on their status page.
- deleted 5y ago[deleted]
- ramesh31 5y agoAuth0 went down for us as well right when AWS did. At least it's not like those two systems run our entire company...
- rwalk 5y agoYup, trouble in us-west-2 for us.
- rychco 5y agoTsheets is also down so I can’t clock my hours LOL
- aswinmohanme 5y agoCouldn't access Notion, so came to check HN, and boom here is the answer.
- alvis 5y agoOh man, not again!
- alvis 5y agoOh man. Not again!
- ramesh31 5y agoAuth0 down as well, right at the same time. There goes any sort of productivity today. Whole company in firefighting mode.
- oriettaxx 5y agoAWS appears to be expensive again
- deleted 5y ago[deleted]
- devin 5y agoCan someone please update the title to be broader than AWS?
- dhruvarora013 5y agoLooks like its taken down SendGrid, NPM, Twitch, Auth0 so far
- 8K832d7tNmiQ 5y agoTwitch video streaming is also down right now: HTTP Error 500 internal server error
- justinc8687 5y agous-west-2 stuff is down for me too
- Sholmesy 5y agoYup, seeing this on us-west-1
- 1123581321 5y agoYes, all our stuff in west-2 went down at 7:15 PT.
- thadjo 5y agoobligatory comment about status page showing seas of green: https://status.aws.amazon.com https://status.aws.amazon.com
- rpadovani 5y agoFor me, it's down as well
- deleted 5y ago[deleted]
- tubignaaso 5y agoThe status page appears to be down now as well.
- thadjo 5y agoyep I'm seeing that too - wow.
- BoiledCabbage 5y agoMaybe they got so much flak last time for it being worthless, that they just decided to pull it this time??
- gregmfoster 5y agoLove to see the manually updated status page not updating
- arsome 5y agoFor this kind of thing it's usually better to just use a user-driven site like: https://downdetector.ca/status/aws-amazon-web-services/ https://downdetector.ca/status/aws-amazon-web-services/ Some users are clueless, but the clueless users average out over time and the spikes make it clear when there are actual issues.
- TheFragenTaken 5y agoAt least Twitch.tv (Amazon subsidiary) and npmjs.com seems to be affected.
- AustinDev 5y agoYeah, I'm getting 2000 player errors in the Twitch video player.
- gregmfoster 5y agoDown for us (graphite.dev) as well, running on us-west-2
- tubignaaso 5y agoSeeing this on us-west-1. us-east-1 appears to be functioning for us.
- CodinM 5y agoI fucking swear to God.
- kp195_ 5y agoWe're having issues connecting to our EC2 bastions and accessing the us-west-1 dashboard too EDIT: Cognito auth seems down for us too EDIT2: our ALBs are timing out as well EDIT3: us-west-1 looks like working now!
- tmvnty 5y agoSome npmjs.com pages are returning 503 Service Unavailable for us
- alecr95 5y agoYep, we're also having issues. Hosted on us-west-2
- rpadovani 5y agoSystems manager in eu-central-1 is giving us some issues now, but I am not sure about their internal architecture for it, so maybe needs some us resources?
- robthebrew 5y agohttps://nolandda.org/images/memes/nuke_from_orbit.gif https://nolandda.org/images/memes/nuke_from_orbit.gif
- nickjj 5y agoI'm seeing outages on us-west-2 too. Customer facing traffic being served through Route53 -> ALB -> EC2 is down and CLI tools are failing to connect to AWS too.
- sheepdog 5y agoI can't log on to the console for us-east-1. But our api gateway seems to be working, so I guess production is still up...for now...
- iamricks 5y agoHow much do you guys think these frequent outages will effect their market share in cloud products? Is this enough of a push for organizations to actually move over their infrastructure to other providers?
- ceejayoz 5y agoNot at all. The other cloud providers have had their own outages.
- bravetraveler 5y agoSadly this, people are entrenched with AWS and the... "We're not the only ones down" thing truly has some effect Organizations can more easily swallow an AWS failure when they aren't the only ones hit. They move elsewhere, those outages look more unique Folks may think multi cloud is a good idea... But you're just as likely to suffer from the extra points of failure as you are to benefit
- tyingq 5y agoMulti-cloud is such an odd idea to me. You're either building abstractions on top of things like cloud-provider specific implementations of CDNs, K8S, S3, Postgres, etc...or using the cloud just for VMs. The latter would be cheaper with just old-school hosting from Equinix, Rackspace, etc. The former feels like a losing battle.
- pm90 5y agoIt’s prompted discussions of building multi regional services in my org but not multi cloud. They would have to really really really screw up for that to happen… maybe be down for like a week or something.
- prakashqwerty 5y agoleetcode.com is also down
- deleted 5y ago[deleted]
- jrs235 5y agoI checked their health status page. All is good. /s https://downdetector.com/status/aws-amazon-web-services/ https://downdetector.com/status/aws-amazon-web-services/
- sgt 5y agoOk, so it can't be down then. This is proof!
- tyingq 5y agoThey did add an update, faster than last time: "7:42 AM PST We are investigating Internet connectivity issues to the US-WEST-2 Region." https://status.aws.amazon.com/ https://status.aws.amazon.com/ Edit: They added US-WEST-1: "7:52 AM PST We are investigating Internet connectivity issues to the US-WEST-1 Region." Edit: Found root case, maybe? "8:01 AM PST We have identified the root cause of the Internet connectivity to the US-WEST-1 Region and have taken steps to restore connectivity. We have seen some improvement to Internet connectivity in the last few minutes but continue to work towards full recovery." "8:01 AM PST We have identified the root cause of the Internet connectivity to the US-WEST-2 Region and have taken steps to restore connectivity. We have seen some improvement to Internet connectivity in the last few minutes but continue to work towards full recovery."
- chasd00 5y agosomeone tripped over the fiber run i bet. Or, a cleaning person unplugged a router to plugin a vacuum (that actually happened but to a minicomputer iirc)
- niks2112 5y agowe are having issue with us-west-1 and us-west-2
- wenbin 5y agoListenNotes.com has servers running on us-west-2. One issue is that outbound requests from our servers us-west-2 timeout. Other than that, it seems that we are running ok so far.
- lukeqsee 5y agoWe lost all public IPv6 in the Linode Newark DC. This appears to be cross-provider. Edit: We have IPv6 back.
- theverything 5y agoSlack seems to be having issues too.
- mrsuprawsm 5y agoSeems like this is affecting Dropbox paper, at least for me.
- deleted 5y ago[deleted]
- FoLeyy 5y agonpmjs returning 503
- zonkd1234 5y agoyes. Having issues as of few mins ago reaching us-west-2 ec2.
- TekMol 5y agoIt is surprising that their status page is down too: https://status.aws.amazon.com https://status.aws.amazon.com Their CDN, CloudFront, always works reliable for me. Couldn't they put the status page on CloudFront?
- deleted 5y ago[deleted]
- electroly 5y agoThe status page is working great for me. Did they make it multi-region after the last failure? I'm on the east coast.
- Sebb767 5y agoCentral EU here, appears to be down.
- Hamuko 5y agoNorthern EU, down as well. AWS Management Console in eu-west-1 opens up just fine though. Edit: Hitting refresh a bunch finally got it open.
- oneplane 5y agoWestern EU here, appears to be up for me. Maybe a peering issue?
- Sebb767 5y agoIt's back up for me, too, right now. Rather slow, though, and traceroute shows 25 hops. So it might really be peering.
- hericium 5y agoWorks for me. It's the usual static page with everything green.
- QuiiBz 5y agoI tried to monitor services status using https://stop.lying.cloud https://stop.lying.cloud, but they are also hosted to AWS, and down too.
- civilized 5y agoIf they're monitoring AWS downtime they might want to rethink this.
- cinntaile 5y agoHow come? It's accurate.
- johnisgood 5y agoTrue, if it is down, then that means AWS is down (not necessarily, obviously). :D But honestly, if they want to monitor AWS, they gotta pick something else for this reason, something that is not down when AWS is.
- jeanlucas 5y agoWell... Yes. Hahahah
- nexuist 5y agoWork smarter, not harder
- civilized 5y agoI guess it depends on whether you like your FALSE's encoded as timeouts :)
- deleted 5y ago[deleted]
- mishftw 5y agoFunny I didn't know that and assumed it was okay
- wirelesspotat 5y agoWe're seeing AWS issues with us-west-2 at [medium-sized tech company]
- deleted 5y ago[deleted]
- Jamie9912 5y agoTwitch seems to have recovered, is it back now for everyone?
- bloaf 5y agoStill getting errors in Houston edit: some streams back up, chat still buggy as of 09:55 local time edit2: appears to be back ~10:00 local time
- deleted 5y ago[deleted]
- yabones 5y agoYep, it's broken again. I was trying to install some Thunderbird extensions, and stuff started breaking halfway through. Never thought of an AWS outage borking my mail client I guess...
- dijit 5y agoIs it AWS or could it be an ISP? AWS seems to be working for me, but I’ve worked with clients in the US and spectrum internet tended to drop connections to us sporadically, which looks like an outage to our clients but is something we obviously can’t control.
- albatross13 5y agoCurrently we're seeing 40kms response times from CloudFront distributions, we can't hit PagerDuty (probably runs on AWS), etc. I guess it could be an ISP thing but I guess we're all assuming 80/20.
- avs733 5y agoI wonder if you really dug into most company's tech stacks, how many of their support tools (e.g., PagerDuty) are reliant on overlapping cloud providers.
- simcop2387 5y agoYea a number of people got hit by that, Louis Rossmann found out that every form of contact to his buisness was reliant on AWS east 1. https://www.youtube.com/watch?v=DE05jXUZ-FY https://www.youtube.com/watch?v=DE05jXUZ-FY
- albatross13 5y agoOh man, it is insane. During the aws incident last week we couldn't build software because bitbucket pipelines were all down, due to them running lambdas in us-east-1 only haha. We've taken a massive turn away from a "decentralized" internet.
- avs733 5y agoit's still decentralized...it's just a centralized version of it right? just like Cavendish bananas are grown in multiple places...
- 5y ago
- anpat 5y agoMy monitoring is on fire, flipping red to green every minute because of connectivity issues with every single LB in us-west-2.
- baskethead 5y agoIt sounds like their systems design interviews aren’t rigorous enough.
- pm90 5y agoAt this point they should hire specifically for config management and rollout. Mostly /s; I wish the aws engineers the best of luck through this.
- lordnacho 5y agoHow about an ad for a "Status Page Engineer"?
- tyingq 5y agoI'm guessing lots of people fled us-east-1 for us-west-2, after the last outage, and overwhelmed something there.
- yellowsir 5y agonpmjs has problems too :(
- yellowsir 5y agoseems to be up again
- rexreed 5y agoIt's not just AWS - check the down reports:https://downdetector.com/ https://downdetector.com/ Cloudflare having some significant issues as well on certain domains.
- the_pwner224 5y agoHN was also (briefly) down around that same time (roughly 1 hour ago from now).
- chasd00 5y agocan confirm i have multiple salesforce instances down.
- nerdjon 5y agoThe list of affected services is a bit all over the place, especially since I highly doubt Xbox Live or Halo is running on AWS.
- iamricks 5y agolol imagine if azure was just AWS in the backend
- nerdjon 5y agoIs it bad that I can almost see that being a quick and dirty MVP to get out the door while you built your own cloud solution? Raises serious migration and cost issues, but... would be interesting.
- vidarh 5y agoI think for some targeted things there might well be "value added" services you could offer to transparently wrap AWS. E.g. a "write-through" S3 wrapper was something I was actually looking at because some clients when I was contracting were very reluctant to trust anything but AWS for durability but at the same time AWS bandwidth costs were so extortionate that renting our own servers from somewhere like Hetzner and then proxying writes both to a local disk and to S3 and serve up from local disk with a fallback to pull a fresh copy from S3 if missing broke even at a quite small number of terabytes transferred each month. The nice part about something like that is that properly wrapped you can change your durable storage as needed, and can easily even selectively pick "cheaper but less trusted" options for less critical data. It also allows you to leverage AWS features to ride closer to the wire. E.g. to take another example than storage, I've used this to cut the cost of managed hosting by being to spill over onto EC2 instances in the past, allowing you to run at much higher utilisation rate than what you can safely on managed / colo / on-prem servers alone - as a result, ironically the ability to spill over onto EC2 makes EC2 far less competitive in terms of cost to actually run stuff on most of the time.
- branon 5y agoYup. Having issues with IT Glue and Duo here.
- rd0 5y agoDuo issues here as well.
- swaraj 5y agoOur IaaS vendor, Aptible, reports us-west-1 is down / throwing errors
- deleted 5y ago[deleted]
- 300bps 5y agoI'm on us-east-1 and everything is fine for me including: * EC2 instances * AWS Workspaces * FSx for Windows * AWS Directory Service * S3 Buckets
- samgranieri 5y agoAt least this still works: https://livemap.pingdom.com/ https://livemap.pingdom.com/
- fy20 5y agoPartially, the stats on the right are wrong. For me it shows: Website outages in the past hour 86,967 Lowest 16,208 Average 16,208 Highest 16,209
- navidkhn1 5y agoMy personal health dashboard on AWS shows "InternetConnectivity operational issue us-west-2" [07:42 AM PST] We are investigating Internet connectivity issues to the US-WEST-2 Region.
- iJohnDoe 5y agoProbably a silly question, but what are you using to get this info?
- pbalau 5y agoA browser most likely... this is the "Personal Health Dashboard" one gets for each AWS account /edit: https://phd.aws.amazon.com/phd/home#/dashboard/open-issues https://phd.aws.amazon.com/phd/home#/dashboard/open-issues
- iJohnDoe 5y agoThanks. Didn’t know if it was a custom dashboard or something provided by AWS.
- dannyw 5y agoPrime video down for me. Australia.
- wirelesspotat 5y agoAWS status page shows an update: > AWS Internet Connectivity (Oregon): 7:42 AM PST We are investigating Internet connectivity issues to the US-WEST-2 Region. Source: https://status.aws.amazon.com https://status.aws.amazon.com
- alvis 5y agoOh. not again...
- curtisblaine 5y agoThe npm registry is down too.
- menmob 5y ago7:42 AM PST We are investigating Internet connectivity issues to the US-WEST-2 Region.
- adnauseum 5y agoSeems like ever since Microsoft bought AWS, it's been going down an awful lot.
- endisneigh 5y ago> Seems like ever since Microsoft bought AWS, it's been going down an awful lot. What?
- rfoo 5y agoSatire. Every time Github went down multiple people post on HN saying "every since they were bought by Microsoft, ...". As annoying as those Rust evangelists on every single memory corruption bug.
- staticassertion 5y ago> As annoying as those Rust evangelists on every single memory corruption bug. First of all, how dare you! Second, shoulda used rust ¯\_(ツ)_/¯
- deleted 5y ago[deleted]
- metaltyphoon 5y agoObviously while using Arch btw
- elasticventures 5y agoI could have written the OP message a year ago -- I used to feel the same way. Plz don't disparage Rust evangelism! Rust is awesome. yes it is complex, frequently annoying, easy to learn difficult to master. I'm speaking from a 30 year dev career. a few months ago I intended to do a quick investigation into RUST to validate my "i really don't need to learn this" specifically for an embedded project. Within a few hours I found I had become a zealot. Rust has too many "omg, i should tell everybody about this" behaviors that I can't even find my favorite aspect yet. It's equivalent to a lost soul finding Christianity and accepting the lords blessing and forgiveness! The weight that is lifted of being forgiven to your sins resulting == no more guilt, it's all forgiven! immediately reduction of cognitive dissonance. in this example with rust, it's pointer tracking and memory management, but it's basically the same thing. Rust is for the pious developer. Those people who are still using C++ for fresh starts are the same folks who love to do things the hard & wrong way, or at least those who don't know any better, infidels, unwashed heathen. Join us. join rUSt.
- iJohnDoe 5y agoYes, seeing it too. Seems to be down in a major way. Lots of various AWS services are down. However, so many things depend on AWS that it could just be EC2 is down and it is causing a rippling affect.
- rakem 5y agoproof: https://twitter.com/thedrunkneteng/status/1471144289476526081 https://twitter.com/thedrunkneteng/status/147114428947652608...
- aaronharnly 5y agoEveryone who spent the past week migrating from us-east-1 to us-west-2: this joke is on you. :)
- DarthNebo 5y ago"US-EAST-1 or bust" being manifested right now.
- monkeybutton 5y agoI really appreciate seeing these threads. Let's me know I haven't lost it.
- deleted 5y ago[deleted]
- mtschopp 5y agoCould it be related to a Log4j issue?
- yawnxyz 5y agoVercel is down too. My sites run on Cloudflare and Vercel, and I can't even log in to those right now. I'm curious — what does Hacker News run on? It seems impervious to any kind of downtime...
- qeternity 5y ago> I'm curious — what does Hacker News run on? It seems impervious to any kind of downtime... On a dirty, disgusting dedicated server.
- yawnxyz 5y ago> On a dirty, disgusting dedicated server. I'm adding "reliable" into that mix. Too bad they're too expensive and hard to setup for side projects, but HN is probably one of the most stable site I frequently visit, and I don't even think about it.
- deleted 5y ago[deleted]
- Nextgrid 5y agoI disagree that they're expensive. Expensive to own maybe, but you can rent them on a monthly basis from something like Hetzner or OVH for a fraction of the cost of AWS (especially when you include bandwidth which is free and unmetered in this case) and they handle hardware maintenance for you. Hard to setup is relative. It all depends on what you're doing and how much reliability you need. For a side project or a dev server you can just start with Debian, stick to packaged software (most language runtimes and services such as Postgres or Redis are available) as much as possible and call it a day. You can even enable auto-updates on such a stable distro. The knowledge you'll gain by dealing with bare-metal is also going to be useful in the cloud even in container environments.
- jjav 5y ago> I'm adding "reliable" into that mix. Too bad they're too expensive and hard to setup for side projects I mean they're not particularly. Unless the use is extremely minimal, it'll always be cheaper to buy a small server even for a side project. I use cloud VMs for projects that can live on $5/mo VMs because at that usage rate I'll never break even to buy a machine. But as soon as your AWS bill is even like $50/mo, worth to start looking at alternatives.
- iso1631 5y agoMust be a Y in the day. It amazes me how many projects exist that don't even have multi-region capability, let alone no single point of failure
- ahallock 5y agoYou're saying that as if it's a walk in the park to set up and not cost prohibitive, in terms of opportunity cost and budget, especially for smaller companies.
- tylerrobinson 5y agoRight. Downtime (or perception of downtime) is bad for business, so AWS is surely working to improve reliability to avoid more black eyes on their uptime. But at the same time, an AWS customer might be considering multi-region functionality in AWS to protect themselves ... from AWS making a mistake. As a customer, it's unclear what the right approach is. Invest more with your vendor who caused the problem in the first place, or trust that they'll improve uptime?
- Spivak 5y agoThis might be a multi-region problem. Auth0 as an example has three US regions and two of them are down.
- staticassertion 5y agoIDK, don't you end up with a bunch of extra costs? Like you're going to literally pay more money because now you have cross region replication charges, and then you're going to pay a latency cost, and then you may end up needing to overprovision your compute, etc. All to go from, idk, 99.9% uptime to 99.95% (throwing out these numbers)? The thing is when AWS goes down so much of the internet goes down that companies don't really get called out individually.
- dilyevsky 5y agoIf you just sat that there and took that 8 hour outage you’re barely even 99.9 for the year
- the_iceman 5y agoConfirmed experiencing significant issues in US-WEST-1 as well
- the_iceman 5y agoExperiencing significant issues in US-WEST-1
- mgbmtl 5y agoQuickBooks Online seems to be down, and they seem to be hosted on AWS.
- deleted 5y ago[deleted]
- zedpm 5y agoWow, yeah, us-west-1 AND us-west-2 are reporting connectivity issues. I'm guessing this is related to the Auth0 outage that's currently going on too.
- streetcat1 5y agoRemember, every 12 secs take one 9.
- deleted 5y ago[deleted]
- andrew_ 5y agoRoot logins are suffering some kind of "captcha outage." The buzz has just begun https://twitter.com/search?q=aws%20captcha&src=typed_query https://twitter.com/search?q=aws%20captcha&src=typed_query
- qwertyuiop_ 5y agoLog4jammed ?
- cebert 5y agoThis outage is extremely frustrating to me. My company hosts all our apps in gov cloud. Gov Cloud West 1 is also down, but the AWS Gov Cloud status page indicates that everything is healthy and green. I thought AWS's incident response to the East outage last week was that they'd update the status page to better reflect reality. Gov Cloud Status Page: https://status.aws.amazon.com/govcloud https://status.aws.amazon.com/govcloud
- texasviking 5y agoWe are in the same boat. Finally updated "We are investigating Internet connectivity issues to the US-GOV-WEST-1 Region"
- chasd00 5y agoi had multiple govcloud hosted salesforce instances down but they appear to be coming back up now.
- clavicat 5y agoWe are barbarians occupying a city built by an advanced civilization, marveling at the hot baths but know nothing about how their builders keep them running. One day, the baths will drain and anyone who remembers how to fill them up will have died.
- tibbar 5y agoThe remarkable thing is that today no one knows how to “fill up the baths”, or to do more than a small part of the job. Teams exist with extremely narrow expertise. But if anything, there are more options today for DIY infrastructure - way easier to be more advanced than “run the Apache on the server box.”
- nic_wilson 5y agoIs this a fifth season reference?
- bawolff 5y agoOn-prem is much more rare, but its hardly non-existent. Plenty of people know how to do this sort of thing.
- vidarh 5y agoOn-prem, maybe, but if you include co-located equipment and managed hosting I don't even think it's more rare in absolute terms. Just smaller as a percentage of overall hosting.
- starfallg 5y agoThere's (still?) a lot of on-prem and managed hosting. It's probably the majority of hosted services. Otherwise VMWare wouldn't be doing as well as it is.
- Waterluvian 5y agoThis has been true for a long time and it is not a bad thing. It’s an easy target to romanticize but realistically, any alternative is basically a way of saying: “let’s stop evolving.”
- stevenhubertron 5y agoYeah. It's inconsistent but a number of my production servers appear to be down. Along with my New Relic logging.
- phgn 5y agoThis also seems to affect NPM, I can't install packages locally :/
- markbnj 5y agoOur systems that talk to S3 in CA and OR are timing out trying to open SSL connections. AWS lists outages in these regions on their status page.
- commandlinefan 5y ago"Hey boss, that thing that took down us-east-1... that can't take down us-west-1 next week, can it?" "No, no, of course not" "Should I check?" "No, don't waste time checking, get back to your TPS reports"
- sam0x17 5y agoThey really need to stop requiring SVPs or higher to show non-green status on the status page, as other HNers have revealed in last week's AWS post. It's effectively not a status page, and they could probably be sued if it can be demonstrated that X service was down but the status page showed green (since the SLA is based on status page). Should be automated and based on sample deployments running in every region and every service. And they should use non-AWS instances to do the sampling, so they can actually sample when, say, we experience the obligatory black friday us-east-1 outage every year.
- ceejayoz 5y agoThey were much faster than usual about updating the AWS Status page.
- JshWright 5y agoOur ~four person ops team shouldn't be able to have our status page updated 15 minutes before the upstream status page...
- Isthatablackgsd 5y agoI thought Status Pages or Health Pages is designed to automate the reporting and checking the status automatically. This was my impression when I came across those status pages. Apparently, it is not automated and only update it manually. What is the point of having a status pages if it cannot be automated? I'm sure FAANG and tech conglomerates don't want it to be automated because of SLA. I'm surprised with FAANG hosted their stuff in their competitors cloud services without providing a fallback cloud service if the primary service is down. Sure it cost money but it would be effective this way than putting all eggs in one basket.
- erhk 5y agoAny public communication is handled by people not machines. No one wants to make an automated status page because theres a shit ton of real noise that users dont need to hear about, nd theres a lot of outages that automation won't accurately catch
- gz5 5y agolooks specific to certain (possibly AWS hosted or partially dependent) services such as Auth0: https://status.auth0.com/ https://status.auth0.com/ e.g. our services running on AWS are fine right now, but new sessions dependent on Auth0 are not.
- tuzemec 5y agoIs that related to the current NPM status (https://status.npmjs.org/ https://status.npmjs.org/)?
- moneywoes 5y agoBack up
- jcoder 5y agoThis is new… Siri hasn’t been able to connect for me since this began
- ghawkescs 5y agoSame thing here.
- BTCOG 5y agoCan't use MFA right now to get into multiple instances due to this outage.
- codercotton 5y ago"Everything is fine." - https://status.aws.amazon.com https://status.aws.amazon.com
- rytrix 5y agoEverything *is* fine now. The status page previously reflected an issue much quicker than last time.
- myth_drannon 5y agoThat's the price of PIP culture and burning out your devs. Now noone wants to work at Amazon and they can only hire new grads.
- Throwawayaerlei 5y agoI hear they do get people who want to be able to get experience at AWS's scale, there's only a few places for that. The thing that really gets me is the reports from the last major outage a few days ago about how pervasive lying inside the company is. This really doesn't work well for engineering and we're possibly seeing the results of that. We should certainly expect to see that becoming visible the more time goes on without a major cultural shift. Which given that the guy who ran AWS now runs all of Amazon.com....
- earthboundkid 5y agoHOST THE GODDAMN STATUS PAGE ON AZURE FOR FUCKS SAKE. There is zero excuse for this shit. Be professional. Acknowledge reality. It is logically impossible to run your own status page. Trying to do so just wastes everyone else on the internet's time when you have an outage.
- aaronharnly 5y agoSeriously.
- boopboopbadoop 5y agoYou don’t even know what the problem is yet. Stop shouting solutions.
- slig 5y agoThe problem is very clear: the status page is not working as it should.
- mark-r 5y agoGiven the legal liabilities Amazon has with their SLAs, it may be working exactly as Amazon thinks it should. Whether anybody would agree with that assessment should be obvious.
- boopboopbadoop 5y agoWhat if the problem is not an AWS problem? My point is that you don’t know what the problem is, you’re assuming.
- yunwal 5y agoThe problem is that AWS can't update their status page to reflect that there's a problem. This happens during every AWS incident without fail.
- hatware 5y ago
- deleted 5y ago[deleted]
- johnisgood 5y agoAnd I kept getting "We're having some trouble serving your request. Sorry!" on HN for the past 10 minutes or something.
- edoceo 5y agoTraffic flood to this site for status reports on AWS
- account758 5y agoAWS Global Accelerator not working correctly anymore as well, connections dropped worldwide. Seems like it is managed from us-west-2 and not redundant.
- electroly 5y agoThis comment taught me about the existence of Global Accelerator and, somewhat ironically given the context, we decided to deploy it today. Pretty neat! I'll have to keep in mind that I learned about it because of a worldwide outage :) Thanks!
- RunOutOfMemory 5y agoout of memory again. ;<
- turtlebits 5y agoThat was fun. Badges weren't working (daily checkin required) so the front desk had to manually activate them. Slack wasn't sending messages and Pagerduty was throwing 500's.
- api 5y ago... because you need to contact a server 1000 miles away to issue badges in your building. This cloud-for-everything-even-local-devices thing is both hilarious and sad. I wonder if anyone had trouble doing their dishes or laundry today, because I'm sure someone thought dish washers and washing machines needed cloud.
- ec109685 5y agoI don't know if you can say an on-premise badge hosting service would be more reliable than the cloud.
- kazen44 5y agowell, atleast you have the agency to do something about it yourself. also, building access systems should be hosted in the building they reside in for security reasons anyways.
- marcosdumay 5y agoThis creates some really fun failure cases on the form of "I need to enter the building so anybody can enter the building". Depending on the cloud is certainly a very stupid decision. keeping everything inside the building is better, but still not ideal.
- jjav 5y agoAny electronic access system like this requires manual backup. As in, some doors with regular locks using physical keys.
- gitfan86 5y agoI'm so glad that I'm not still the CTO of a startup. I would be getting dozens of e-mails from people without engineering backgrounds asking "Are we multi-cloud", "why didn't you make us multi-cloud"?
- necovek 5y agoWell, why didn't you? :) The response is that this actually works well enough, so the investment required has not pushed anyone to do it (with that meaning building the core infrastructure to make that easy).
- NicoJuicy 5y agoI get the feeling that Havoc will happen when a tornado would reach us-east-1
- cblconfederate 5y agoReminder that the internet was literally invented to avoid this kind of nuclear attack. But i guess people are herdish animals and prefer to die as a group
- throw_m239339 5y agoMore like ultimately all these companies buy into a certain form of vendor lock-in and they have no competence or willingness to migrate or even consider the competition. It's starts with "oh I'm just renting a remote virtual server" and in no time it's "Oh, all my stack is tied to AWS proprietary products" because convenience. That's what Amazon wants.
- doublepg23 5y agoSeems like the Internet level networking is quite robust at this point.
- Graffur 5y agoI thought the whole point of AWS was that you could fail over to a different location?
- CoastalCoder 5y agoAsking as a non-cloud-developer: why would Crunchyroll's recovery [0] lag so much behind AWS's recovery [1]? [0] https://downdetector.com/status/crunchyroll/ https://downdetector.com/status/crunchyroll/ [1] https://downdetector.com/status/aws-amazon-web-services/ https://downdetector.com/status/aws-amazon-web-services/
- spenczar5 5y agoI don't know for sure, but this is generally common because caches get cold. A lot of websites use a cache in front of databases (or template rendering engines, or many other systems). That cache might evict entries based on time - after 5 minutes, the entry is considered invalid. But that means that if you have no traffic for 10 minutes, the cache completely empties. Then when traffic returns, it all skips the cache and actually triggers a real hit to the backend - which is now overwhelmed with traffic. The cache protects the backend in normal behavior, but now it's not doing its job, so the backend has many more requests than usual. In the worst case, those requests are enqueued in a big serial sequence... but the ones at the back of the queue may time out. The client may do something like say "it's taken me 5 seconds and I still don't have a response - I'll abort and retry!" and now you have even _more_ traffic to deal with. So cold caches and retries can conspire to keep a service down for a long time even after the root cause is fixed.
- CoastalCoder 5y agoI'm accustomed with cache-eviction policies based on LRU, age, etc. But in my systems, eviction happens only when (a) the content is known to be invalid, or (b) there's competition for cache space. IIUC the parent comment, it's describing a policy that evicts entries even (a) and (b) are false. Is that common in the web-hosting / CDN world? Or is age considered a proxy for stale?
- spenczar5 5y agoRight, age is used as a proxy for stale, because we often don't have anything better. A lot of web systems work this way - DNS records for example use a "TTL" which means "time to live." If the TTL is 60, then you throw it out of the cache after 60 seconds even if you have room in the cache, and you have no reason to believe it's invalid. This lets independent entities (like a DNS authority) make a change and get it rolled out everywhere. I think the reason this is common is that proving cache invalidity is so hard, especially with the typical "dumb" cache appliances that are widely used. They just do stuff like cache the response bytes for a particular URL; they might not even understand HTTP beyond interpreting the request's headers, and certainly don't really understand the response.
- mysql 5y agoIt's bad that I come here first to see if I am crazy or AWS is actually down.
- joelbondurant 5y agoAWS is the McDonald's of computer hardware. Billions ov hamburgers served, so dey must have da best hamburger cooks n da world. Decades of corporate insistence on outsourcing every shred of hardware talent has left the software industry filled with imbeciles.
- waynecochran 5y agoThere was a brief period of time back in the early 90's where I felt I understood how Linux worked -- the kernel, startup scripts, drivers, processors, boot tools, etc... I could actually work on all levels of the system to some degree. Those days are long gone. I am far removed from many details of the systems I use today. I used to do a lot of assembly programming on multiple systems. Today I am not sure how most of the systems works in much detail.
- cle 5y agoTo an extent, this is one of the goals, to free up engineers to work on higher level things. Whether it meets that goal in some cases is debatable, and it’s certainly not ideal for us engineers who like to get to the bottom of things.
- someguydave 5y ago“working on higher level things” currently implies that depending on many layers of opaque and unreliable lower level hardware and software abstractions is a good idea. I think it is a mistake.
- cle 5y agoThe best conclusion I can come to is "sometimes it works, sometimes it doesn't". Depends on the context. I've seen cases where it works great and other times where it's a huge hassle.
- 10000truths 5y agoFunny, I feel the exact opposite way. The low level stuff is where all the magic happens, where performance improvements can scale by orders of magnitude rather than linearly with a CTO’s budget. I’d much rather figure out how to condense some over-engineered distributed solution down to one machine with resources to spare.
- yottalove 5y agoEven as a software engineer, I think I could build from primitive materials a couple of battery operated transceivers to replace the signal flags or horsemen for critical communications. A little basic physics and materials science goes a long way.
- hdjjhhvvhga 5y agoAn honest question. Why do you guys use AWS instead of dedicated servers? It's terribly expensive in comparison, nowadays equally complex, scalability is not magic and you need proper configuration either way, plus now the outages become more and more common. Frankly, I see no reason.
- dsr_ 5y agoOnce you have committed to a certain way of doing things, the transition costs can be very high. Let's consider RockCo and CloudCo. They both provide a B2B SAAS that is mostly used interactively during the working day, and mostly used via API calls for the rest of the working week. Demand is very much lower on weekends. Both RockCo and CloudCo were founded with a team of six people: a CEO who does sales, a CTO who can do lots of technology things, three general software developers, and one person who manages cloud services (for CloudCo) or wrangles systems and hosting (for RockCo). In the first year, CloudCo spends less on computing than RockCo does, because CloudCo can buy spot instances of VMs in a few minutes and then stop paying for them when the job is done. RockCo needs a month to signficantly change capacity, but once they've bought it, it is relatively cheap to maintain. In the second year, they are both growing. CloudCo buys more average capacity, but is still seeing lots of dynamic changes. RockCo keeps growing capacity. In the third year, they're still growing. CloudCo is noticing that their bills are really high, but all of their infrastructure is oriented to dynamic allocation. They start finding places where it makes sense to keep more VMs around all the time, which cuts the costs a little. RockCo can't absorb a dynamic swing, but their bills are now significantly lower every month than CloudCo's bills, and the machines that they bought two years ago are still quite competitive. A four year replacement cycle is deemed reasonable, with capacity still growing. And bandwidth for RockCo is much cheaper than the same bandwidth for CloudCo. Who's going to win? Well, you can't tell. If they both got unexpectedly sudden growth surges, RockCo might not have been able to keep up. If they both got unexpected lulls, CloudCo might have been able to reduce spending temporarily. RockCo spent more up front but much less over the long term. CloudCo could have avoided hiring their cloud administrator for several months at the beginning. RockCo's systems and network engineer is not cheap. And so on, and so forth.
- tomlagier 5y agoI wonder if AWS will make more or less money from these outages? Will large players flee because of excessive instability? Or will smaller players go from single-AZ to more expensive multi-AZ? My guess is that no-one will leave and lots of single-AZ tenants who should be multi-AZ will use this as the impetus to do it. Honestly, having events like this is probably good for the overall resilience of distributed systems. It's like an immune system, you don't usually fail in the same way repeatedly.
- andy_ppp 5y ago* Free chaos monkey installed in every AZ
- jjav 5y ago> * Free chaos monkey installed in every AZ Only during this beta period, AWS will start charging for this feature soon enough.
- jedberg 5y agoWe (Netflix) begged them for years to create a Chaos Monkey that we could pay for. There were things we just couldn't do ourselves, like simulate a power pull or just drop all network packets on the bare metal. I guess not enough people asked.
- kortex 5y agoCMaaS sounds amazing for resiliency engineering. There's so much I want to be doing to perturb our stack, but I don't know all the ways stuff can go wrong. Sure I can ddos it, kick services and servers offline, etc, but that's what, a few dozen failure modes? Expertise in chaos would be valuable by itself. Not to mention being able to shake parts of the system I normally can't touch. Side note: terraform is pretty good for causing various kinds of chaos, deliberately or otherwise.
- s_dev 5y ago
- pjf 5y agoKentik data on the outage: https://twitter.com/DougMadory/status/1471162450649223173 https://twitter.com/DougMadory/status/1471162450649223173
- belter 5y agoAWS Outage Analysis - December 15, 2021: https://www.thousandeyes.com/blog/aws-outage-analysis-december-15-2021 https://www.thousandeyes.com/blog/aws-outage-analysis-decemb... https://azycqgvwjz.share.thousandeyes.com/view/tests/?roundId=1639581900&metric=availability&scenarioId=httpServer https://azycqgvwjz.share.thousandeyes.com/view/tests/?roundI...
- blueside 5y agoThe vehement defenders of AWS are starting to remind me of the cryptobros
- alberth 5y agoIt appears AWS Status Page is hosted at AWS [0]. Seems like a really bad idea. [0] https://hostingchecker.com/ https://hostingchecker.com/
- redety 5y agowow
- deleted 5y ago[deleted]
- evilhackerdude 5y ago4 hours in, our AWS IoT endpoint (not ATS, Symantec) in us-west-2 is still down according to monitoring, PHD and support.
- mattjaynes 5y agoTangentially related: On Friday Backblaze and B2 were down for 10+ hours to update their systems for the log4j2 vulnerability. Seemed noteworthy for the HN crowd and I posted a link to their announcement when the outage began. However, the post was quickly flagged and disappeared. Genuinely curious, why is announcing some outages ok and others not?
- qaq 5y agoWhat would be the ratio of HNers who are Backblaze customers vs those who are AWS customers. I bet Backblaze number is small enough where Backblaze employees on HN can downvote you enough for it to matter.
- xondono 5y agoAt which point this outages are a sign that something inside AWS is deeply broken and pretty much unfixable?