33 ms·
Tell HN: AWS appears to be down again
Console is flickering between "website is unavailable" and being up for my team. This is happening very frequently just now, reliability seems to have taken a hit.
- biznickman 5y agoWhy isn't Heroku showing a status error despite being offline?
- mikece 5y agoBecause it's built on AWS and uses the AWS status page for it's status info?
- jakub_g 5y agoWhere are you located? "X is down" without location is only moderately useful. I'm having issues with Slack from central EU (Poland) -- can't upload images, or send emoji reactions to post; curiously, text works fine). Wondering if linked
- riknox 5y agoAWS Console runs in us-east-1 so that points to at least that region having issues IIRC. I am also having Slack issues in EU.
- hdjjhhvvhga 5y agoYou should complain to Slack then. It's their problem to choose a reliable provider, and AWS seems to have trouble with keeping this status.
- exabrial 5y agoStat That.
- allocate 5y agoAlso running a big production app in east-1 and we're experiencing issues.
- sprite 5y agoI'm also in east-1 and completely down.
- throwaway81523 5y agoOk, enough AWS outages to say I'm tired of hearing about low end stuff being flaky.
- henriquez 5y agoHeroku isn’t “low end,” it’s a PaaS built on top of AWS. So you’re really just hearing about another AWS outage lol
- christophilus 5y agoThey're not saying Heroku is low end. They're saying, "I'm tired of hearing that it's irresponsible to run your own servers." At least, that's what I understood.
- ryanbrunner 5y agoAny place I've worked at that managed their own servers (to be fair, the last time I worked at a place like that was 2010) definitely had more protracted downtimes than AWS - it just felt not as bad because we were in control of the situation, but at the end of the day that didn't get us up any faster. Another side benefit of being with AWS is when you do have an outage, a lot of other people have outages, and so you sort of blend in with the noise. It's not great to be down, but if you're down and also "big service X" who's also an AWS customer is down, it makes your downtime look less like a lack of competence and more like an unavoidable force of nature.
- dijit 5y agoI guess it's extremely dependent on an org to org basis. I worked at a company that's bread and butter was online services (e-commerce SaaS platform, similar to Netsuite) and we had significantly fewer outages than AWS had. But we had redundancies built in to most things, I'm not saying it was perfect but it worked. The major difference might be that almost nobody is willing to spend 20% of what they spend on AWS/GCP to have a self-hosted solution. The reason "cloud is so expensive" is because they're essentially telling you what the price will be and even if they only spend 40% of that on actual hardware and operations: it's more than most companies would invest in themselves. This is absurd, of course, but it's absolutely true.
- sprite 5y agoMy app running on AWS is currently down. Having intermittent problems with console as well.
- dolibasija 5y agoOne of our EC2 instances in us-east-1c is unavailable and stuck in "stopping" state after a force stop. Interestingly enough, EC2 instances in us-east-1b don't seem to be affected. The console is throwing errors from time to time. As usual no information on AWS status page.
- chrishynes 5y agoI had the same issue with unavailable, but on an instance in us-east-1b. Finally just got the force stop to go through a minute ago and it's now running and available again.
- mike-cardwell 5y agoYour us-east-1b may be the parents us-east-1c. The letters are randomised per AWS account so that instances are spread evenly and biases to certain letters don't lead to biases to certain zones.
- chrishynes 5y agoHuh, that's interesting. Didn't know that, but makes sense.
- thrtythreeforty 5y agoIt's pretty cool. If I recall, they call it "shuffle sharding."
- ciceryadam 5y agoYou can check which availability zone is with: aws ec2 describe-availability-zones --region us-east-1
- throwaway984393 5y agoI'm not sure if we should say "AWS is down" if only us-east-1 is down. That region is more unstable than Marjorie Taylor Greene on a one-legged stool.
- lukeqsee 5y agoI can't get to the console either, receiving a "Temporarily unavailable" notice without branding.
- pawelduda 5y agoBitbucket is affected, pages randomly take forever to load or return 500
- Pandabob 5y agoYep, just botched a merge likely because of this.
- el_duderino 5y agoBitbucket just completed their migration to AWS too. Rough start.
- sprite 5y agoMy Elastic Beanstalk instances are completely unreachable. Seems at the very least ELB is down. Looking @ down detector it looks like this is taking a bunch of sites down with it. As usual AWS status page shows all green.
- RobertKerans 5y agoAssuming crates.io is AWS-backed? Getting fun situation where direct dependencies of an application are downloading but then the sub-dependencies aren't.
- lukeqsee 5y agocrates.io is directly hosted on GitHub, but I'm sure some dependencies use S3 or other AWS services for things.
- RobertKerans 5y agoYep, S3 possibly the villain here
- withinboredom 5y agoI wonder if there's an s3 compatible service with similar pricing that can be used as a fallback? Are digital ocean s3 compatible storage accounts's backed by real s3?
- RobertKerans 5y agoafaik there's nothing tying it specifically to GH (where the metatada is), and then the actual code is just in an S3 bucket, so in theory should be reasonably easy [ha!] to just host anywhere. In theory, I mean that's a massive lump of stuff, and surely wherever it gets hosted is going to face exactly the same issues (though if it does become very widely used, then you'd think every major provider that controls infra could easily have a mirror)
- Ancapistani 5y agoWould Wasabi.com meet your requirements? I’m not affiliated with them, and haven’t even really used them other than to explore a bit. They come highly recommended by my acquaintances, though.
- RobertKerans 5y ago
- sh4un 5y agoDamn you all eggs in one basket.
- potas 5y agoSlack seems to have some issues because of that - I'm not sure if anyone is receiving messages, as it became completely silent for the last 15 minutes or so.
- Pandabob 5y agoUploading images doesn't work for me.
- jenoer 5y agoSending and receiving messages works here, but editing them does not, it throws an error. Statuses such as "calling" also do not seem to be updated any longer. Edit: Restarting Slack does update the edited messages. Edit 15:24 CET: Slack is back up.
- jakub_g 5y agoSame: only normal text seems kinda working - edits failing or working with big lag; - "Threads" view slow; - can't emoji-react; - can't upload images; - people also say they can't join new channels.
- deleted 5y ago[deleted]
- oneeyedpigeon 5y agoNew messages seem to be ok for me, but editing old ones and uploading images both seem to be broken right now.
- jakub_g 5y agohttps://status.slack.com/2021-12/a17eae991fdc437d https://status.slack.com/2021-12/a17eae991fdc437d > We are experiencing issues with file uploads, message editing, and other services. We're currently investigating the issue and will provide a status update once we have more information. > Dec 22, 1:58 PM GMT+1
- aden1ne 5y ago
- hnarn 5y agoIs there a history of AWS downtimes available somewhere? This makes what, three times in as many months? edit: The question isn't necessarily AWS specific, just any data on amount of downtime per cloud provider on a timeline would be nice.
- LuciusVerus 5y agoI'd say three times in as many weeks, give it or take
- MatteoFrigo 5y agoI don't know about AWS, but both Google Cloud and Oracle Cloud maintain at least a high level history of past outages. See https://status.cloud.google.com/summary https://status.cloud.google.com/summary and https://ocistatus.oraclecloud.com/history https://ocistatus.oraclecloud.com/history
- dijit 5y agoGiven the hilariously awful reputation of the AWS status page I would hazard a guess that such a page would also be incredibly inaccurate. If you can’t even admit you’re having an issue how can you keep an accurate record?
- cassianoleal 5y agoSimilar with GCP. We had a pretty bad outage once where the status page was showing all green. Google informed us that because the actual issue was further down the stack and didn't trigger any internal SLOs the status didn't get an update. It took them hours to acknowledge and fix it.
- dijit 5y agoAssuming you have a support contract the rep should send out a post-mortem page. This is what happens when we've been affected by outages (even without involving support).
- fipar 5y agohttps://downdetector.com/status/aws-amazon-web-services/ https://downdetector.com/status/aws-amazon-web-services/
- darkwater 5y agoFields of green here https://status.aws.amazon.com/ https://status.aws.amazon.com/ Anyway I can access the web console with no issue (eu-west)
- temp0826 5y agoChanges to this page require very high level management approvals (source: used to work at aws)
- hnarn 5y agoI think it's pretty widely accepted that AWS' own status pages are utterly useless.
- darkwater 5y agoYeah, it was just to confirm that this time was no different :)
- hdjjhhvvhga 5y agoIn Russia they have a specific name for it: https://en.wikipedia.org/wiki/Potemkin_village https://en.wikipedia.org/wiki/Potemkin_village
- s_dev 5y agoYou would think that but there always a few contrarian AWS evangelists in the comments going on about the "difficulty" in operating a status page as though it were trying to conjure a N=NP proof. Like how come down detector can do a superb job of detecting when AWS goes down and AWS can't? Because AWS doesn't want account managers of SLAs asking for credits for the uptime they're paying for but not getting. https://downdetector.co.uk/status/aws-amazon-web-services/ https://downdetector.co.uk/status/aws-amazon-web-services/
- lordnacho 5y agoThe elite DevOps teams are always assigned to the status page
- schnebbau 5y agoSo, how many execs are going to push to move to self-managed hosting in the new year? Packaging a way to migrate off AWS could be a unicorn idea.
- mikece 5y agoWould need one hell of a compressional algorithm to keep the data exfiltration costs down.
- pm90 5y agoPied Piper
- qwertyuiop_ 5y agoNone. Amazon hired all ex VPS, CTOs, Directors of small, medium large companies with Rolodexes.
- adamm255 5y agoAnyone using VMware Cloud services is probably laughing. Just chuck it at Azure or GCP or back on prem.
- wallacoloo 5y agoAWS has its Outpost product for on-prem hosting. not 100% self-managed, but maybe enough to satisfy the execs and make your market a bit smaller.
- Nextgrid 5y agoDoes it come with its own locally-hosted console or does it still rely on the main AWS control plane? If the latter then it could be affected too.
- dehrmann 5y agoDepends on how many customers are ready to move to a different vendor. I suspect most customers are forgiving because either they were also down or half the services they use were down. You don't get fired for hosting in AWS.
- captn3m0 5y ago4:35 AM PST We are investigating increased EC2 launched failures and networking connectivity issues for some instances in a single Availability Zone (USE1-AZ4) in the US-EAST-1 Region. Other Availability Zones within the US-EAST-1 Region are not affected by this issue. via https://stop.lying.cloud/ https://stop.lying.cloud/
- junon 5y agoCan anyone explain the affiliation of stop.lying.cloud to Amazon? All of the legalese in the header/footer seem to indicate it's actually owned and run by Amazon. If so... why? Why not just... use the real status page? I mean I'm glad it exists, don't get me wrong. Just weird that they'd have two status pages, one seemingly existing only to sort of 'mock' themselves...
- taspeotis 5y agoThe people who maintain the unofficial site would have, at some point, used their CTRL and C keys followed not immediately, but closely by, their CTRL and V keys.
- junon 5y agoBut that is copyright infringement. You're not allowed to copy some work, modify it, then slap the original copyright on it. This is an illegal website, prone to being taken down by AWS. It's just strange.
- Acestus 5y agoI would be worried because getting taken down is Amazon’s speciality.
- bsenftner 5y agoActually, having a satire site taken down over copyright is one of the best ways to extort large amounts of money from the copyright holder, because constitutional attorneys will seep from the floorboards and appear in your shower trying to be an attorney on that case. Satire is extremely protected speech.
- sreitshamer 5y agoConsole is sluggish for me, but S3 (us-east-1) seems to work fine.
- loudtieblahblah 5y agoYay! Adult snowday!
- RobertKerans 5y agoApropos of nothing, but a few Christmasses ago the place I worked had a dedicated fibre line that some workmen doing gas line repairs sawed straight through, took out everything; I was just drone worker at the time & it was a beautiful thing
- sswaner 5y agoNot down as of 7:40 EST. US-EAST-1 hosted site (athene.com). Cognito, API Gateway, Lambda, S3, DynamoDB, RDS, S3, Cloudfront.
- throwaway875487 5y agoOur RDS instances have completely packed up. Hell knows what's going on. Here come the customer support tickets.
- IceWreck 5y agoHonestly my server at home has more uptime than US-East-1
- BossingAround 5y agoDoes your server at home handle similar traffic to that of US-East-1 since you're comparing uptime? Simiarly, my laptop, if I keep it plugged in the wall, and enable httpd on localhost, will surely have better uptime than any of the top clouds. I'd bet that it'd have 100% uptime if I plugged in a UPS and cared for traffic on my local network only.
- IceWreck 5y agoNo but I access my home-server remotely from my university all the time and it hasn't gone down once. Better uptime than paying for EC2 on AWS US-East-1. Obviously this approach isn't scalable but it serves me well.
- amelius 5y ago> Obviously this approach isn't scalable but it serves me well. It's perfectly scalable. Just give everybody their own home server.
- Sammi 5y ago> Does your server at home handle similar traffic to that of US-East-1 since you're comparing uptime? Of course it doesn't. Why are you asking antagonistic questions?
- vegai_ 5y ago5ish years ago it was common knowledge that us-east-1 is generally the worst place to put anything that needs to be reliable. I guess this is still true?
- beermonster 5y agous-east-1 seems to be AWS’s not so well kept little dark secret! In all seriousness though - even non-regional AWS services seem to have ties to us-east-1 as evidenced by the recent outages. So you might be impacted even if it looks like (on paper at least) you’re not using any services tied to that region.
- taf2 5y agoI don't know about that. It was more like common knowledge that one availability zone in us-east-1 was a problem - you would have to figure out which one it was usually by spinning up instances in all 4 zones (now 6)... and that it was the largest of all regions making it ideal place to put your service if you wanted to be close to other vendors/partners in AWS...
- thow-58d4e8b 5y agoUnfortunately, the fact that us-east-1 is roughly 10% cheaper than other regions usually overrides any other concerns
- rswail 5y agoSo why are people not migrating out of us-east-1? Operating in ap-southeast, we weren't that affected by the us-east-1 down time, although our system is reasonably static and doesn't make lots of IAM calls (which seems to be a large SPOF from us-east-1).
- taf2 5y agolatency. us-east-1 is positioned very nicely relative to many large businesses in North America and Europe. This gives you pretty good access to a very large percentage of the economies of the world with good latency... while not requiring you to architect your application around multiple regions...
- dijit 5y agoSome “global” systems run in us-east1 even if you’re not hosted there a service you depend on might be. Notably: cognito, r53 and the default web UI. (You can work around the webui one I’m told, by passing a different domain instead of just console.aws.amazon.com)
- watermelon0 5y agoDon't forget about CloudFront, which can only be configured via us-east-1.
- bognition 5y agoWhat a way to start my day
- omosubi 5y agoI do wonder if the great resignation has anything to do with this. My team (no affiliation with Amazon) was cut in half from last year and we are struggling to keep up with all the work
- clavicat 5y agoHow much more frequent do these outages need to become before it starts triggering SLA limits?
- streamofdigits 5y agoSomebody call the IT department
- 300bps 5y agoCan we please stop saying, “AWS is down”? AWS consists of over 200 services offered in 86 availability zones in 26 regions each with their own availability. If one service in one availability zone being impaired equals a post about “AWS is down” we might as well auto-post that every day.
- sawmurai 5y agoIt's like my grandma saying "Honey, the internet is broken again." xD
- KptMarchewa 5y agoWould be cool if this wasn't the region where AWS hosts their internals, making other regions unusable, right?
- satya71 5y agoSeems enough services in us-east-1 are down to cause most apps to fail. My simple app uses 10s of AWS services, at least some of which are out.
- 300bps 5y agoI may have seen more of these posts than you. The last one I saw where “AWS is down” was us-west-1.
- omh2 5y agoAWS doesn't follow their own advice about hosting multi-regional so every time us-east-1 has significant issues pretty much every AZ and region is affected. Specifically large parts of the management API, and IAM service are seemingly centrally hosted in us-east-1. If your infrastructure is static you'll largely avoid the fallout, but if you rely on API calls or dynamically created resources you can get caught in the blast regardless of region
- mule1 5y agoFeel for devops peeps who are just trying to chill for Christmas
- andyjih_ 5y agoThe most hilarious irony of not being able to acknowledge a 4AM page in the PagerDuty mobile app because AWS is down.
- exikyut 5y ago(Which was about AWS being down?)
- antihero 5y agoI wonder how many 9s AWS is going for. Can't be a lot of 9s anymore.
- sascha_sl 5y agoquay.io is also dead, as well as giphy, some parts of slack just the weekly internet apocalypse, happy holdidays fellow SREs
- izietto 5y agoI guess that's why I'm experiencing weird issues with Heroku: remote: Compressing source files... done. remote: Building source: remote: remote: ! Heroku Git error, please try again shortly. remote: ! See http://status.heroku.com for current Heroku platform status. remote: ! If the problem persists, please open a ticket remote: ! on https://help.heroku.com/tickets/new
- dijit 5y agoYes. Another thread: https://news.ycombinator.com/item?id=29648325 https://news.ycombinator.com/item?id=29648325
- anshumankmr 5y agoIf AWS, GCP and Azure go down, we will be back in the stone ages, right?
- dijit 5y agoThe only stuff that will work will probably depend on things in AWS in some form. That, or people never took the “if AWS goes down then lots of people will have a problem, so we’ll be fine” line seriously; there are few such cases.
- deleted 5y ago[deleted]
- deleted 5y ago[deleted]
- rsp1984 5y agoBitbucket having issues too: https://bitbucket.status.atlassian.com/ https://bitbucket.status.atlassian.com/
- debarshri 5y agoHubspot seems to be down too [0]. [0] https://status.hubspot.com/ https://status.hubspot.com/
- Demcox 5y agoImgur is suffering from this too, I think.
- stunt 5y agoIt seems that it's due to powerloss. [05:01 AM PST] We can confirm a loss of power within a single data center within a single Availability Zone (USE1-AZ4) in the US-EAST-1 Region. This is affecting availability and connectivity to EC2 instances that are part of the affected data center within the affected Availability Zone. We are also experiencing elevated RunInstance API error rates for launches within the affected Availability Zone. Connectivity and power to other data centers within the affected Availability Zone, or other Availability Zones within the US-EAST-1 Region are not affected by this issue, but we would recommend failing away from the affected Availability Zone (USE1-AZ4) if you are able to do so. We continue to work to address the issue and restore power within the affected data center.
- aledalgrande 5y agoIf you haven't seen yet, news is it was a power loss: > 5:01 AM PST We can confirm a loss of power within a single data center within a single Availability Zone (USE1-AZ4) in the US-EAST-1 Region. This is affecting availability and connectivity to EC2 instances that are part of the affected data center within the affected Availability Zone. We are also experiencing elevated RunInstance API error rates for launches within the affected Availability Zone. Connectivity and power to other data centers within the affected Availability Zone, or other Availability Zones within the US-EAST-1 Region are not affected by this issue, but we would recommend failing away from the affected Availability Zone (USE1-AZ4) if you are able to do so. We continue to work to address the issue and restore power within the affected data center.
- codeduck 5y agoanother example of a single dc in a single AZ rendering an entire region almost unusable. This has shades of eu-central-1 all over again.
- nightpool 5y agoAmazon is claiming the failure is limited to a single AZ. Are you seeing failures for instances outside of that AZ? If not, how has this rendered "the entire region almost unusable"?
- londons_explore 5y agoA lot of people will automatically fail over jobs to other AZ's. That often involves spinning up lots more EC2 instances and moving PB's of data. The end result is all capacity on other AZ's gets used up, and networks get full to capacity, and even if those other zones are technically working, practically they aren't really usable.
- Godel_unicode 5y agoThat doesn't appear to have happened though, I haven't seen issues outside az4
- networkisfine 5y agoIsn't the point of the design of an availability zone having multiple data centers so that if a single data center in the availability zone fails, services aren't affected?
- temptemptemp111 5y ago
- RONROC 5y agoThe prevailing wisdom throughout the last couple of years was: “ditch your on-prem infrastructure and migrate to a major cloud provider” And its starting to seem like it could be something like: “ditch your on-prem infrastructure and spin up your own managed cloud” This is probably untenable for larger orgs where convenience gets the blank check treatment, but for smaller operations that can’t realize that value at scale and are spooked by these outages, what are the alternatives?
- paulryanrogers 5y agoSpread the risk? Smaller on prem and cloud / rented bare metal?
- Spivak 5y agoNah, it's actually better to concentrate the risk in this case. If your app depends on a few 3rd party services -- SendGrid, Twilio, Okta and they're all hosted on different infra then congrats! You're gonna have issues when any one of them are down, yayyy. Also the marketing benefit can't be downplayed. If your postmortem is "AWS was having issues" then your execs and customers just accept that as the cost of doing business because there's a built-in assumption that AWS, Azure, GCP are world class and any in-house team couldn't do it better.
- aflag 5y ago> Also the marketing benefit can't be downplayed. If your postmortem is "AWS was having issues" then your execs and customers just accept that as the cost of doing business because there's a built-in assumption that AWS, Azure, GCP are world class and any in-house team couldn't do it better. In my experience, execs and customers don't treat an outage differently because AWS is at fault. Though the developers do often have the attitude that it's "someone else's problem", which can actually can make execs more worried than if the problem was well known and under the company's control.
- Victerius 5y agoI'm tempted to found a startup to help businesses migrate from cloud providers to on-prem infrastructure.
- CaptRon 5y agoAt least HN works.
- sctgrhm 5y agoInvision image uploads are down too because of this : https://status.invisionapp.com/ https://status.invisionapp.com/
- reactive55 5y agoBitbucket is down as well because of this. https://bitbucket.status.atlassian.com/incidents/r8kyb5w606g5 https://bitbucket.status.atlassian.com/incidents/r8kyb5w606g...
- reactive55 5y agoBitbucket is down as well
- bobviolier 5y agoSeems unlogical that this is just a single region in a single US region We are having issues pulling images from public.ecr.aws from an EU region.
- saxonww 5y agoI don't know what's still true, but at one point us-east-1 seemed more critical than other regions because there were some things that had to be there. One thing that comes to mind is ACM certificates used with things like API Gateway (probably Cloudfront), they had to be in us-east-1 no matter where the rest of your infrastructure was. So it's not shocking to me that something going down in us-east-1 could have impact on other regions.
- Hippocrates 5y agoEvery time a major cloud provider has an outage, Infra people and execs cry foul and say we need to move to <the other one>. But does anyone really have an objective measure of how clouds stack up reliability-wise? I doubt it, since outages and their effects are nuanced. The other move is that they want to go multi-cloud... But I’ve been involved in enough multi-cloud initiatives to know how much time and effort those soak up, not to mention the overhead costs of maintaining two sets of infra sub-optimally. I would say that for most businesses, these costs far exceed that occasional six-hour-long outage.
- metadat 5y agoI know the Oracle OCI cloud has a reputation for never going hard-down, but also realize HN seems to loathe Big Red (understandably, to a degree, though OCI is pretty nice IME and _very_ predictable).
- SixDouble5321 5y agoI don't think it's unfair. They aren't the worst villain, but they are up there.
- mongrelion 5y agoI agree with you. I think that having multi-AZ is the first thing to figure out before wanting to do multi-cloud, which is just another buzzword taken out of management's bullshit bucket :)
- Hippocrates 5y agoAgree, and multi AZ is usually easy. IME with AWS and GCP the control plane is the same, the scaling works across AZ, bandwidth is free and latency is near zero. The level of effort to do that is simply ticking the right boxes at setup time IME.
- Jweb_Guru 5y agoCross-AZ bandwidth is far from free and the biggest reason companies avoid it (IMO). Also latency is not near zero but I don't think that's the primary reason.
- devoutsalsa 5y agoWe'll never really know the answer, but I have to wonder what percentage of comments on this thread are from Amazon downplaying the severity & other cloud providers hyping it up.
- mongrelion 5y agoYou give HN too much credit.
- temptemptemp111 5y ago
- iso1631 5y agoAhh, the cloud https://imgflip.com/i/5yrt24 https://imgflip.com/i/5yrt24
- ItsBob 5y agoI've built out many 42U racks in DC's in my time and there were a couple of rules that we never skipped: 1. Dual power in each server/device - One PSU was powered by one outlet, the other PSU by a different one with a different source meaning that we can lose a single power supply/circuit and nothing happens 2. Dual network (at minimum) - For the same reasons as above since the switches didn't always have dual power in them. I've only had a DC fail once when the engineer was performing work on the power circuitry for the DC and thought he was taking down one, but was in fact the wrong one and took both power circuits down at the same time. However, a power cut (in the traditional sense where the supplier has a failure so nothing comes in over the wire) should have literally zero effect! What am I missing? I've never worked anywhere with Amazon's budget so why are they not handling this? Is it more than just the imcoming supply being down?
- Bluecobra 5y ago> What am I missing? My guess is that they cheaped out in having redundant PSUs to get you to use multiple availability zones. (More zones = more revenue) Even a single PSU shouldn’t be an issue if they plugged in an ATS switch though.
- Godel_unicode 5y agoUnless the ATS breaks, which happens.
- mnordhoff 5y agoYup. I'm still upset (but not angry) about https://status.linode.com/incidents/kqhypy8v5cm8 https://status.linode.com/incidents/kqhypy8v5cm8.
- Bluecobra 5y agoFor sure, in my context I meant a ATS in single rack/cabinet. If that went bad the blast radius would be contained to a single cabinet. But yeah, anything can and will happen. At another place I worked at, a site UPS took down an entire server room. It was pretty nice Eaton system but there was some event that fried the whole thing. Eaton had to send an specialist to investigate the matter as those events are pretty rare.
- sydthrowaway 5y agoSwitch to Azure
- camdenreslink 5y agoWho needs chaos monkey? Just host on AWS for a similar effect.
- anonu 5y agoBetter polish off your BCP docs. People will be asking for them quite a bit more in the new year.
- gtsop 5y agoQuestion to the sysadmins here: Is it really that outrageous of amazon to have such issues or are people way to spoiled to appreciate the effort that goes into maintaining such a service? Edit: Not supporting amazon, i generally dislike the company. I just don't understand the extend to which the criticism is justified
- dsr_ 5y agoThe issue is in three parts: 1. Did AMZN build an appropriate architecture? 2. Did AMZN properly represent that architecture in both documentation and sales efforts? 3. What the heck is going on with AMZN? Let's say that they build an environment in which power is not fully redundant and tested at the rack level, but is fully redundant and tested across multiple availability zones. Did they then issue statements of reliability to their prospective and existing customers saying that a single availability zone does not have redundant power, and customers must duplicate functionality in at least 2 AZs to survive a SPOF?
- quantumfissure 5y agoMe: Hesitation at last job moving absolutely everything (including backups) to AWS because if it goes down it's a problem I'm a firm believer in some kind of physical/easily accessible backup. Coworkers: "You're an f'n idiot. Amazon and Facebook don't go down, you're holding us back!" <-Quite literally their words. Me: leaves cause that treatment was the final straw Amazon and Facebook both go down within a month of each other, and supposedly they needed backups Them: shocked pikachu face
- dookahku 5y agoSend Your former colleagues a group email asking how it is
- xwdv 5y agoYou’re still in the wrong, don’t be so smug. These few downtimes are no big deal in the grand scheme of things, and your proposed solution would have been more work and headaches for little to no realizable gains, and not to mention the cybersecurity ramifications. Quite frankly, they are probably glad that you’re gone and not around to gloat about every trivial bit of downtime.
- locallost 5y agoThey're not gloating and also not smug. There's not even a 'hehe' in the post.
- lmilcin 5y agoThink about it this way: 1) Can you make your on prem infrastructure go down less than Amazon's? 2) Is it worth it? In my experience most people grossly underestimate how expensive it is to create reliable infrastructure and at the same time overestimate how important it is for their services to run uninterrupted. -- EDIT: I am not arguing you shouldn't build your more reliable infrastructure. AWS is just a point on a spectrum of possible compromises between cost and reliability. It might not be right for you. If it is too expensive -- go for cheaper options with less reliability. If it is too unreliable -- go build your own yourself, but make sure you are not making huge mistake because you may not understand what it actually costs to build to AWSs level. For example, personally, not having to focus on infra reliability makes it possible for me to focus on other things that are more important to my company. Do I care about outages? Of course I do, but I understand doing this better than AWS has would cost me huge amount of focus on something that is not core goal of what we are doing. I would rather spend that time thinking how to hire/retain better people and how to make my product better. And adding all that complexity of running this infra to my company would cause entire organisation be less flexible, which is also a cost. So you can't look at cost of running the infra like a bill of materials for parts and services. And if there is an outage it is good to know there is huge organisation there trying to fix it while my small organisation can focus preparing for what to do when it comes back up.
- ChrisMarshallNY 5y agoI can't play Borderlands 3 this morning (Epic). Wonder if it's connected?
- ClumsyPilot 5y agoNow that everyone and their dog is on AWS, it is not just 'a website stops working', half the world, from telephones to security doors and Iot equipment, stops working? I am not sure if the movement the cloud has reduced amount of failures, but it definitely has made these failures more catastrophic. Our profession is busy makin the world less reliable and more fragile, we will have our reconning just like the shipping industry did.
- madeofpalk 5y agoall I've noticed is slack was a bit unreliable for a little bit, but i just carried on and otherwise ignored it. my world did not stop working.
- ClumsyPilot 5y agoMy apartment block has a dialing system, that, instead if using a cale that goes to your apartment, relies on IP telephony and calls your mobile phone. It stos working if there is no internet, or your phone is out of battery, or you are not home but your wife is.
- deleted 5y ago[deleted]
- KronisLV 5y agoSame, maybe that was a related issue. Today, on Slack i could not edit messages, could not edit statuses and could not post attachments. Pretty annoying!
- dehrmann 5y agoIt's more like it's making downtimes correlated rather than random. For everything other than urgent communication, I'm not sure if this is a big deal.
- kingsloi 5y agoOf all the AWS outage, my team and I have dodged them all, except this one. 3 instances down and unavailable > Due to this degradation your instance could already be unreachable >:(
- electroly 5y agoFWIW I don't think that message has anything to do with this outage. I think it's just a coincidence that you got some degraded hosts. They didn't send out emails like that for this AZ outage (nor would I expect them to -- that email is for when host machines die).
- exogenousdata 5y agoLooks like the SEC's Edgar website is affected. This is the site the SEC uses to post the filings of public companies. Normally there are a hundred or more company filings in the morning starting at 6am ET. This morning there are two. https://www.sec.gov/cgi-bin/browse-edgar?action=getcurrent https://www.sec.gov/cgi-bin/browse-edgar?action=getcurrent
- JCM9 5y agoAWS didn’t “go down”. They had an outage in one AZ, which is why there are multiple AZs in each region. If your app went down then you should be blaming your developers on this one, not AWS. Those having issues are discovering gaps in their HA designs. Obviously it’s not good for an AZ to go down but it does happen and why any production workload should be architected to have seamless failover and recover to other AZs, typically by just dropping nodes in the down AZ. People commenting that servers shouldn’t go down ect don’t understand how true HA architectures work. You should expect and build for stuff to fail like this. Otherwise it’s like complaining that you lost data because a disk failed. Disks fail… build architecture where that won’t take you down.
- bencoder 5y agoOur API is just appsync (graphql) + lambdas + dynamoDB so, theoretically, we shouldn't have been affected. But about 1 in 3 requests was just hanging and timing out. As others have said, they are not being forthright about the severity of the issue, as is standard.
- dkryptr 5y ago100% agree. I'm actually surprised AWS hasn't built in a Chaos Monkey into their APIs/console so people can test their resiliency regularly if an AZ goes down. edit: of course, AWS does have this: AWS Fault Injection Simulator
- stingraycharles 5y agoBecause then people would complain about AWS being less reliable than Azure / GCP.
- biohax2015 5y agoAWS Fault Injection Simulator does this.
- lljk_kennedy 5y agoIs that what they call us-east-1 nowadays?
- exabrial 5y agoAs an industry, can we please stop making products like vacuums that can't operate unless someone else's computer is working in a field in Virgina? There's literally no reason for it.
- bob1029 5y ago2 of our servers are fucked right now. VOIP services down. Only with AWS and Github do I seem get panicked text messages on my phone first thing in the morning... Our workloads on Azure typically only have faults when everyone is in bed.
- 13daug 5y agoThis S3 how you gonna get you investment back from it
- pkulak 5y agoI used to think it was silly to have your own hardware (like a NAS) in your house. What makes you think you can do it better than AWS? Santa is bringing me a Synology in three days.
- darkstar999 5y agoWhy not both? I just got a Synology NAS and it makes cloud sync dead simple. Now the most important things are on my PC, mirrored on 2 drives in my NAS, and on AWS S3 (or any other cloud storage).
- pkulak 5y agoOh yeah. My plan is to migrate everything to the NAS, then have that back up to Glacier and/or Rsync.net. By S3, do you mean Glacier?
- darkstar999 5y agoI have some in glacier, some in Infrequent Access.
- jorgeudajer 5y ago
- richardfey 5y agoAs far as I understood a whole availability zone went down; today is also the day a lot of people understand why "multi-AZ" matters, so I don't think it's fair to say that services are down because the whole AWS is down.
- whoomp12342 5y agothe cloud is great they said...
- joelbondurant 5y ago
- l0b0 5y agoMeta: I posted a "PyPI is down" link a few days ago, and the post got insta-flagged. Is there some rule about this sort of thing?
- amai 5y agoA problem with log4j/logshell?
- amai 5y agoThank goodness we host all IT services in the same cloud. Imagine the chaos we had if everything would not fail at the same time.
- kemals 5y agoHere is The Internet Report episode on the topic of recent AWS outages that covers outage and root causes: https://youtu.be/N68pQy8r1DI https://youtu.be/N68pQy8r1DI
- joe_chip 5y ago
- tomerbd 5y agoRumble was up all this time.
- j10c 5y agoI also had problem with loading youtube at the same time(for 10-15 minutes) . It looks like a coincidence, but who knows if google uses some of the infrastructure from aws.