12 ms·
AWS IAM is having issues again
- rootforce 6y agoThis probably needs a better link, but the AWS status page shows everything up. UPDATE: Status page now shows it https://status.aws.amazon.com/# https://status.aws.amazon.com/#
- jjoonathan 6y agoThe AWS status page misses lots of "bump in the night" outages.
- temp0826 6y agoThose statuses are updated by people (after bureaucracy and only with very high level approval), not by anything based in reality
- privacy-matters 6y agoThat is super messed up. That is not how a status board should work.
- QuinnyPig 6y agohttps://stop.lying.cloud https://stop.lying.cloud if you (like me) keep getting the order of the aws, amazon, and status confused.
- tgsovlerkhgsel 6y agoIs this just a gimmick run by AWS and the "honest" in the logo is just a play on the host name, or is it some third-party version that adds information that AWS isn't reporting?
- outworlder 6y agoThird party obviously. Check whois.
- jacques_chester 6y agoThe page footer says "© 2020, Amazon Web Services, Inc. or its affiliates. All rights reserved."
- tgsovlerkhgsel 6y agoDoes it add/edit any information, or is it just a proxy?
- tuananh 6y agoprobably just cname.
- QuinnyPig 6y agoIt's mine. It's an old Chrome extension shoved into a Lambda@Edge function that dynamically transforms the actual status page to cut out a lot of the "sea of green bubbles." I should look into updating it; it used to automatically upgrade the severity one level as well but AWS changed something. https://gaslighting.me https://gaslighting.me is the non-editorialized version.
- BillinghamJ 6y agoLooks to be affecting all regions - at least within the standard aws partition. Not sure about aws-cn and aws-us-gov
- riffic 6y agoIAM is a global service.
- bythe4mile 6y agoIAM is a global service. So it kind of makes sense that all regions would be affected by an outage.
- BillinghamJ 6y agoWell yes but it's entirely possible for it to be locally affected
- bythe4mile 6y agoIAM is seems to be run out of North Virginia (other than Gov Cloud)[0]. However they probably may be doing some magic under the covers to reduce latency across the globe. It would also explain why GovCloud isnt affected. [0] - https://aws.amazon.com/about-aws/global-infrastructure/regional-product-services/#AWS_Identity_and_Access_Management_.28IAM.29 https://aws.amazon.com/about-aws/global-infrastructure/regio...
- BillinghamJ 6y agoYeah certainly the existence of credentials etc must be replicated to deal with cross-region connectivity issues etc. - which I guess explains why authentication is still working fine GovCloud I believe is technically separated in this regard, I don't think IAM credentials can operate across partitions. Though it does seem to be possible to link the existence of a standard AWS account with GovCloud, the same is not true of the China partition
- ArchOversight 6y ago
- holler 6y agomy site is down on all environments (us-east-1). oy vey
- leesalminen 6y agoIs your site dependent on IAM?
- deleted 6y ago[deleted]
- andrewxdiamond 6y agoPretty hard to not be dependent on IAM. Authentication and authorization are some of the most core concepts you can have
- cperciva 6y agoAmazon says that "The issue continues to affect create, describe, modify or delete of IAM accounts and roles. [...] Authentication using IAM accounts and roles are not affected." So it's entirely possible to depend on IAM but not in a way which this is breaking.
- andrewxdiamond 6y agoFor sure, that’s just not what the parent comment asked
- sk5t 6y agoThis makes sense, observing services that rely on IAM/STS very, very routinely--but without changing IAM properties--and no alerts popped up during this outage.
- temp667 6y agoI thought IAM authentication was working, but the create / remove / etc steps were down. This is a HUGE impact if all IAM role access is down globally for AWS. Please confirm and I will work to spread the news - can I cite you as the source?
- tus88 6y agoIt sirtainly is. And I came in this morning specifically to creating some Lambda roles to test. Fark.
- TazeTSchnitzel 6y agoOff-topic: I hadn't heard of nitter.net before, it seems pretty cool.
- rootforce 6y agoI like it a lot better than the current twitter experience, and it's open source.
- boring_twenties 6y agoIt's self-hostable, too
- DC-3 6y agoIt makes me sad to see these reminders of how fast websites are perfectly capable of being.
- dang 6y agoSorry to disappoint, but we've changed the URL from https://nitter.net/RyanGartin/status/1306352941964701696#m https://nitter.net/RyanGartin/status/1306352941964701696#m to the original source, as the site guidelines ask: https://news.ycombinator.com/newsguidelines.html https://news.ycombinator.com/newsguidelines.html.
- rootforce 6y agoYep, sorry about that. Thanks for the update.
- gscho 6y agoLet's all move to serverless!
- banana_giraffe 6y agoWhat a perfect time to onboard some new employees. Blah.
- jaaames 6y agoI guess I'm going outside today.
- runawaybottle 6y agoWhat’s everyone’s back up plans? Just a ‘We’ll be right back’ page?
- snazz 6y agoSwitching to that sort of page is probably the most cost-effective solution for small businesses. I've heard of larger companies running a completely redundant hot standby on another independent cloud platform and switching DNS over to the standby when something goes wrong. With auto-scaling, you're not paying to have the standby running at full throttle. Of course, you have to exclusively use services that have equivalents on the other cloud provider.
- redis_mlc 6y ago> What’s everyone’s back up plans? Just a ‘We’ll be right back’ page? I've done a lot of cloud site HA work on large sites. Waiting out the cloud outage so far ends up being the best solution for almost all companies, from both engineering and business standpoints. Eat the outage, but continue with a known-working site afterwards. You just blame the cloud provider for the downtime. > I've heard of larger companies running a completely redundant hot standby on another independent cloud platform and switching DNS over to the standby when something goes wrong. In theory, that makes sense. It practise, it almost never works. If by "independent cloud platform" you mean another AZ or region in the same cloud, that is often attempted and can work reasonably well. If you mean failover from AWS to GCP, then that's unlikely, since everything is different. An example is whenever DynDNS goes down for 2-3 hours, and everybody builds out flaky failover tools that are less reliable than their original DNS provider - and have to be maintained forever. Might work with one or two domains, becomes a huge ongoing problem with dozens of them. Also, DNS mgmt. APIs are flaky in several dimensions (availability, versioning, parameters, etc.) Another is that you can't failover to another location that doesn't have all your data, current certificates, monitoring, etc., and the failover site needs the capacity of the original site to work. That costs ongoing money and time, and you never know how well the failover will work or how it will perform. An example is Heartland, one of the biggest US payment providers, who failed over to another location and took a 5 day outage. Or gitlab, who took a one day outage because of database isues. I have (automatically) failed over a large site for a publicly-traded company from one AWS region to another, but that took a year of work to setup, and I understood almost all aspects of the site. Afterwards, I realized that almost nobody really has time to organize that either at a conceptual or engineering level. And organizations don't recognize Herculean efforts like that, so think twice beforehand. Key point: always involve your DBA from the beginning when doing a project like this.
- sebmellen 6y agoI can always rely on HN for finding out why AWS is broken... Is it time to switch to GCP?
- bfieidhbrjr 6y agoAzure. The web panel is a bazillion times better and they know AD... pretty well.
- dijit 6y agoI think this is sarcasm. But genuinely AzureAD is really good. Authentication and Authorization systems are difficult, AD has always been begrudgingly the most well rounded (yes, it has warts) but AzureAD really is nice. I especially like that my local admin can delegate Enterprise Apps to me so I can create SAML/OAUTH2 SSO links between stuff we use without needing the keys to the kingdom. I recently set up Enterprise Federation with GCP and it took less than a day. (compared with many months in my last company which used on-prem AD)
- sargun 6y agoDoes azure AD actually do active directory yet (LDAP and all)? Or is it still just a name?
- dijit 6y agoAll of our windows desktop machines use it as their login controller, but I’m not 100% sure of the implementation- it is likely there needs to be a local login server because from what I remember of Windows it’s using WINS to find a logon server, and that I think must be local.
- easton 6y agoIt’s actually a new-ish “credential provider” system for logging into Windows. It’s not just Azure AD either, Gsuite can be used for logging into and managing Windows now. There is no longer a need for a local domain controller either. https://docs.microsoft.com/en-us/windows/win32/secauthn/credential-providers-in-windows https://docs.microsoft.com/en-us/windows/win32/secauthn/cred...
- peterthehacker 6y agoIt’s funny how on Twitter this currently has 3 retweets and 17 likes but on HN it has 110 points!
- specialp 6y agoI started seeing this while running Terraform and immediately went to Twitter to see if it is down for everyone or just me. Same experience. From now on I will go to HN :)
- kohtatsu 6y agohttps://hckrnews.com https://hckrnews.com is an alt-reader that sorts the front page chronologically which helps.
- simonkafan 6y agoMaybe because the audience on Twitter and Hackernews is different?
- dvtrn 6y agoDid this link get changed to a twitter post from something else?
- rootforce 6y agoYes, it was a link to nitter.net(an alternative twitter front end) due to HN guidelines for posting links to original sources.
- panny 6y agoI noticed (pre outage) IAM console won't work at all if I --disable-reading-from-canvas in my launch args to prevent fingerprinting. All the other service consoles I use work. I have to have a special config for my browser just for AWS because of it. Wishful thinking, but maybe they're fixing that just for me.
- deleted 6y ago[deleted]
- TheYahiaBakou 6y agoI just feel bad for the oncalls...
- igetspam 6y agoWherever possible, don't use us-east-1. It's one of the older regions and parts are aging. Yes, I know there are things that are only available in the old regions but most services are globally available. I've worked with a few ex AWS SWEs and SREs. They drink the kool-aid and won't say anything bad about us-east-1 but they also won't launch net-new services there. YMMV
- nerdbaggy 6y agoSo for east coast you would say go Ohio?
- PopeDotNinja 6y agoIf that's they case, maybe AWS should raise the pricing on us-east-1 to fund improvements and/or encourage people to either move off to newer data centers.
- Cthulhu_ 6y agoThey're one of the wealthiest companies in the world and you suggest they should raise prices to raise money to improve one of their older money printing locations? What planet are you from? They can rebuild that whole datacenter a hundred times over and still not feel it in their wallet.
- Aperocky 6y ago> things that are only available in the old regions And usually it's old things, like really really old if something is available in us-east-1 and not us-east-2. I would imagine some ancient ec2 instance types.
- s09dfhks 6y agoI modified some of our IAM policies earlier this afternoon, followed by the pages that some of our teams were having IAM issues, caused me great discomfort
- rootsudo 6y agoCongratulations, this time it wasn't you! Maybe, possibly, hopefully...
- NathanWilliams 6y agoNoticed something odd today I think is connected to this. The other day we started using Access Advisor, and we found some of our KMS key policies with a Principal of '*'. It wasn't marked as globally open, so we planned to fix them a little later. This morning we found that status had changed. While we were in the wrong to begin with, it was a little surprising to find the interpretation of the key policy changing overnight. Of course it became our top priority and is now fixed. Something to look out for...