8 ms·
Interesting day. I've been on an incident bridge since 3AM. Our systems have mostly recovered now with a few back office stragglers fighting for compute. The b
by time0ut 1y ago
Interesting day. I've been on an incident bridge since 3AM. Our systems have mostly recovered now with a few back office stragglers fighting for compute.
The biggest miss on our side is that, although we designed a multi-region capable application, we could not run the failover process because our security org migrated us to Identity Center and only put it in us-east-1, hard locking the entire company out of the AWS control plane. By the time we'd gotten the root credentials out of the vault, things were coming back up.
Good reminder that you are only as strong as your weakest link.
- shawabawa3 1y agofor what it's worth, we were unable to login with root credentials anyway i don't think any method of auth was working for accessing the AWS console
- kondro 1y agoSure it was, you just needed to login to the console via a different regional endpoint. No problems accessing systems from ap-southeast-2 for us during this entire event, just couldn’t access the management planes that are hosted exclusively in us-east-1.
- nijave 1y agoLike the other poster said, you need to use a different region. The default region (of course) sends you to us-east-1 e.x. https://us-east-2.console.aws.amazon.com/console/home https://us-east-2.console.aws.amazon.com/console/home
- 1970-01-01 1y agoI remember Facebook had a similar story when they botched their BGP update and couldn't even access the vault. If you have circular auth, you don't have anything when somebody breaks DNS.
- crote 1y agoWasn't there an issue where they required physical access to the data center to fix the network, which meant having to tap in with a keycard to get in, which didn't work because the keycard server was down, due to the network being down?
- lenerdenator 1y agoNot to speak for the other poster, but yes, they had people experiencing difficulties getting into the data centers to fix the problems. I remember seeing a meme for a cover of "Meta Data Center Simulator 2021" where hands were holding an angle grinder with rows of server racks in the background. "Meta Data Center Simulator 2021: As Real As It Gets (TM)"
- junon 1y agoYep. And their internal comms were on the same server if memory serves. They were also down.
- simplyluke 1y agoI was there at the time, for anyone outside of the core networking teams it was functionally a snow day. I had my manager's phone number, and basically established that everyone was in the same boat and went to the park. Core services teams had backup communication systems in place prior to that though. IIRC it was a private IRC on separate infra specifically for that type of scenario.
- junon 1y agoThanks for the correction, that sounds right. I thought I had remembered IRC but wasn't sure.
- prmoustache 1y agoI remember working for a company who insisted all teams had to usr whatever corp instant messaging/chat app but our sysadmin+network team maintained a jabber server + a bunch of core documentation synchronized on a vps in a totally different infrastructure just in case and sure enough there was that a day it came handy.
- vladvasiliu 1y ago> Identity Center and only put it in us-east-1 Is it possible to have it in multiple regions? Last I checked, it only accepted one region. You needed to remove it first if you wanted to move it.
- AndrewKemendo 1y agoCorrect. That does make it a centralized failure mode and everyone is in the same boat on that. I’m unaware of any common and popular distributed IDAM that is reliable
- fheisler 1y agoNot sure if this counts fully as 'distributed' here, but we (Authentik Security) help many companies self-host authentik multi-region or in (private cloud + on-prem) to allow for quick IAM failover and more reliability than IAMaaS. There's also "identity orchestration" tools like Strata that let you use multiple IdPs in multiple clouds, but then your new weakest link is the orchestration platform.
- mooreds 11mo agoDisclosure: I work for FusionAuth, a competitor of Authentik. Curious. Is your solution active-active or active-passive? We've implemented multi-region active-passive CIAM/IAM in our hosted solution[0]. We've found that meets needs of many of our clients. I'm only aware of one CIAM solution that seems to have active-active: Ory. And even then I think they shard the user data[1]. 0: https://fusionauth.io/docs/get-started/run-in-the-cloud/disaster-recovery https://fusionauth.io/docs/get-started/run-in-the-cloud/disa... 1: https://www.ory.com/blog/global-identity-and-access-management-multi-region https://www.ory.com/blog/global-identity-and-access-manageme... is the only doc I've found and it's a bit vague, tbh.
- vinckr 11mo agoHey Dan, appreciate the discussion! Ory’s setup is indeed true multi-region active-active; not just sharded or active-passive failover. Each region runs a full stack capable of handling both read and write operations, with global data consistency and locality guarantees. We’ll soon publish a case study with a customer that uses this setup that goes deeper into how Ory handles multi-region deployments in production (latency, data residency, and HA patterns). It’ll include some of the technical details missing from that earlier blog post you linked. Keep an eye out! There are also some details mentioned here: https://www.ory.com/blog/personal-data-storage https://www.ory.com/blog/personal-data-storage
- ttul 1y agoThere is always that point you reach where someone has to get on a plane with their hardware token and fly to another data centre to reset the thing that maintains the thing that gives keys to the thing that makes the whole world go round.
- SOLAR_FIELDS 1y agoThis reminds me of the time that Google’s Paris data center flooded and caught on fire a few years ago. We weren’t actually hosting compute there, but we were hosting compute in AWS EU datacenter nearby and it just so happened that the dns resolver for our Google services elsewhere happened to be hosted in Paris (or more accurately it routed to Paris first because it was the closest). The temp fix was pretty fun, that was the day I found out that /etc/hosts of deployments can be globally modified in Kubernetes easily AND it was compelling enough to want to do that. Normally you would never want to have an /etc/hosts entry controlling routing in kube like this but this temporary kludge shim was the perfect level of abstraction for the problem at hand.
- citizenpaul 1y ago> temporary kludge shim was the perfect level of abstraction for the problem at hand. Thats some nice manager deactivating jargon.
- SOLAR_FIELDS 1y agoYeah that sentence betrays my BigCorp experience it’s pulling from the corporate bullshit generator for sure
- johndubchak 1y ago+1...hee hee
- LPisGood 1y agoManager deactivating jargon is a great phrase - it’s broadly applicable and also specific.
- ct_list 1y ago[dead]
- jordanb 1y agoCouldn't you just patch your coredns deployment to specify different forwarders?
- hinkley 1y agoToo much armor makes you immobile. Will your security org be held to task for this? This should permanently slow down all of their future initiatives because it’s clear they have been running “faster than possible” for some time. Who watches the watchers.
- barbazoo 1y agoWow, you really *have* to exercise the region failover to know if it works, eh? And that confidence gets weaker the longer it’s been since the last failover I imagine too. Thanks for sharing what you learned.
- shdjhdfh 1y agoYou should assume it will not work unless you test it regularly. That's a big part of why having active/active multi-region is attractive, even though it's much more complex.
- ej_campbell 1y agoThat wouldn't have even caught that, most likely unless they verified they had no incidental tie ins with us-east-1.
- jpollock 1y agoThe last place I worked actively switched traffic over to the backup nodes regularly (at least monthly) to ensure we could do it when necessary. We learned that lesson by having to do emergency failovers and having some problems. :)
- ej_campbell 1y agoTotally ridiculous that AWS wouldn't by default make it multi-region and warn you heavily that your multi-region service is tied to a single region for identity. The usability of AWS is so poor.
- skywhopper 1y agoThey don’t charge anything for Identity Center and so it’s not considered an important priority for the revenue counters.
- reenorap 1y agoIt's a good reminder actually that if you don't test the failover process, you have no failover process. The CTO or VP of Engineering should be held accountable for not making sure that the failover process is tested multiple times a month and should be seamless.
- sroussey 1y agoIf you don’t regularly restore a backup, you don’t have one.
- ct520 1y agoI always find it interesting how many large enterprises have all these DR guidelines but fail to ever test. Glad to hear that everything came back alright
- ozim 1y agoSounds like a lot of companies need to update their BCP after this incident.
- ct_list 1y ago[dead]
- saltserv 1y ago[dead]
- michaelcampbell 1y ago"If you're able to do your job, InfoSec isn't doing theirs"
- ransom1538 1y agoPeople will continue to purchase Mutli-AZ and multi-region even though you have proved what a scam it is. If east region goes down, ALL amazon goes down, feel free to change my mind. STOP paying double rates for multi region.