5 ms·
This should be a huge shitshow for Microsoft. This never should have happened. Azure AD taking down pretty much the entirety of Azure / O365 / Teams is frankly
by king_magic 6y ago
This should be a huge shitshow for Microsoft. This never should have happened.
Azure AD taking down pretty much the entirety of Azure / O365 / Teams is frankly inexcusable. Astounding incompetence, and an astonishing single point of failure that needs to be re-architected.
- macintux 6y agoIn fairness to Microsoft, there's no way to make authentication not a single point of failure. You have to make the service as resilient as possible, which clearly they failed at, but you can't very well fail over from your AD service to something else.
- king_magic 6y agoMicrosoft is rolling back some update they pushed that borked Azure AD. Early signs are not pointing to this as a service resilience issue - this appears to be sheer incompetence - pushing an update that was likely not tested well enough that broke pretty much everything. More than that, why aren't updates to Azure AD being rolled out regionally? Why is Azure AD architected in such a way that the entire thing going down can do so much damage? It's pretty hard to be understanding with a preventable screwup with this kind of global scale.
- adflux 6y agoYou are making a whole bunch of assumptions here without ANY info about how Microsoft rolls out updates to its Azure services.
- king_magic 6y agoSure, but those assumptions are largely backed up by a single update to Azure AD knocking it out worldwide. End of the day, whatever they pushed was not nearly well tested enough, and they paid for it with a massive global outage that borked pretty much everything they sell.
- regecks 6y agoTo add another datapoint, we recently faced an issue where Microsoft rolled out a change to their OIDC endpoints, which broke our client's tenant, but not other tenants. (It was crashing with an HTTP 500 if the Accept request header was a certain value: the default value in OpenJDK 8). Some days later and after our complaints, Microsoft rolled back the change. I would call this evidence that they do practice gradual deployments of changes. I would guess that this outage was caused by something more complex than MS not slow-rolling their deployment.
- extrapickles 6y agoFor something as critical as auth, I would expect most changes to be rolled out over a few days just to make sure something like this cannot happen to everyone. Ideally a change would be shadowed for a few % of traffic before it was allowed to start rolling out.
- hsbauauvhabzb 6y agoThey could quite easily have progressive migrations, blue/green deployments, or a failover environment..
- macintux 6y agoTrue. I guess it’s semantics whether to call that resiliency within a service or failover to another service.
- wallacoloo 6y ago> In fairness to Microsoft, there's no way to make authentication not a single point of failure. Huh? x509-type authentication works without a single point of failure. I can authenticate myself to any client via public-key cryptography without any external dependencies, after we've agreed on a root of trust. Active Directory chose a solution that requires centralized availability. They didn't have to do it this way (though it does make certain administrative tasks, like revocation, simpler). Note that even AD itself does not require a single point of failure though! You can define secondary (i.e. failover) DCs (usually via DNS). Then DNS becomes your point of failure -- and again most DNS implementations support failover, such that any single point of failure gets pushed lower to the stack (e.g. NICs), which also support failover or multipathing, etc, if you really need high availability. Anyway, it's totally possible to make authentication not be a single point of failure. Of course, you have to make some tradeoff to do it (see: the CAP theorem).
- macintux 6y agoMy point was poorly stated, based on one line of the GP comment: > Azure AD taking down pretty much the entirety of Azure / O365 / Teams is frankly inexcusable My objection was that the Microsoft Azure AD service is down, there's nothing that Microsoft can do to protect the services that rely on it. Setting aside your valid x509 counter-argument, the remaining options you listed to increase availability, to me, have to do with resilience of the Azure AD service. They're not alternative services. It depends on where you draw the boundary around the service. Again, clearly Microsoft has failed to provide sufficient resiliency, not arguing against that at all. But once the service is down, practically by definition Teams and Exchange and everything else is doomed.