5 ms·
A much better UX would be clear error messaging informing users that the service is down and there is no problem with their individual account. This would prev
by mfrommil 3y ago
A much better UX would be clear error messaging informing users that the service is down and there is no problem with their individual account.
This would prevent people from panicking they've been hacked and/or unnecessarily resetting their password.
- eurekin 3y agoThere was for a brief moment. I got that once
- Kalium 3y agoYou are absolutely correct. That would be a much better experience. That said, getting there strikes me as pretty challenging. Automatically detecting a down state is difficult and any detection is inevitably both error-prone and only works for things people have thought of to check for. The more complex the systems in question, the greater the odds of things going haywire. At Meta's scale, that is likely to be nearly a daily event. The obvious way to avoid those issues is a manual process. Problem there tends to be that the same service disruptions also tend to disrupt manual processes. So you're right, but also I strongly suspect it's a much more difficult problem than it sounds like on the surface.
- matsemann 3y agoBut there's something off here. I wouldn't expecting to be shown as logged out when the services are down. I'd expect calls to fail with something aka 500 and an error showing "something happen edited on our side". Not all the apps going haywire.
- Kalium 3y agoAt the scale of Meta, "down" is a nuanced concept. You are very unlikely to get every piece of functionality seizing up at once. What you are likely to get is some services ceasing to function and other services doing error-handling. For example, if the service that authenticates a user stops working but the service that shows the login form works, then you get a complex interaction. The resulting messaging - and thus user experience - depend entirely on how the login page service was coded to handle whatever failure the authentication service offered up. If that happens to be indistinguishable from a failure to authenticate due to incorrect credentials from the perspective of the login form service, well, here we are. At Meta's scale, there's likely quite a few underlying services. Which means we could be getting something a dozen or more complex interactions away from wherever the failures are happening.
- shkkmo 3y ago> If that happens to be indistinguishable from a failure to authenticate due to incorrect credentials from the perspective of the login form service, well, here we are. If you can't distinguish those, then that is bad software design.
- sandspar 3y agoSlightly unrelated question, but just how "Big" is Meta? I know it's vast, but as an outsider I have trouble grokking the scale of it.
- pixl97 3y agoWhen most people talk about serving thousands and maybe millions of requests per second, Meta talks about billions of requests per second. https://read.engineerscodex.com/p/how-facebook-scaled-memcached https://read.engineerscodex.com/p/how-facebook-scaled-memcac...
- jessriedel 3y agoIsn't this just the standard problem of reporting useful error messages? Like, yes, there are academic situations where you can't distinguish between two possible error sources, but the vast majority of insufficiently informative error messages in the real world arise because low effort was applied to doing so.
- Kalium 3y agoYes and no. Yes, with the additions of sheer scale, a vast number of services, multiple layers, and the difficulty of defining "down" added in. I think the difficulty of reporting useful error messages is proportional to the number of places an error can reasonably happen and the number of connections it can happen over, and by any metric Meta's got a lot of those. No, in that detecting when you should be reporting a useful error message is itself a complex problem. If a service you call gives you a nonsense response, what do you surface to the user? If a service times out, what do you report? How do you do all this without confusing, intimidating, and terrifying users to whom the phrase "service timeout" is technobabble?
- lanstin 3y agoCome on use a little imagination. DNS lookup for the db holding the shard with the user credentials disappears. Code isn’t expecting this, throws a generic 4xx because security instead of a generic 5xx (plenty of people writing auth code will take the stance all failures are presented the same as a bad password or non-existing username); caller interprets this a login failure. Same auth system system used to validate logins to the bastions that have access to DNS. Voilá.
- shkkmo 3y ago> plenty of people writing auth code will take the stance all failures are presented the same as a bad password or non-existing username Those people would be wrong. You can take all unexpected errors and stick them behind a generic error message like "something went wrong" but you should not lie to your users with your error message.
- jtuple 3y agoIt's about not leaking sensitive information. If you have different messages for invalid username vs invalid password, you can exploit that to determine if a user has an account at a particular service. "Invalid credentials" for either case solves this problem. But sure, let's report infra failures different as "unexpected error" Now, what happens if the unexpected error is only when checking passwords, but not usernames? Do you report "invalid credentials" when given an invalid username, but "unexpected error" when given a valid name but invalid password? If so, you're leaking information again and I can determine valid usernames. So, safe approach is to report "invalid credentials" for either invalid data or partial unexpected errors. Only time you could safely report "unexpected error" is if both username check and password check are failing, which is so rare that it's almost not worth handling. Esp. at the risk of doing wrong and leaking info again.
- ambichook 3y agowhat if, if one service doesnt respond at all or responds with something that doesnt fit an expected format that it would if working correctly, the whole thing just says "sorry, we had an error, try again later"? if it has to check both at the same time, and cant check them independently, wouldn't that solve the vulnerability? or am i missing something? totally understandable if i am, i just want to learn /gen
- seppel 3y ago> That said, getting there strikes me as pretty challenging. Automatically detecting a down state is difficult and any detection is inevitably both error-prone and only works for things people have thought of to check for. The more complex the systems in question, the greater the odds of things going haywire. At Meta's scale, that is likely to be nearly a daily event. Well, in principle, the frontend just has to distinguish between HTTP status 500 (something broken in the backend, not the fault of the user) and some HTTP status code 4xx (the user did something wrong).
- Kalium 3y agoYes, assuming the responses are usefully different, accurate, and you get responses in a timely manner.
- seppel 3y agoThe "your username/password is wrong" message came in a timely manner. So someone transformed "some unforeseen error" into a clear but wrong error message. And this caused a lot of extra trouble on top of the incident.
- boring_twenties 3y agoWell you can't expect to hire engineers with half a brain for the pitiful compensation Meta offers, can you?
- barbazoo 3y agoNot the worst thing that a bunch of Facebook users are resetting their passwords.
- sandspar 3y agoThat's the sound of millions of "password" becoming "passwordnew"
- mns 3y agoThere are quite some harsh comments here below. You can't plan for every possible failure point, who knows what part of a system/infra out of everything that they have went down and triggered this behaviour. Some things you just can't catch/predict. Especially in huge systems like theirs. I would expect people here to understand things like these and not just call people names for something like this, we all know things seem simple/clear from the outside, but the job of debugging and fixing something like this take quite some effort.
- anigbrowl 3y agoThis is a company with one of the largest digital infrastructures in the world. An outage is understandable, inability to tell they're having an outage and inform users appropriately is not. Stop making excuses for people who are literally awash in resources.
- dylan604 3y agoIt is always better for the company's rep for the issue to have been on your end. Admitting fault comes with a potential liability. It's gaslighting written as an SLA
- edanm 3y ago> Stop making excuses for people who are literally awash in resources. This is a pretty weird outlook to have - looking at any group awash with resources, whether it be governments or other companies, and you can clearly see that even with those resources, failures still happen. You can jump up and down and pretend that this is solvable, or you can look at reality, look at all the evidence of this happening over and over to almost everyone, and conclude with some humility that these things just happen to everyone. (Looking this reality in the face is one of the things motivating my beliefs around e.g. AI safety, climate change, etc.)
- anigbrowl 3y ago> An outage is understandable
- jonnycomputer 3y agoIt would be a better UX, but, depending on the outage, that might be a really hard behavior to guarantee.