2 ms·
I work at a big tech company (not MS, though). The short answer is it really all comes down to prioritization. In my experience, we do have the same kinds of t
by Kilenaitor 5y ago
I work at a big tech company (not MS, though). The short answer is it really all comes down to prioritization.
In my experience, we do have the same kinds of tools for surfacing those issues. Granted, they do tend to differ per function or even per org, although there are some standardized ones you can plug into to get e.g. issue aggregation, stack traces + metadata, etc. Granted... a lot of the time that requires a team to 1. actually plug into this infrastructure (otherwise it can either bubble up as an unknown exception or, worst case, get swallowed by a try/catch) and 2. actually pay attention to things once they do plug in.
The other problem is not just the high scale of traffic (or 5XXs or errors in general) but that high traffic is usually due to a myriad of products, features, etc. And there's just simply too much work for too few people. So, we prioritize based on severity of the inherent issue (e.g. privacy/security issues go to the top) or, if that's equivalent, number of affected users take priority. Which just means low-severity, low-impact bugs almost never get fixed. Especially if they have a workaround that can be surfaced e.g. when someone reports the issue to support.
Happy to elaborate further if that doesn't make sense. I also want to clarify that this explanation isn't meant to be some perfectly valid reason as to why this kind of work doesn't get prioritized. I agree a lot of times big companies fall short here, including where I work. This is just one of the many tradeoff that get made when there is more work than people.