3 ms·
I am proud of my team at AWS that took backwards compatibility very seriously. Even introducing a temporary backwards incompatibility was a no-go in design revi
by redditor98654 4y ago
I am proud of my team at AWS that took backwards compatibility very seriously. Even introducing a temporary backwards incompatibility was a no-go in design reviews.
We had a service that had a list API that was paginated. It returned a nextToken to specify the start of the next page of the results.
Internally we were doing a database migration to a completely different system and migrating one customer at a time. The problem was that if a customer was in the middle of a list call and had the next token with them which was generated from the previous database system, after migrating to the new database the older token would not be able to start from exactly where it should had the customer not been migrated. This was because it would not have all the information of the service to resume at the exact offset.
One option was to throw an error and let the customer retry the request; another option was to return some possibly duplicate items in the next page; none of these were good enough for both engineers and PMs and instead we decide to take up a bunch of additional work so that no customer would be impacted. This was 10+ weeks of additional work for the whole team but we did it because culturally it felt the right thing to do for the customer.
Note that the impact would have been tiny if at all. A customer would have to be in the middle of a paginated request and out migration system would have had to migrate that particular customer at that exact time and the impact would have been a few possibly duplicate items. But we didn’t know the actual impact of those temporary duplicated for a single call and we all agreed breaking changes like this are unexpected and cause customer to lose trust with us.
- jiggawatts 4y agoI read stuff like this about AWS and I wonder what the guys over at Azure are smoking. I can deploy trivial architectures with ordinary resources and hit four or five “broken by design” issues and at least one or two Heisenbugs that aren’t reproducible in other identical setups due to fun things like “legacy” subscriptions — an invisible internal setting.
- NickGerleman 4y agoI previously worked at Microsoft in an unrelated area, but had a friend who worked on CosmosDB, and later a different part of Azure. There are some Microsoft products I genuinely love, but some are terrible. To an extent it is a reflection of the inconsistency in internal teams. Culture, values, skill level, and quality bar are all over the place depending on who you talk to, even compared to other large companies. From what I had heard, CosmosDB was not a healthy team, and I would not consider using it as a product.
- tkk23 4y agoWith the funding of MS, how could this be turned around? Is it necessary to build a new team or is it enough to exchange leadership and let them bring in new members?
- NickGerleman 4y agoIt's a good question well above my pay grade :). Changing culture isn't an easy problem. I suspect the way Microsoft does interviewing and performance management (very local to the specific team) contributes to the inconsistency. MSFT has also been fairly open to its employees that it does not try to compete with competitors like Google, Meta, or even Amazon, in terms of compensation. So it isn't really trying to get the best engineers, so long as it can continue to print money. There are still folks there who are incredible, but the floor is shockingly low at times. Folks will self-select, so you will then get teams which are more homogeneously good or bad.
- aoetalks 4y ago>I suspect the way Microsoft does interviewing and performance management (very local to the specific team) contributes to the inconsistency. I think pay is certainly a factor in talent bleeding to G/Meta in general, which they refuse to address. I imagine this isn't unique to MSFT.
- stingraycharles 4y agoIt’s interesting that Microsoft doesn’t have this type of attitude as well for Azure, as they’re famous for their extreme care around backwards compatibility for Windows as well. Perhaps it’s because Microsoft is still catching up with Azure, and as such prefers moving fast and occasionally breaking things?
- aoetalks 4y agoI work for Azure (different product) and we do care about breaking customer. We have to keep GA APIs around for years even as they’re being deprecated. The AzureRM module was deprecated over a year ago I think, and it will work until 2024. This really feels like a bug to me, and probably didn’t trip monitors due to 2 reasons: 1) Given that portal does the right thing (and probably ARM template samples), this was very small percentage of traffic. 2) the failures would look like client side errors, making it less likely to trip monitors. *PS my comment is not an official response (I don’t even remotely work on CosmosDB) but I’ll forward this internally