4 ms·
Something that can be extremely useful as well in situations like this, is doing API brownouts, toward the later stages of the process. Disable the API for shor
by pocketarc 2y ago
Something that can be extremely useful as well in situations like this, is doing API brownouts, toward the later stages of the process. Disable the API for short stretches of time, on the way to disabling it entirely, to give consumers who might not be keeping up with changes a way to be alerted (they'll notice the downtime).
- simonw 2y agoI really like the API brownouts trick. GitHub have been doing this for years, a few examples: - https://developer.github.com/changes/2018-11-05-github-services-brownout/ https://developer.github.com/changes/2018-11-05-github-servi... - https://github.blog/changelog/2021-08-10-brownout-notice-api-authentication-via-query-parameters-for-48-hours/ https://github.blog/changelog/2021-08-10-brownout-notice-api...
- mhink 2y agoI was thinking about this kind of thing. Another idea might be to introduce artificial latency and gradually increase it over time? Maybe dial up rate-limiting? I'm not really sure if this is a better or worse idea, though.
- pocketarc 2y agoThat is a good idea, but I don't think it does enough to serve the purpose - the point of brownouts is to trigger their error/alert system. If the API is just being slower than usual, it won't trigger anything. Even if a human was reviewing it manually (which is quite unlikely), they would only think "oh, their API's really getting slower these days, sad". There'd be nothing that indicates "The API is going to get shut down in a month and I need to move off of it ASAP!". Random, intermittent API failures would lead you to go check the API status out, and in the process you'd find out "oh, this API is going away". Edit: On the point of rate limiting, I think the problem with it is that it'll affect everyone using the API all the time, not just during the brownout period. It effectively shuts the API down for everyone still using it (if the rate limit is too low, and if it isn't, then it won't be noticed by low-use consumers).
- hooverd 2y agoHey, some of us look at P95 latency.
- Joel_Mckay 2y agoOr instead of losing paying customers: 1. design API accounts to include a preferred server and default server (handy to explicitly load balance, or dynamically bounce users to specified servers.) 2. design clients to have a timed service lifecycle (expiry 2 weeks prior to cert expiry). Then enter a semi-dormant mode until valid signed updates succeed. 3. add a random timed daily update check, and begin reassigning the users to the new API after updates install properly. Also, warn users the migration will happen 2 months before it is set to launch (do a few random A/B tests the first day). 4. Never rely on people to act, or not act for keeping infrastructure running. You are not going to be able to manually update 30000 legacy hosts with a single team. Worst case scenario you must auto-reconfigure the clients for a standalone offline mode... so the next user of the IP doesn't get hammered by failed connection retry attempts. Brownouts won't work for cached-edge systems designed to reconcile month long intermittent outages. i.e. systems that were designed to handle DDoS, worms, and acts of clod... Have a nice day, =3
- bogdan-lab 2y agoYes, I completely agree. The story sounds like fairytale: migration was announced to happen in 3 months and in 4 months it was done by removing old API. Where are all those clients, who are happy with current API performance and do not want to spend their money on making API owners life better? What happened to them? Did the company just decide to let those clients go?
- Joel_Mckay 2y agoCollecting clients other people hosed is an easy business. Except, entrenched incompetence may still pine for the convenience of a quick sometimes-broken kludge (some folks expect everything to be glitched half the time). I definitely understand why some techs just stop caring about customer opinions. You'll know when you are in a senior role when one starts to fantasize about being a Plumber. =3
- simonw 2y agoYeah, one of the biggest problems with API deprecation is that you have zero control over the roadmap of your clients. If they can't spare the engineering time in the next six months to carry out the upgrade (and you aren't 100% mission critical to their business) they're not going to do that no matter how much you bug them. Depending on how old the integration is they may not even still employ the engineers who built the first version, which makes it even harder for them to roadmap the work.