8 ms·
Deactivating an API, one step at a time
- pocketarc 2y agoSomething that can be extremely useful as well in situations like this, is doing API brownouts, toward the later stages of the process. Disable the API for short stretches of time, on the way to disabling it entirely, to give consumers who might not be keeping up with changes a way to be alerted (they'll notice the downtime).
- simonw 2y agoI really like the API brownouts trick. GitHub have been doing this for years, a few examples: - https://developer.github.com/changes/2018-11-05-github-services-brownout/ https://developer.github.com/changes/2018-11-05-github-servi... - https://github.blog/changelog/2021-08-10-brownout-notice-api-authentication-via-query-parameters-for-48-hours/ https://github.blog/changelog/2021-08-10-brownout-notice-api...
- mhink 2y agoI was thinking about this kind of thing. Another idea might be to introduce artificial latency and gradually increase it over time? Maybe dial up rate-limiting? I'm not really sure if this is a better or worse idea, though.
- pocketarc 2y agoThat is a good idea, but I don't think it does enough to serve the purpose - the point of brownouts is to trigger their error/alert system. If the API is just being slower than usual, it won't trigger anything. Even if a human was reviewing it manually (which is quite unlikely), they would only think "oh, their API's really getting slower these days, sad". There'd be nothing that indicates "The API is going to get shut down in a month and I need to move off of it ASAP!". Random, intermittent API failures would lead you to go check the API status out, and in the process you'd find out "oh, this API is going away". Edit: On the point of rate limiting, I think the problem with it is that it'll affect everyone using the API all the time, not just during the brownout period. It effectively shuts the API down for everyone still using it (if the rate limit is too low, and if it isn't, then it won't be noticed by low-use consumers).
- hooverd 2y agoHey, some of us look at P95 latency.
- Joel_Mckay 2y agoOr instead of losing paying customers: 1. design API accounts to include a preferred server and default server (handy to explicitly load balance, or dynamically bounce users to specified servers.) 2. design clients to have a timed service lifecycle (expiry 2 weeks prior to cert expiry). Then enter a semi-dormant mode until valid signed updates succeed. 3. add a random timed daily update check, and begin reassigning the users to the new API after updates install properly. Also, warn users the migration will happen 2 months before it is set to launch (do a few random A/B tests the first day). 4. Never rely on people to act, or not act for keeping infrastructure running. You are not going to be able to manually update 30000 legacy hosts with a single team. Worst case scenario you must auto-reconfigure the clients for a standalone offline mode... so the next user of the IP doesn't get hammered by failed connection retry attempts. Brownouts won't work for cached-edge systems designed to reconcile month long intermittent outages. i.e. systems that were designed to handle DDoS, worms, and acts of clod... Have a nice day, =3
- bogdan-lab 2y agoYes, I completely agree. The story sounds like fairytale: migration was announced to happen in 3 months and in 4 months it was done by removing old API. Where are all those clients, who are happy with current API performance and do not want to spend their money on making API owners life better? What happened to them? Did the company just decide to let those clients go?
- Joel_Mckay 2y agoCollecting clients other people hosed is an easy business. Except, entrenched incompetence may still pine for the convenience of a quick sometimes-broken kludge (some folks expect everything to be glitched half the time). I definitely understand why some techs just stop caring about customer opinions. You'll know when you are in a senior role when one starts to fantasize about being a Plumber. =3
- simonw 2y agoYeah, one of the biggest problems with API deprecation is that you have zero control over the roadmap of your clients. If they can't spare the engineering time in the next six months to carry out the upgrade (and you aren't 100% mission critical to their business) they're not going to do that no matter how much you bug them. Depending on how old the integration is they may not even still employ the engineers who built the first version, which makes it even harder for them to roadmap the work.
- michaelt 2y ago> It might be that you want to replace it with a new, more capable version If you're truly replacing your API with a new, more capable version there's a much better option, in my experience. Roll out your new API, and replace your old API's implementation with a proxy that calls through to the new API. The proxy will need very little maintenance, as all it's doing is connecting one fixed, stable API (your old one) to another fixed, stable API (your new one). Lock it down to only your old customers, if you want. The support costs will be basically zero, and your existing paying customers will thank you for respecting their time with a dependable, low-churn API.
- n4r9 2y agoWhat if the new API is more capable precisely because the parameters have been completely redesigned?
- cinntaile 2y agoSurely you can keep it backwards compatible AND give people access to the new good stuff?
- smallnix 2y agoNot if the API reflects a fundamental change in flows. E.g. not fun to proxy a sync API into the new async one which splits a single operation into three.
- AlotOfReading 2y agoThat's a reason not to make fundamental changes to your flows, but the example given isn't that difficult. It's probably how the developer on the customer's side is going to deal with the issue, so you may as well do it for them. Trying to convince any meaningful percentage of customers to rewrite fundamental aspects of their existing integration to support your product roadmap isn't going to happen, so don't expect it.
- petesergeant 2y ago> In addition to offering human-understandable communication, I asked the API producer to add the Deprecation HTTP header field to all responses Cute, but, I question the value
- bpedro 2y agoGreat question. As a consumer, you can set up an alert when any of your API requests has a deprecation or sunset HTTP header. You'd know immediately if any of the APIs you depend on is about to be deactivated.
- compootr 2y agoYou're assuming downstream companies will parse & utilize this header If they were competent enough to do that, they'd probably pay attention to updates sent via, say, email. Brownouts would make people notice, like "oh shit, $API doesn't work!"
- chipdart 2y ago> Cute, but, I question the value I came here just to say that. What an half-baked idea. It might be trivial to mindlessly bolt on response headers, but if the goal is that the mechanism needs to be impactful and have consequences then the response header is a big red herring as you're actually relying on clients to implement support for sunsetting the endpoints. If that's the case then you already have meaningful mechanisms, such as passing this sort of metadata in responses to requests for the root resource. This is something that pretty much any HATEOAS spec already supports.
- lazyasciiart 2y ago> If that's the case then you already have meaningful mechanisms, such as passing this sort of metadata in responses to requests for the root resource. This is something that pretty much any HATEOAS spec already supports. Yes, they support it with headers.
- athenot 2y agoOther strategies that could come in handy before completely shutting down the API: - Rate-limit the API, with increasing aggressiveness until you're down to 0 requests per unit of time; - Introduce latency in serving the requests (assuming your edge can handle the increased volume of open connections). Both of these introduce gradual degradation of the old API, without outright killing the business functionality that recalcitrant customers are nevertheless reliant upon. It helps spur a bit of urgency to switch to the new API, while remaining nice. Many (enterprise) customers will wait until the last minute to switch: essentially they are having to put in work without any tangible feature gain—at least from their perspective. Regardless of the strategy, one other point I'd add is to monitor actual use of the API; if important customers are still actively using the old API, it would be unwise to shut it off.
- croemer 2y agoI think GitHub actions cause errors for a few minutes to alert, one could do that also with an API.
- Suppafly 2y agoMan I really hate the idea of "Let's make a thing that works shittier so that people switch to the new thing". If you're in a situation where you have customer's just be honest about the changes coming and give them enough runway to get the changes made.
- beachy 2y agoYep. If you gradually degrade performance, and word of your intentions fails to make it to the right people, the victim may well have to devote significant effort to trying to work out what's going on. Boy will they be pissed when they find out.
- bongodongobob 2y agoThis is standard practice though. You can only communicate it so much. If your customers ignore it, that's on them. You either have a hard cutoff date or you introduce failures or throttling. The latter is friendlier.
- its_ethan 2y agoIt would be really cool to see a graph of the API usage over time with markers showing when each "stage" was occurring. I'm wondering if there were significant dips shortly after each stage, or if it was more of a gradual decline? It would also be interesting to know if the API was still being used up until the final stage? Were there any ramifications/ angry customers at the door after that?
- bpedro 2y agoThese are great suggestions. I'll write about it the next time I'm in a similar situation.
- bornfreddy 2y agoGiven the timeline of just 4 months, I would expect 90% of customers not migrating in time. The percent of those who migrated to the new API after that depends on how vital this service is to them, it would be interesting to see the rate of lost businesses. I know that if a service provider did this to me, I would prefer migrating to their competition, all things being equal.
- robogration 2y agoI think one thing the author did not mention is telemetry. It is very dangerous to deactivate the service when there are still incoming traffic. Normally people only deactivate the service when there is absolutely zero incoming requests in an extended period of time. Using old API as proxy as new API has its merit in that the callers needs to do little to pick up the logic change. However, often the old API are badly designed. Now you basically have two sets of APIs to maintain. When business requests come in, you might need to enhance the old API and the new API to support it. Of course, one could push the callers to switch to new API, but in a large organization, this often involves a lot of politics. The difficulty of migrating/deactivating an API also depends on the implementation. If the API is a C++ library, it can be hard to tell how often the API gets called as you cannot tell whether the code path that contains your API is executed or not. For a HTTP or RPC service, at least you could put in some logs/metrics to check who is actually sending requests.
- beefnugs 2y agoThis kind of thing is important to consider, and could really change the world for the better in terms of more reliable world... Should we consider adding "user suggestion" as a feature to all APIs whenever possible? Clients could support by showing arbitrary messages to users, servers could say "this api will deprecate on 2024-07-10" But capitalism crushes this to hell, imagine all the extra work to anticipate future problems, when "good enough, barely works" triggers immediate SHIP IT BOYS, then most engineers will just choose "someday shit stops working, oh well its not like this will be deployed in a nuclear reactor we hope they read the eula about that not being allowed"