7 ms·
It is a good write up of the options, and sad to say there is no "right" answer. That being said, I think early companies (like mine with an API) need to go th
by gbourne 5y ago
It is a good write up of the options, and sad to say there is no "right" answer.
That being said, I think early companies (like mine with an API) need to go the "Migrations a’la Stripe" mode. Basically you take the brunt of translating old request to new ones. We've done this several times and update our docs with the new calls/parameters so new users start using the latest version. Current users have no breakage (you hope) and new user are on the latest API version. Old users are also encouraged to move to the latest version since only it contains new functionality.
It works well...however, the downside it you now have a lot of translations in your code and some users are on the new and some on the old. This means you eventually have a bit of a mess on your hands. The way we plan on mitigating this is tracking the translations and over time informing our users of the deprecation.
Larger companies with dozens or hundreds of hands in the code, or with legacy code, likely can't use this technique. Perhaps why Facebook over time move to the /v1/ technique.
- tomschlick 5y agoThe key to doing migrations on the request is having the migrations live in single classes/files. It makes it very clean to have the app migrate from version 3 to version 12 with the 9 or so classes handing the response from one to the other. I made a package a few years ago in PHP/Laravel that focuses on this: https://github.com/tomschlick/request-migrations https://github.com/tomschlick/request-migrations
- derefr 5y agoOr doing the migrations as load-balancer-side request rewrites. (Nginx is particularly good for this use-case; it’s what sets it apart from simpler LBs like HAProxy.) Then the cruft of old versions doesn’t have to live in your codebase at all, but can live in the same infrastructure-config repo that holds your backend-service-unifying route-map; your SSL config; your rate-limiting setup; etc.
- tomschlick 5y agoMy main concern there would be the management and orchestration of the releases for that since its detached from the codebase that is changing. It sounds very performant though. Have any examples of this setup?
- derefr 5y agoI don’t, but that’s because I keep my own infrastructure config living inside the same repo the code lives in. We do k8s GitOps with [a simulacra of] Google Kubernetes Engine’s “Application Delivery” (https://cloud.google.com/kubernetes-engine/docs/concepts/add-on/application-delivery https://cloud.google.com/kubernetes-engine/docs/concepts/add...). In this approach, you keep all your k8s manifests relating to the app in the app’s repo, and then the app is “released” by running a command that does the following: 1. tags a particular commit, which locks down the release as being of a particular codebase + a particular target converged k8s state; 2. generates a Docker image, tags it with the commit tag, and pushes it; 3. makes a copy of the manifests; 4. burns the build-image’s SHA into the copy of the manifests; 5. compiles the source manifests (which are using Kustomize) into a single static manifest, which is a complete definition of the new target converged cluster state for the app's k8s namespace; 6. commits that static manifest to a separate tagged “deployment” repo. A separate "deploy" command is then used to reach out to a cluster-side converger component and tells it to pull a particular commit from the "deployment" repo and converge to it. ----- We do have subcomponents that live outside this repo, though; we manage them by: 1. running a “release” within the subcomponent repo (which does create a git tag + push a build-image tagged with that tag; but doesn’t generate/commit any k8s manifests, since the subcomponent repos don’t have any k8s config of their own); and then 2. manually (for now) taking the git tag of the built image, and updating the main app repo’s k8s per-env Kustomization.yaml file with it, i.e. updating a stanza like this with a new value for "newTag": images: - name: gcr.io/our-project/subcomponent newTag: v20210409094011 In theory, our main component could also be tracked as a subcomponent in this manner, such that all our infra config would live in its own basically-empty "app" repo; but that would lose one of the main advantages of GitOps, which is being able to see from the git log exactly what went into a deployed release, both in terms of build-image and infra-config. Our subcomponent services don't evolve nearly as quickly as our main app does, so we don't lose much in the way of release comprehensibility by having them "symbolically linked" to the release like this. As it is, though, our app's LB config for api.example.com lives as a k8s Ingress manifest at /.appctl/config/base/unified-api/ingress.yaml within our app repo. (We're using ingress-nginx, so there's a pretty direct 1:1 mapping between this manifest and an Nginx server{} block.)
- gregmac 5y agoI've done this in a codebase as well, and it worked quite well. The "main" code was always updated to be the latest stuff, and the old methods were moved to a class named like `WhateverApi_v1_0` (based on the last version it was available in). A routing engine picked up the proper controller based on the requested API version. The other thing we had to help make this easier was automated API "shape" tests. These are per-version tests that basically just call every API method and check that the response matches a specific schema. For every release, we made a new directory containing a copy of all the tests from the last version, and then never touched any of the old directories. If an API ever changed in a backwards-incompatible way, one of these tests would catch it. All this was relatively painless to maintain, and we also never had any API regression issues over dozens of releases spanning years. I'll note we did sometimes "cheat" and add new properties to an existing model (without making the backwards-compatible controller stuff), but this doesn't break any consumers because of the nature of JSON (at least we never had anyone complain about it, and I'm not aware of any language where that would happen).
- tomschlick 5y agoYup thats exactly what the package I linked does as well and it has worked without a hitch on the project I have implemented it on. All you have to do to test a specific version is send that header in the integration test and its good to go.
- derefr 5y ago> informing users of the deprecation If you can get users to take notice from the start of the fact that your API has the possibility of spitting out certain ‘temporary errors’ that “MUST” trigger a client-side retry with backoff (e.g. 429 errors) — and your API does actually emit these errors sometimes, such that clients' codebases are very likely to have this error-handling code in place — then you’re in a much better situation here: you can give deprecations like these technical force, by making deprecated APIs begin to randomly emit spurious failures, with the failure self-documenting with a response error message like “usage of this API has been deprecated and will gradually cease to function. [link to blog post about transitioning to new API]” Start off with allowing 99% of requests through; then lower it following a sigmoid, e.g. 95%; 90%; 66%; etc; until eventually it’s at some low number like 1%. Wait a few months at 1%, and then turn it off. (And, obviously, email every user who your metrics say are still using the deprecated API, each time you ratchet it down, to nudge them once again to update their code.) Any app developer who doesn’t notice that your API is now requiring ~100 retries to contact successfully, either isn’t using the results for anything; or just plain isn’t around any more to update their app, such that the app is now abandonware. Either way, you’re likely safe to shut off the API at that point, and finally clean up that code. Of course, you can also make side-deals with any big enterprise user who needs more time, putting their API keys on a whitelist so that they'll get 100% success from the API until the very end. (Try not to allow them to slip the sunset date for the API, though; that would force you to keep the code around longer, which is what you're trying to avoid!)
- cpeterso 5y agoOr gradually increase the API latency so the system still produces correct results but the users are increasingly motivated to upgrade to the new API. I like the idea of encouraging clients to handle API errors more robustly, but the type of client developers that are slow to upgrade to the new API might also write sloppy code. Their client will break but they’re likely to blame your service.
- derefr 5y agoThe problem with just adding artificial latency, is that it makes it unclear why latency is increasing. Users may think "maybe the upstream service is just gaining users and failing to scale." The great thing about an error, is that it gets your users' users upset (without really impacting them materially, if you're only doing them 1% of the time), in a way that tends to get your user to sit down and debug the problem. When they open their logs, they'll see your service's deprecation notice in each error message. If you're already thinking about this when designing your API response envelope format, you could have a "warnings" field to put information like that in. But a lot of devs are going to miss that entirely, because they've got your API wrapped in a gateway object that strips out everything but the data they're interested in for "successful" responses. They'll only actually pass any of the envelope stuff through in the case of an error.