4 ms·
My advice whenever you have large, complex, or slow to perform (eg hours/days) migrations, that will take some time to process, don't put yourself into a positi
by plasma 5y ago
My advice whenever you have large, complex, or slow to perform (eg hours/days) migrations, that will take some time to process, don't put yourself into a position of "we can't go back", until the very last moment, after you have been assured as much as possible things are fine.
Things to keep in mind include:
* If migrating data to a new column, keep the old data for a while
* After migrating the data (with new code writing old/new data), you may want to do an audit query to validate everything was migrated correctly
* If you need to mutate data inline in the column, add a new column and just copy the (pre-migrated) data to it as a backup (so worst case you have a pristine copy from before), then drop that column some time later
* Runtime config flag in your app that lets you read the new column/truth source as authoritative for a decision, or put the flag back to false to revert it. This could include a percentage so that you slowly creep load onto a new service and build trust in it functioning correctly.
* If the data migrated is quite complicated, help yourself to easily switch between old/new data sources for your own manual review in your apps code paths at runtime (eg, render this analytics from Source A or Source B via a runtime flag, query string parameter etc)
- cletus 5y agoDid you even read this article? That’s literally what it said.
- jameshart 5y agoThe article didn’t mention having a rollback strategy. This post is a useful addition to the discussion.
- hackerfromthefu 5y agoThe authors reply to this comment also assumed that the data schema migration will always succeed - by saying the new schema usually 'has more information' that you can reconstruct the original from. All in all the authors writeup and reply is only covering the happy path, not the real world nuance that mistakes or errors can happen!!
- rtpg 5y agoI mean... I understand the notion of fucking up a migration, of actually introducing data loss etc. I guess this is partly my fault, cuz I said this list was pedantic but I didn't go in deep enough ("rollback if something is clearly blowing up" for example). I mean yeah, don't delete the old data until you're sure! But you know what? If I'm moving two booleans into an enum with 4 states (or less states but I checked the production database to confirm a case didn't exist), I feel pretty confident about doing that data move and not having some secret data being missed. You can use logic/math/data analysis to determine you are handling every case! And yeah, you can make mistakes, but the variety of mistakes become much smaller. And this is all in the framework of having backups of your data, having reviews, letting time pass to reduce uncertainty... just run of the mill "run stuff on computers and editing multi-tenant data" stuff.
- wdb 5y agoI am wondering how the long running migrations are done. Are they running as a scheduled job outside the app? I can imagine it can't be part of the app itself as that would block deploying a new version of the app (/docker image)
- rtpg 5y agoOriginal author here. I do think it’s good to think about rollback strategies (especially given that your change might go out with something else that might need rollbacks). There is one thing though, and that is that a lot of these transformations end with more information, so you can usually recover the old state if you ever needed to (or write some compatibility shim for whatever reason) I had a whole part about rollbacks (and an intuitive proof about why each step was needed and not being able to skip any) in my notes for this, but I decided to leave it out cuz I couldn’t do it justice just yet. There are so many little details to cover in this space, and honestly the vocabulary and tooling are a bit poor so it’s hard to be succinct
- hinkley 5y ago> new column/truth source I think we need to be a bit more explicit in this part. Database columns are your system of record. If the code treats them as the source of truth, you need to break that first. If your code is factored properly you only need a toggle in the retrieval code. If it's not factored properly, now you understand why some of your peers are pushy about separating data retrieval from data use - it concentrates any mutation to a point in time before first use where you can easily find it when surprises happen and where there can be no concurrency issues. If that's too daunting you can settle for the lesser solution of replacing use with a function call. Once everything is using the new source of truth, you can put a toggle there. And after the migration you can decide if you want to keep them separate or fold them back together.
- jrockway 5y agoI strongly agree with this. Relatively recently, I did a major migration to clean up some weird account states that we didn't want to support anymore, so we could launch a new feature. I spent some time writing all the SQL queries to do the migration, went into staging, ran them... and found that there were a lot of edge cases that weren't handled correctly. It was terrible. I decided to abort the SQL statement approach, and just wrote a somewhat-complicated program with extensive unit tests to handle the migration. The tests exercised the edge cases and ensured that they were accounted for correctly in the migration. (And we had time to come up with a real plan for each of them.) When the time came to do the cut-over in production, we had a war room, ran the program, and it was all over in seconds. Everything worked perfectly and I honestly felt a little dejected how boring the most major change we had ever made as a team went. I had hyped myself up for some stressful debugging, hair pulling, restoring things from backups... but it was all over and done with perfectly in a few seconds. I didn't really know what to do with the rest of the day. It worked! Major new feature! Back to work. I try to make all future "major" migrations equally boring. The time you spend debugging things in development is time you don't spend with production being down. Worth it every time.