3 ms·
I have no inside information for TSB, but I have worked on other large migrations before at $DAYJOB. Based on the original planned 50-hour outage, I suspect th
by isp 8y ago
I have no inside information for TSB, but I have worked on other large migrations before at $DAYJOB.
Based on the original planned 50-hour outage, I suspect that TSB chose a "big bang" migration: all-or-nothing, no rollback possible. These are technically easier (so cheaper) to develop, but far far riskier than taking a phased approach.
I have successfully argued against such approaches in the past, due to the high risk of catastrophic failure if it goes wrong (i.e., exactly what has happened to TSB).
I regard TSB's failure as an in-the-making textbook example of "how not to do it".
- mcroft 8y agoOften, these kinds of migrations are the only option due to previous cost cutting/money saving decisions - likely due to pressure from the business and nothing to do with any technical reason. I remember one migration which involved lots of internal services which were all tightly coupled, meaning updates across the whole backend, not just the part that needed it and no way to do a phased rollout. It was made worse by the fact that the update was to the platform and that depended on a DB upgrade, but that upgrade was incompatible with the old version which meant everything had to be done at once. They had no disaster recovery plan. They got it done, but I'm sure a lot of people involved went grey early thanks to that nightmare. Of course the business had no idea of the utter mess that caused all of this and continued to make harmful decisions in the name of saving money and further underfunding the IT department.
- isp 8y agoFair general point. (Though it is often - if not always - possible to do some sort of dual-running, to mitigate risk. Albeit at much greater expense.) In the specific case of TSB, I have my suspicions about the reasons for cutting corners in this migration, given: https://www.thetimes.co.uk/article/missed-deadlines-leave-1m-hole-in-tsb-boss-paul-pester-s-pocket-zpzdvbt7k https://www.thetimes.co.uk/article/missed-deadlines-leave-1m...
- ams6110 8y agoEven with tightly coupled systems, ultimately it's all bits on the wire. You can always develop shims or intermediaries to perform a phased migration with rollback options. Of course that takes time and costs money, which is why some people take the damn the torpedoes approach.
- dhimes 8y agoIt's bad enough to not be able to roll it back, but to also not thoroughly test? Unfathomable.
- dx034 8y agoThe problem with core banking systems is that they're nearly important to test. There's no distributed system, you always need one single source of truth for the whole bank (the ledger). Moving part of that is only possible if you keep shadow accounts in the old system. That's also why live migrations aren't possible and need at least one night downtime.
- dhimes 8y agoCan't you set up fake accounts with fake data? Random numbers? I do that all the time is a different context. I have a bunch of scripts I run that populate a database with nonsense. So not shadow accounts, but fake accounts.