4 ms·
Figuring out whether today is a working day is actually a dynamic problem. It is different for every locale, and it changes more frequently than one might expec
by eigenrick 7y ago
Figuring out whether today is a working day is actually a dynamic problem. It is different for every locale, and it changes more frequently than one might expect.
It's actually a great example of something that changes rather frequently, and it's a concern that cuts across dozens (or in Monzo's case, maybe hundreds) of services.
Causing a roll-out to hundreds of critical business functions (whether hosted in a monolith or microservices) because Uganda added a holiday seems quite excessive.
The rule I and my company follows is that the shape of a module should follow its deployment model and scope. Some services are global in scope (literally targeted at the earth) others are scoped to a small subset of our inner network. They should have a different code repo because they have a different rollout schedule.
In the case of the working-days services. Only 1 rollout needs to occur, and it seems like a fairly safe rollout.
Let's consider another service, the CriticalTransactionService, or CTS. It executes critical business transactions in a very stateful way. When deploying a new version, in order to avoid any loss of availability, often a special rollout dance must be executed. E.g. switching the master transaction writer to another region, which means changing databases from passive replication to active master, and vice-versa.
One might consider this rollout "risky" and therefore only limits it to happen at 3am on a Saturday. It seems like a good idea to limit the scope of this 3am Saturday rollout only to the CriticalTransactionService, and most other services are free to deploy to prod whenever they wish.
- trollied 7y agoI guess what I don't get is that the service needs to be rolled out/restarted too. Doesn't the data just live in a database & there's a cache going on that can be invalidated? Or is everything just so overcomplicated these days?
- eigenrick 7y agoThis particular case could be a couple tables in a database. To follow proper encapsulation guidelines, it would have to be front-ended by some stored procedures, so that the underlying representation is free to change. How would you roll out this change to production? If you just swapped out tables and stored procedures, you'd have downtime. You can't just install a parallel table and stored procedures, because there is no way to tell all of the consumers to use the v2 of your functions. So you'd have to temporarily remove availability of the WorkingDays functionality. If it were deployed as a microservice, you are free to deploy a v2 of WorkingDays in parallel. When it's live, the old one goes away, and there is no loss of availability.
- pbalau 7y agoThere might be db schema changes, migrations, who knows what else. Your service might even have multiple databases it interacts with. Your service might need to precache some stuff, that will add startup time. You can't simply shut down your service, you need to finish answering the requests already started.
- jlokier 7y agoA database table with proactively-invalidatable caches in front is already a little bit complicated. Some of the things that can go wrong are: - One or more caches failing to be invalidated by your "update business-day database" function. - Temporary loss of connectivity to a cache from the updater, resulting in failure to proactively invalidate, so incorrect business-day results used by some services depending on which cache they read. - Additional caches you didn't realise someone had added to the application or library code, that don't get invalidate. - Behavioural inconsistencies as different functions in the distributed system read either the database or different caches, and get inconsistent results shortly after an update (after DB write, before proactive invalidation) ... In other words, the usual problems with distributed systems. Also some logic issues, which a microservice is more likely to detect and log: - No entry in the database for some far-future or far-past date, that application code assumed was ok to query, but nobody filled out in the DB. Anyway, a database table can be thought of as a kind of microservice, in the sense that it's also remote network call, and it's also something you need to treat as an API contract with careful updates. If you aren't using a consistently distributed database, it suffers from multi-zone latency just like microservices calls. Another reason why something like "business day" might be a function of stored data rather than just database data is that's actually a bit of a complicated and evolving logic as the product evolves: Maybe the first application logic assumed a business day to be "Monday to Friday except for national holidays in the DB", did a DB lookup for the holiday flag, and combined with Mon-Fri. Later, we expanded to new countries and find it's different in other countries. Later, we found country was insufficient and it needed further details about administrative regions. After each step, the previous database query and logic would still work but be incorrect due to missing a query field. So at each step, new logic needed to be rolled out globally across all services. For behaviour consistency across the system, all at the same time (similar to a DB schema change). With a microservice, the logic update would be global and immediately consistent. Each part of the application could be updated separately, on their own development schedules, to provide the additional query field, long before the microservice was switched over to requiring the field. After that the microservice would fail any requests without the additional query field, making certain all functions and services acted consistently with the new logic, or failed. ...Having said all that, my favourite isn't microservices at all. In principle, logic updates can be rolled out with ACID-like properties and transactions, all the way to the edge of a system, just like data updates and cached data. There is no inherent reason why application and library code needs to suffer from complicated updates across a large business system. Business logic can be dynamic, replicated, and guaranteed consistent, while at the same time being available as fast, direct function calls. But I've rarely seen this implemented.