4 ms·
The author has clearly an opinion. I am getting more and more the feeling, that IT is a field where opinions are strongly hold, mostly because of past personal
by yokaze 5y ago
The author has clearly an opinion. I am getting more and more the feeling, that IT is a field where opinions are strongly hold, mostly because of past personal bad experiences, and not with much evidence.
There are many ways to approach a problem, and there are good and bad ones, but often you have to make a trade-off. I would appreciate a more differentiated view.
The author makes the math on "high cost of maintenance windows", but leaves out the expenditure side to it. And what is the hand-rule for adding a nine behind the availability? I vaguely remember exponential cost increase. "No downtime, ever", requires next to infinite cost.
The author suggest a false dichotomy. You can have maintenance windows and while aiming for zero downtime.
The maintenance window serves the customer a higher predictability on when to expect failure.
No failures are obviously preferable, but it is never the question if to cause failures or not, but where you spend your resources on.
- spiffytech 5y agoThe author also cherry-picks some examples of planned downtime that are anomalously egregious. I occasionally get notices of planned downtime from services I use. It's almost never an issue for me - the outage windows are usually relatively brief, infrequent, and off-hours. Not every service has the luxury of even having "off-hours", but simply making planned downtime infrequent goes a long ways.
- rbarrois 5y agoIndeed — when you look at this from other engineering disciplines: having your train track built for "no downtime, ever" means that you have a third track available, with its dedicated platforms, so that you can work on one track while traffic goes on. This might make sense for your inner city loop, with trains passing by every 30 seconds (and it's gonna help when one train inevitably breaks down). However, it's a totally unreasonable cost for a station deep in the country where trains only stop 4 times a day.
- zahllos 5y agoA maintenance window doesn't necessarily have to mean downtime either, it could just mean time you can make changes where people are on hand to fix it if things go wrong. To extend your analogy, it might be a terrible idea to start doing maintenance on a single line track that sees high usage in a holiday season just before those holidays where everyone wants to get away or get home as do your staff. So you might implement what is usually called a change freeze. On the other hand you might have a different time where demand is low. Shutting down the line if it comes to that will annoy some people but not as many. So you plan your maintenance then and have people ready in case things go wrong.
- drewcoo 5y agoAdopting metaphorical analogs of all the impediments of real engineers does not make us real engineers. Software is cheap to produce and quick to change. Building a "third track" on the fly, as needed, is exactly the kind of thing we can do that actual engineers can't. They probably wish they had the ability to do things like that.
- obscura 5y agoAs with so many opinions these days, it's at one extreme and ignores the middle ground. For some apps, the big issue is when you schedule the downtime to happen. If you are truly customer-centric, you should schedule for a time that causes the least disruption for your users. Of course, a decision like this depends on other factors - the number of users likely to be affected, how critical the app changes are, costs, etc. I've noticed a number of companies in my country taking sites and apps offline in the middle of the workday when they're very likely to be in use. To me, this is unacceptable - the work on the app should be done after hours.
- jvvw 5y agoSurely this is up to the company? They make a decision as to whether to pay somebody to work after-hours versus the business they lose by not doing so. I specifically took a job that didn't require after-hours work (I have a family which I prioritise) and we did upgrades in working hours. Once in a blue moon a site would be down for a minute or two, but the trade-off was obviously worth it for my public sector employer which generally struggled with recruiting engineers.
- sysadmindotfail 5y ago>IT is a field where opinions are strongly hold, mostly because of past personal bad experiences, and not with much evidence. This is spot-on. I see this attitude often when there is a lack of data being exposed for observation and it "just feels" like ____.
- dncornholio 5y agoIt's clearly an AWS marketing blog post.
- jameshart 5y ago> The maintenance window serves the customer a higher predictability on when to expect failure. I don't get this - who wants 'predictable failure'? If you're a B2C business, this is completely unacceptable - consumers don't plan their google searches around your preannounced downtime windows. If you're a B2B SaaS business, well... your customers downstream have their own downtime to manage. If they have two vendors, each of which have 'scheduled maintenance windows' that don't overlap, then their combined availability just dropped dramatically. Far better to build systems that are generally resilient to being down - queues for offline processing, idempotent operations, retriability... you'll wind up needing them to handle scheduled downtime anyway, and once you have them, you can switch to 'as needed' maintenance with no loss of service.
- CJefferson 5y agoMost people don't plan to have downtime, but turns out even the biggest companies in the world can't achieve zero downtime. Also, some changes are one off and significant (like some database changes), they will hopefully go smoothly, but it is hard to fully prepare for every possible issue.