6 ms·
No downtime is acceptable, but they have only one server? What if a technical failure happen? What if there's a fire in the server room? What if there is an ea
by Milank 6y ago
No downtime is acceptable, but they have only one server?
What if a technical failure happen? What if there's a fire in the server room? What if there is an earthquake and the building collapses? What if... many things can happen that can result in a long, long downtime with this tactics.
If uptime is so crucial, the system should be setup in such way that moving one server should be a peace of cake, not a spec-ops mission.
- redwood 6y agoRemind me of how IBM positions mainframes: they are so highly available that you simply never let them shut down.
- lasereyes136 6y agoIBM Mainframes are designed to be serviced while running so if you have multiple CPUs you can offline one at a time for upgrade it without the whole mainframe going down. Big Sun Solaris boxes where built like at as well. If your mainframe had only one CPU, you did have to turn it off in order to service it. But you could upgrade the OS without turning it off. While they aren't cool tech now, mainframes are a marvel of hardware engineering.
- chasd00 6y agoplus, i would imagine turning them on and bringing them online isn't just a press of a button.
- MrMorden 6y agoIt's not. https://web.archive.org/web/20190324191654/https://www.ibm.com/support/knowledgecenter/zosbasics/com.ibm.zos.zsysprog/zsysprogc_initialization.htm https://web.archive.org/web/20190324191654/https://www.ibm.c... (archive.org link because ibm.com apparently isn't hosted on a mainframe.)
- walrus01 6y agoFrom an ISP perspective this seems like the sort of company that orders one $250 a month business DIA circuit (at a price point where there is no ISP ROI for building a true ring topology to feed a stub customer) and has no backup circuit. Then the inevitable happens like a dump truck 2km away with a raised dump driving through aerial fiber and causing an 18 hour outage. Some circuits might average 5 to 7 nines of uptime over a year, but the next year is dump truck time... You can never truly be certain.
- momokoko 6y agoYou’d be shocked how rare downtime is with modern hardware. A redundant power supply and SSDs in the right RAID configuration typically will not have any issues for years until it can be replaced by a newer model. Also, hardware monitoring is significantly improved to the point where you’ll typically know if something will fail and can schedule the maintenance. In the past power supplies and spinning disc hard drives would fail much more often. It’s basically a solved problem, outside of extremely mission critical, 5 nines kind of stuff, that we all forgot because of AWS. HN ran, and may still run, on a single bare metal server.
- paulie_a 6y agoQuality hardware has existed for years. At a ford motor plant they were doing an inventory and couldn't locate a 10 ton mainframe. It was working so well for 15 or so years the tribal knowledge of where it was physically located was lost.
- Milank 6y agoThis is all true, but you still can't rely on increased hardware quality if you can't afford any downtime due to moving (a one-time event) a server. Also, that doesn't cover other problems mentioned here, like natural disasters, ISP problems, etc.
- hnlmorg 6y agoOften these kinds of SLAs are decided upon based on blame rather than what is reasonably required by the customers of that system. In this case, moving offices means the downtime is due to internal reasons. But if an ISP goes down or there is a natural disaster, then that isn't in their control. Also cost does come in play as well. Multiple physical links in would be very expensive for what sounds like internal services. Likewise a natural disaster might cause bigger issues to the company than those internal services going down. They might still have offsite back ups (I'd hope they would!) so at least they can recover the services but the cost of having a live redundancy system off site might not justify those risk factors. The customers requires are definitely unreasonable though. I'd hope those systems are regularly patched, in which case when is downtime for that scheduled and why is that acceptable but not when you're physically moving the server? I doesn't really make much sense; but then "not making much sense" also quite a common problem when providing IT services for others.
- galoisgirl 6y ago> Should have been a 5 minute job if done correctly. Owner ended up paying for over 10 hours of work. Stupidest thing I've ever had to do. You can see the common sense ship has sailed.
- coldcode 6y agoI worked at my last job for a place with a single rack mounted set of Windows servers at a data center - with no backup power supply, no backups of any kind for that matter, no UPS and no redundancy of any system, plus they didn't even have an admin for 6 months. The CEO refused to spend money on a 2nd anything. The company has 2000 employees. One server held all of the companies photos (which is basically the core of the business) and of course was not backed up.
- Milank 6y agoOf course it can work, you can get far with one server and no spending on anything like backups, UPS, etc. Whether it's smart and good for your business/reputation is a different question.
- elliekelly 6y agoThis is the kind of company that could benefit immensely from a ransomware attack.
- closetohome 6y agoMy boss refused to use UPSs for years because he bought one once and couldn't get it to stop beeping.
- gear54rus 6y agoHe should have taken it offline without notifying this brain-dead manager. Probably wouldn't have noticed lol. And then charge for those 5 hours for good measure. In general, this stupid trend of wanting 0 downtime makes no sense to me. If you're not NASA, police or other emergency service you 100% can afford a few hours of downtime with scheduling it be forehead.
- icedchai 6y agoNever mind these less common scenarios... What do they do about Windows updates?
- YetAnotherNick 6y agoYou wrote one server but describe the failure modes of having one data center. I think it is very very uncommon and hard to allow for data center level issue. After all Instagram and 100 other site failed when one AWS data center went down. I would interested to know how/whether anyone's backend will work if any data center and its databases completely fails due to fire/earthquake/networking etc. Second thing is having multiple machines for server. In theory it might help in increasing the availability but in practice I haven't seen any random issue due to machine which occurs just based on probability. I think almost all failure modes that exist, they are correlated between machines. eg suppose you have data loss on one machine, you could more likely than not, blame it on code and it would be similar across machines.
- toast0 6y agoRe: single datacenter. At the basic level, you need a second datacenter with enough machines to provide your service (or a emergency version at least), replication of data, and a way to switch traffic. It's doable, but expensive in capital and development. If you're dependant on outsourced services, they also need to be available from both datacenters and not served from only one. In an ideal world, your two datacenters would be managed by different companies, so you would avoid any one company's global routing failure (IBM had one recently). Re: multiple servers. Power supplies fail, memory modules fail, cpus fail, fans fail, storage drives fail. Sometimes those are correlated --- the HP SSDs that failed when the power on hours hit a limit (two separate models) are going to be pretty correlated if they were purchased new and stuck into servers at a similar time and then on 24/7. Most of those failures aren't that correlated though. Software failures would be more likely to be correlated though, of course. The key thing is to really think about what the cost for being down is, how long is acceptable/desirable to be down, and how much you're willing to spend to hit those goals.
- YetAnotherNick 6y ago> In an ideal world, your two datacenters would be managed by different companies, so you would avoid any one company's global routing failure I can't understand this. I think transferring servers would be the the least of problems. Its the transferring of database and maintaining consistent version of databases in both the locations. Moving the snapshots after every X minutes doesn't maintain consistency. I would like to read about any company that is able to do this, as honestly it sounds really hard to me. Is there any writeup of IBM thing you mentioned?
- misiti3780 6y agoor even better, how do they apply OS patches?
- wastedhours 6y agoWe used to have one server for a website I was a content guy on - it was in a standard PC case, plugged into a switch in the IT team's office (this was not a tech-centered org). The main IT guy went on holiday and one of the cover guys from another office decided to tidy up. He unplugged the server and thought (and told me after his thought process) "if anyone was using it, they'll let us know". This was the one, single box for the whole website - no one else was monitoring (even though the central office had a proper, dedicated web team) and the assumption was I was sysadmin. An hour later I'm sprinting down the corridor to find out what the hell happened and why I can't even SSH into the box. We put a sticker on the case saying not to unplug it after that...