4 ms·
I recall sometime in the mid 2000s there was a fever for achieving five-9s (99.999% uptime, I think -- it became fodder for a few episodes of Mr. Robot). Not t
by devchix 6y ago
I recall sometime in the mid 2000s there was a fever for achieving five-9s (99.999% uptime, I think -- it became fodder for a few episodes of Mr. Robot). Not that the metric ever went away, but back then a lot of BigIron(TM) vendors advertised achieving five-9s by replacing hardware while the OS remained running and continuing service. Sun 15K and 25K series (Gilfoyle had a used one in the garage running his network) were behemoths whose mem/cpu boards you could swap out wholesale while the entire frame and backplane was powered on, and while the OS the board came out of remains functioning. There were many caveats around the procedure but it worked. Execs and sales guys loved those demos. These monsters were expensive and banks and energy conglomerates were buying them by the dozens. There was also a big todo about hot swappable drives. The idea that you could be doing hardware maintenance while the machine was still running was a novelty, something like brain surgery while the patient was not only awake, but awake and eating, driving his car, talking on the phone, etc.
A decade later I look back with deep surprise that we didn't think to abstract out the service instead of the hardware. I don't know how many of those behemoths are still being bought, now I work almost exclusively with small server instances that can come and go on the fly. Micro services and AWS have taken five-9s in a different direction. I frequently think of Sun as a failed Hephaestus, in a Christopher Nolan film he would be brilliant but could only turn out clumsy tools because of his deformity, he hates the things he makes so he throws them away before completion. Men find these cast-offs and temper and refine them.
- chasd00 6y agoi remember those. A friend of mine was a network engineer at a local datacenter ( UUNET then MCI pre-scandal ) and said companies were buying Suns for everything no matter how trivial. He worked a night shift and i use to go hang out with him in the noc and download movies (residential bandwidth was not what it is today). Odd he nor i ever get in any trouble for that heh.
- tlack 6y agoWell as I recall there were a few reasons that people focused on reliability in hardware in the late 90s: 1. Shared state storage systems that supported replication were rare (I think Oracle and Informix maybe?) 2. Virtualization software was in its infancy (did SunOS have something before Solaris?) 3. RAM and hardware were waaaaay more expensive, meaning you often had to buy more pure metal just to answer questions fast enough At least that's my take on it based on my dim faded memories
- theevilsharpie 6y ago> [In the mid-2000's], a lot of BigIron(TM) vendors advertised achieving five-9s by replacing hardware while the OS remained running and continuing service... A decade later I look back with deep surprise that we didn't think to abstract out the service instead of the hardware.... Micro services and AWS have taken five-9s in a different direction. In the mid-2000s, enterprises were (and in many cases, still are) running proprietary software with proprietary RPC protocols that had no available source code or other means of modification, and most had no support for application-level high availability, access control, or any other operational quality-of-life feature that people take for granted today. Rather, that functionality was handled at the infrastructure level, through things like the aforementioned Big Iron. The world looks different today, but those machines made sense for the environment at the time.
- ClumsyPilot 6y agoI think it kind of makes sence in general, and the obly questuon is whether it could be achieved at lower cost. Complexity of Todays commodity machines is conparable to big iron kf yesteryear
- larrik 6y ago> a lot of BigIron(TM) vendors advertised achieving five-9s by replacing hardware while the OS remained running and continuing service. AS/400's were capable of that in the 90's (possibly the 80's as well). Heck, they'd call IBM for replacement parts on their own. You'd show up for work and there'd be an IBM guy waiting to be let in. He'd swap out a part with no downtime, and be gone. I've seen machines with uptimes of over a decade with zero on-site IT.
- blhack 6y agoWe had one of these at an old office of mine. I actually think it's really cool.
- pmiller2 6y agoI recall working on some Sun machines with hot-swappable CPUs (and, I assume, disks and other peripherals). If they somehow made memory hot swappable (I'm sure it's possible, just uncommon and/or verrry expensive), with hot swap CPUs and disks, and redundant power supplies, you could tear the machine half apart and it would still keep running. Of course, at that point, once everything is hot swappable, there are generally multiples of everything, so your one machine is really more like multiple machines inside one box than a single discrete machine.
- bcrosby95 6y agoAWS single region SLA isn't 5 9s though. If you want 5 9s in the cloud, not even multiple AZ is enough - you need to go multi region or even multi cloud.
- closetohome 6y ago> Multicloud This sounds like something they'd make up on NCIS.
- vxNsr 6y agoIt means using aws+azure+do+gCloud
- techslave 6y ago> a novelty not a novelty. the true bigiron vendors (not sun) had been doing this for decades. mainframe reliability puts the upstart unix systems to shame.
- devchix 6y agoUp to that moment, hardware maintenance meant having to power cycle the server. True HPC systems like those that ran at the US National Labs didn't find its way to the general market, and still haven't as far as I know.
- techslave 6y agoeven the AS/400 didn’t require power down. this is actually so ancient it’s hard to find docs. here’s something from 1976. (the report is 1990 but the hardwares dates to ‘76). https://www.hpl.hp.com/techreports/tandem/TR-90.5.pdf https://www.hpl.hp.com/techreports/tandem/TR-90.5.pdf