10 ms·
There is another problem that wasn't covered in the article. The 10+ years of stability leads to behaviors and outcomes that remind me of the long-lived SSL ce
by darkhelmet 3y ago
There is another problem that wasn't covered in the article. The 10+ years of stability leads to behaviors and outcomes that remind me of the long-lived SSL certificate problem. Updating is done so infrequently that the "how?" is forgotten. As the 10 year support limit approaches, most of the old team members who did it last time are gone, tech debt is through the roof, few people know where everything is or how to build it, and so on. Enterprise Linux "stability" enables all sorts of bad behavior if your company is inclined that way.
LetsEncrypt did us a huge favor by forcing automation vs having the guy who knows how to update the SSL certs every 4.9 years and left 6 months ago. I'd like to see the RHEL stability model go away too and force people to complete their automation and solve the problems of being able to rebuilding on demand - and actually doing it.
(I know, most HN folk are well disciplined but there are a lot of corporate cultures that are not.)
- derekp7 3y agoDon't you end up with the same problem with the automation that has been running fine for 5 years, then suddenly breaks? And the person that set it up is either gone, or has no clue how they did it 5 years ago.
- bonzini 3y agoThe idea is that you deploy from scratch all the infrastructure every 6 months, first to testing and then to production.
- retbull 3y agoEvery 6 months? That seems like a pretty long window for tribal knowledge to get lost. Is 6 months arbitrary or is there some reasoning behind that cadence?
- noizejoy 3y agoArguably, tribal knowledge and the dependence on it need to be managed as much as anything else.
- postmodest 3y agoAll you damn kids work in a different industry than I do.
- histriosum 3y agoThat's an amazing amount of effort expended on something that provides exactly zero revenue. I understand the concept, but I've never been fortunate enough to work in a business where that was practical.
- U2EF1 3y agoLE by default successfully ran 2 months prior. 2 months and 5 years are two completely different worlds in terms of bit rot. That and there are many generic tutorials and scripts and knowledgeable devs for configuring LE fresh.
- justinsaccount 3y agoBefore LE almost no one automated SSL cert refresh. Depending on your SSL cert vendor you couldn't automate things even if you wanted to. It's not that the automation ran fine for five years, it's that you'd be lucky if the manual process last done 5 years ago was even documented. SSLMate is about as old as LE, they both started around the same time.
- btown 3y agoRecently saw a (thankfully not mission-critical) old k8s cluster fall down with absurd incompatibilities between node versions, cluster versions, and cert-manager versions - all of which only support upgrades one version at a time. Even infrastructure-as-code doesn’t save you if you need to upgrade something but don’t have the time and expertise (and esoteric changelog knowledge!) to reliably upgrade everything else.
- bogantech 3y ago> I'd like to see the RHEL stability model go away too and force people to complete their automation and solve the problems of being able to rebuilding on demand - and actually doing it. Whenever there's a new distribution release it invariably breaks a bunch of things with the automation and you spend more time massaging your playbook so it works again than it would have taken to do it by hand
- geerlingguy 3y agosystemd threw the biggest wrench, by far, in my automation workflows (this was before containers came to dominate a lot of the landscape, so everything was managed with init scripts). I like it now, but it also broke things quite frequently in the early days, and there was a looong period of time when you had to shim software to work with systemd. But even still, things like snaps, the way Debian handles system Python, and various little changes that have an outsize effect on automated deployment do cause a good amount of churn with automation.
- nijave 3y agoYup, enterprise Linux insulates you from unneeded change (in the business context). For most companies, systemd will have no impact on the bottom line vs sysvinit vs whatever. However, paying an extra engineer to sort through all the changes possibly will. On the other hand, there's some interesting trends like monokernels and minimal OS images that leverage services running off-machine instead of expecting so many local services removing some of the complexity/volatility (DNS, SMTP, federated login)
- midoridensha 3y ago>things like snaps, the way Debian handles system Python Both these things should not be an issue for anyone, just one or the other.
- mananaysiempre 3y ago> Whenever there's a new distribution release it invariably breaks a bunch of things with the automation and you spend more time massaging your playbook so it works again than it would have taken to do it by hand The rolling-release life is that things break constantly, during each week’s upgrade, but only a little bit at a time (and hopefully in staging). I don’t know if this is better for system administration, necessarily, but if you’re used to a stable-release dynamic of heavy discrete breakage and piles of backported patches, then you might be imagining the same scale of breakage every upgrade, which is isn’t the experience at all. So don’t discard the rolling-release option because of this preconception.
- NegativeK 3y ago> complete their automation There are too many places where the current guard is going to have to die off before they even _start_ automation. So we're looking at 20-30 years. LTS is the actual sane solution for these places, despite how utterly insane it is.
- picozeta 3y agoHow is the problem of TLS certificates related to Redhat Linux or Enterprise Linux? I think these are orthogonal problems.
- microtonal 3y agoI think it was an analogy. If you don't do a thing for a long time (updating SSL certificates, updating a Linux system to a major new version), the knowledge of how the systems were maintained/built gets lost. If you have an automated, repeatable process that moves with the times, it is more likely that the process is codified (either in documentation or in infrastructure as data) and easy to repeat.
- myself248 3y agoLikewise, it's been suggested that the 19.6-year (1024-week) GPS epoch is pessimal. Rollover is infrequent enough to be ignored, but frequent enough to actually happen and cause problems. Folks who know such things better than I do, have suggested that it would've been far better at like a 64-week rollover (or just chop it to 52 and leave part of the code space unused), that way everyone would have to have a plan for it. Nobody could claim they don't expect their hardware to be in use 64 weeks in the future therefore they can ignore rollover.
- ghaff 3y agoFunnily enough, I had the Unix epoch time question come up with a customer (who makes very long-lived pseudo-embedded systems) come up in discussion last week.
- thesuperbigfrog 3y ago>> I'd like to see the RHEL stability model go away too and force people to complete their automation and solve the problems of being able to rebuilding on demand - and actually doing it. So how could this be accomplished? A Nix-OS style approach coupled with an immutable OS core? It is much harder to offer stability guarantees than to just publish updates in a rolling release fashion. And yet big organizations pay big money for that stability or the support to poke an enterprise software provider to get stuff working for their needs.
- sam_lowry_ 3y agoArchLinux-style rather thsn NixOS style. Just roll the updates when they are ready into your very own test, int, acc, and finally prod.
- thesuperbigfrog 3y ago>> ArchLinux-style rather thsn NixOS style. Just roll the updates when they are ready into your very own test, int, acc, and finally prod. The issue is that a rolling-release approach does not have stability guarantees and forces everything to be upgrading all the time. This does not work very well if you have specialized hardware or scientific equipment. If the drivers for your lab equipment work with a given release of an enterprise linux, you can't just jump on the next release until you have working drivers ready. The same is true if you are working with some enterprise software which is only certified to work with a given release of an enterprise linux. Would you really want to run business critical software on a version of the operating system which is not (yet) supported by the vendor?
- AnonymousPlanet 3y agoAll those hours hunting for the reasons why something suddenly stopped working every two or three weeks need to be paid. So maintenance cost for Linux servers would either skyrocket or no updates would ever be done for years. There are very good reasons why rolling releases in infrastructure are basically a no-go.
- 3y ago
- michaelt 3y ago> I'd like to see the RHEL stability model go away too and force people to complete their automation and solve the problems of being able to rebuilding on demand - and actually doing it. In this model, what happens when the next Python2->Python3 breaking change comes along?
- microtonal 3y agoUsing whatever Python your distribution needed is bad practice. Own your application environment, there are plenty of ways to do this, such as Nix and Docker, which make your Python environments reproducible across systems. Also, Enterprise Linus is one of the reasons (definitely not the only) that the migration took such a long time. Too many enterprise shops that stuck with Python 2 because it's the lazy thing to do. The tech debt grows every year you don't move with the ecosystem.
- curt15 3y ago>Using whatever Python your distribution needed is bad practice. Own your application environment, there are plenty of ways to do this, such as Nix and Docker, which make your Python environments reproducible across systems. How far down does "own your application environment" extend? How about libc? What is the role of the underlying OS?
- thesuperbigfrog 3y ago>> How far down does "own your application environment" extend? It depends on the needs of your application. >> How about libc? If you need to make sure the underlying libc has what you need, you must either bring your own libc or have sufficient feature test macros and adapters to account for possible differences. >> What is the role of the underlying OS? It depends on what the application requires. What operating system features, if any, do you require? Do you have any timing or scheduling requirements that are sensitive for your application? Do you need real-time responsiveness? How does the operating system handle failure scenarios? What guarantees, if any, does it make when hardware fails? Is it okay for your application to crash if a portion of the computer's memory or disk borks?
- gjsman-1000 3y agoI wonder if there would be a market for an enterprise-grade server microkernel OS. It's not the 90s anymore - Nintendo and QNX are shipping tens of millions of microkernel installs every year; and hardware is fast enough that choosing correctness and security over speed is a valid tradeoff. Maybe if I win the lottery...
- vbezhenar 3y agoKaspersky recently developed their own proprietary microkernel OS. AFAIK they target it for IoT, but kernel is kernel, probably could be used with ordinary servers as well. Main issue is drivers, of course. It's hard to beat Linux. It contains open source drivers and server vendors usually target Linux and Windows with their driver efforts.
- TylerE 3y agoThese things tend to trade 200% performance for 10% security, though. That's not a tradeoff I am comfortable with in anything like all situations.
- gjsman-1000 3y agoNot necessarily if you build them right. Nintendo’s Switch is a true microkernel and, if it cost 200% performance, there’s no way it would be viable on a 2015 Tegra X1. The 200% thing is kind of a myth that doesn’t apply to modern practice - now it’s more like 10%. As for 10% security - it’s more than 10%. Take my same example, the Switch. No bugs have been found to launch unapproved software in the last 4 years. There’s always the Secure Boot bug by NVIDIA in earlier consoles, but not even a WebKit bug will get you homebrew on a Switch. Kind of a big deal… Another example of this would be Microsoft’s experiments with what would happen if an OS was built with all apps running in managed code - no compiled apps. Performance cost? They got it down to just 7% (though, admittedly, Midori never shipped, but it did host Bing in a few countries for a few years.)
- pavon 3y agoThat is a good point, however I've not heard of too many cases where organizations intentionally skip RHEL releases. Systems that are being actively developed do regularly upgrade through each RHEL release, and the 10 year support just lets them be lazy about how quickly they do so. The only systems I see intentionally riding out the 10+ year support are deprecated systems that are already announced to sunset by the time RHEL support ends. The five year reign of RHEL7 was too long and did result in the very issues you bring up, but the ~3 year duration of RHEL 5,6 & 8 was short enough to avoid problems due to attrition in enterprise settings (unlike startups which have higher turnover, and not counting bus factors of one - no release cycle can't solve that). And like others have pointed out, automation doesn't help as much when moving between releases. We have everything configuration controlled with kickstart and ansible and/or docker, and it is great for reproducibility within a release cycle, but it doesn't save much time or knowledge between releases. And Ubuntu is even worse in that regard despite having a shorter release cycle.
- darkhelmet 3y agoIt's one of the things that ground Yahoo to a halt. We spent years migrating from RHEL-4 to 6, then RHEL-6 to RHEL-7, and by the time the projects were pretty much complete, the next sunset was approaching. My cynicism comes from seeing the bad things that "Enterprise Linux" enabled there. Admittedly, Yahoo was an extreme case. It never solved the really building problem - the culture from the early days was to compile, ship and forget. Once a RHEL-6 package was pushed to our dist/yinst system (packages), it would never be rebuilt unless it was 1) necessary, or 2) It was time to try and figure out how to build it on RHEL-7. A lot of effort was spent in the later years to try and address this (by burning the old tech stack to the ground), but the culture was pervasive for the longest time. If 10-year-RHEL didn't exist we would have been forced to address the building processes. If it's hard or error prone, then do it frequently until you get the process nailed down.
- gjvc 3y agoIf it's hard or error prone, then do it frequently until you get the process nailed down. Major life lesson -- practice makes perfect.
- bluGill 3y agoYears back (late 1990s) I worked with some mainframe people. They bragged that everything is hot fixable on the fly so they can apply security updates or replace broken hardware without rebooting. Then they admitted they schedule a reboot every 6 months anyway. Turns out the redundant backup power supply failed in at the same time and one hot patch was not applied to startup scripts and it took a week to figure out what was missing so the system booted again. By rebooting every 6 months they remember everything and so can get the system back up. I probably have some details wrong in the story above. I worked with those people, but never on the mainframe. I think the point stands though, if you don't do something often it can't be done.
- wmf 3y agoThis doesn't surprise me. Mainframes aren't just about never failing; they have a whole culture, including ops, around providing availability in ways that actually work.
- pjmlp 3y agoStarting by having systems programming languages that actually have proper strings, arrays and bounds checking.
- deleted 3y ago[deleted]
- reverius42 3y agoWhat are some of those languages? I'm curious to learn more.
- pjmlp 3y agoSeveral PL/I dialects e.g. PL/S and PL.8, BLISS, Modula-2, ESPOL/NEWP for example. Also Pascal and BASIC compilers with several extensions, e.g. VMS Pascal and VMS BASIC.
- maxclark 3y agoI have a customer with systems so old they weren’t at risk for heartbleed. He was excited about that.
- geerlingguy 3y agoHeh, same but for the Java Log4j vulnerability. "We haven't upgraded in 10 years, and it's secure from that!"
- bandrami 3y agoThat was anybody who used Debian stable
- mfer 3y agoThink of all the places Linux runs. Planes, trains, and automobiles. Medical equipment. So many other places. Many places that don't have readily available network access. Yet, many "enterprises" need support here. If a medical device or train works but needs support for years and years, should someone be constantly updating Linux? What about the software that runs on Linux and is tested there? Considering just the modern cloud environment really limits where enterprise Linux runs and is useful. And, where there are calls for really long support contracts.
- voakbasda 3y ago> If a medical device or train works but needs support for years and years, should someone be constantly updating Linux? What about the software that runs on Linux and is tested there? Yes. If you have embedded software in the field and it is running on hardware that has not reached its EOL, then you absolutely should be fixing bugs and vulnerabilities, doubly so when that hardware is attached to any kind of network, and triply so when the software talks to some kind of cloud services. For most products, the customer should have the power decide when that hardware reaches EOL. In other words, it should be illegal (and severely punished) to disable or downgrade devices remotely, whether by abdicating the responsibility to maintain their software or by shutting down network services that those devices require to operate fully. At the very least, that would prevent the proliferation of pervasive networking features that have no business communicating with anyone except their owners.
- civilitty 3y ago> If you have embedded software in the field and it is running on hardware that has not reached its EOL, then you absolutely should be fixing bugs and vulnerabilities, doubly so when that hardware is attached to any kind of network, and triply so when the software talks to some kind of cloud services. Talk to Qualcomm and NXP and MediaTek and Broadcom and STMicroelectronics and.... Seriously, most of this is completely out of the product engineers' control. We can't update the hardware because our vendors (ALL of our vendors) control the kernel selection, almost never upstream, and never keep the board support packages up to date. Never. IME the only exceptions to that rule are RaspberryPi (kinda sorta), AMD, and Intel. I only recently managed to finally extricate all my projects from Linux 2.6 which I considered a minor miracle. In 2023.
- _ea1k 3y agoHaving worked in both kinds of cultures, I tend to agree. Keeping up is ultimately less pain than trying to upgrade things in huge chunks. But it can be really hard to change the culture at a place that has a long history of "ain't broke, don't fix it" engineering.
- wkat4242 3y agoLess pain yes but more efficient? I'm not sure. The places that don't do constant upgrades also don't usually have teams looking after that. If they time it right they can do with less people. Of course it's less reliable not having as much active knowledge but I do think it can be cheaper if nothing goes wrong.
- _ea1k 3y agoYes, that can be the tradeoff, and a reasonable one in the right circumstances. Some projects are like that where there is really no team dedicated to anything more than keeping the lights on. In my comment, I was thinking of well staffed (or at least close-enough to well staffed) teams making deliberate decisions to defer.
- 2OEH8eoCRo0 3y agoI agree but I don't think that's SUSE's or Red Hat's problem. If you deliver a solid and stable product humans will get complacent.
- wkat4242 3y agoIn my experience updates are not forgotten. Even automated. Upgrades however are a different story. The major version changes require a ton of testing and manual massaging. This is why enterprises like to have that infrequently. For the systems that are easier you can still choose to follow the releases quickly. Because security patches are being back ported is usually not a real issue.
- rightbyte 3y agoIf certificate renewal is automatic why do it at all ...
- vbezhenar 3y agoTo ensure that those who possess the certificate, still control the domain. The main issue is that certificates should really be automated by every web server by default. At least for those with public IP addresses. There are servers like Caddy which implemented it, but it should be basic feature that just works without any additional configuration.
- nijave 3y agoLong term support allows bad behavior but I think it's still useful to reduce the amount of feature/breaking changes happening to software. Those problems can also be mitigated with mandatory environment rebuilds which is trivial for a lot of setups with infrastructure as code. At the extreme end, you have Kubernetes/CNCF where 6 months go by and you're many versions behind with a huge changelog of breaking changes you have to fix first. Stable APIs and stable ABIs are very useful here (which enterprise Linux provides).
- beefield 3y agoSome time ago I was listening a guy talking about (operational) risk management in financial industry. One of his main points was that the systems in financial organizations should not be completely bug-free and automated. Because when something eventually happens, if there is no-one who has had to fix issues in the systems regularly, there is nobody around who can fix the system efficiently. An example of argument that belongs to a weird class of arguments you at the same time want to agree and disagree.
- sidpatil 3y ago> There is another problem that wasn't covered in the article. The 10+ years of stability leads to behaviors and outcomes that remind me of the long-lived SSL certificate problem. Updating is done so infrequently that the "how?" is forgotten. As the 10 year support limit approaches, most of the old team members who did it last time are gone, tech debt is through the roof, few people know where everything is or how to build it, and so on. This is known as the out-of-the-loop performance problem. [1] [1] https://en.wikipedia.org/wiki/Out-of-the-loop_performance_problem https://en.wikipedia.org/wiki/Out-of-the-loop_performance_pr...
- otabdeveloper4 3y ago> LetsEncrypt did us a huge favor by forcing automation vs having the guy who knows how to update the SSL certs every 4.9 years and left 6 months ago. Not in my experience. There's still a guy who goes around and updates (manually) all the LetsEncrypt certificates every year.
- vanviegen 3y agoShouldn't he be going around every couple of months?
- bogeholm 3y agoWe truly live in amazing times! We have language models that sound human and internet from space, but never bothered to schedule that script for updating TLS certs. Or put it in version control for that matter. Sounds like my org :)
- outworlder 3y ago> Not in my experience. There's still a guy who goes around and updates (manually) all the LetsEncrypt certificates every year. LetsEncrypt certificates don't last for one year, they only last for 90 days, no exceptions. You may be thinking about something different.
- fragmede 3y agoThey may be talking about the certbot software itself, which does the updating of certs.