8 ms·
EPYC 7002 CPUs may hang after 1042 days of uptime
- dale_glass 3y agoA machine staying up for almost 3 years is irresponsible in this day and age. Yeah, I remember people having uptime competitions on Slashdot and the like some decades back, but you only need to look at the ssh logs of a 5 minutes old machine to realize this is a terrible idea in modern times.
- sokoloff 3y agoAir gapped machines and kernel live patching both exist.
- cpach 3y agoAnd how many people use that? Most servers today are not air-gapped.
- j16sdiz 3y agoMost server don't do that, but those that do are not crazy
- agentgumshoe 3y agoHow many examples will you need before you say "oh ok, I can see some valid concerns."? I've worked in places where expensive Lab equipment is running off outdated PCs/servers because updates aren't available and they will absolutely stay on for as long as possible. We're not all silicon valley, things can be expensive and difficult to replace...
- BasedAnon 3y agoI have kernel live patching on my mother's computer because it means she has to know how to do less
- mnw21cam 3y agoKernel bugs are rare. Most (almost every single) vulnerability can be patched without rebooting.
- 0x0 3y agoThey're not that rare. Also, there are a lot of other updates that in practice should be followed up with a reboot. For example, any library consumed by systemd (such as openssl) usually requires pid1 to relaunch. For example, debian released an openssl update just yesterday. You can run "checkrestart -v" to try to figure out how to restart every affected app but you'll quickly run into systemd's init process running with the old vulnerable library loaded, and then you might as well just reboot to get a clean "checkrestart -v". Even just relaunching non-pid-1 applications like dbus can quickly create a mess where sshd logins get a delay if you're not careful to also reload everything that depends on it.
- Denvercoder9 3y ago> For example, any library consumed by systemd (such as openssl) usually requires pid1 to relaunch. That does not require a reboot, `systemctl daemon-reexec` is enough.
- znpy 3y agonice username, can i ask you what did you see?
- oynqr 3y agoYou don't need full system resets to get security updates. Kexec, live patching, userspace reboot.
- hardware2win 3y agoI dont understand opinons like this Just because it would be dangerous for your nodejs web_app.exe running on ubuntu behind apache fully exposed on the internet then there are billion other ways to use computers, like even air gapped systems. So, dont try to justify obvious flaw
- dale_glass 3y agoI mean, hardware is cheap enough that any server of importance should be individually disposable. Yeah, you can do stuff to maximize uptime but if it needs to stay up that badly you have to consider the case of the hardware needing to be turned off at some point. > So, dont try to justify obvious flaw I'm not, it's a bug and should be fixed. But I think if anything is powered for 3 years straight it's a bit concerning. Otherwise you're liable to find things like that somebody started something by hand 2 years ago, and at a critical moment nobody quite remember what the command was.
- defrost 3y ago> But I think if anything is powered for 3 years straight it's a bit concerning. Pretty much why Pawsey has an Annual High Voltage inspection shutdown [1] > Otherwise you're liable to find things like [..] TBH that's not really been an issue of note at any of the big iron farms I've been around since the 1980s .. generally there's a disciplined approach to maintaining 24/7/365 operation (that includes scheduled downtime for equipment checks) part of which is process documentation and justification and soft means of freezing | migrating processes+data etc. [1] https://status.pawsey.org.au/incidents/tk5n5y965r5j https://status.pawsey.org.au/incidents/tk5n5y965r5j
- PedroBatista 3y agoYou live in your own World with other people. Please just keep in mind there are many other Worlds with other people and laws of the Universe. I don't know if you're young or don't know much about history but what you describe is a fairly recent way of looking at things, it's not the only one and I guarantee you it will become "out of fashion".
- shpx 3y ago1042 days ought to be enough for anybody
- kuratkull 3y agoAre you perhaps a Windows user? In the Linux world updates don't necessarily require reboots.
- dale_glass 3y agoActually as of late, Linux has been moving towards rebooting for update. Yeah, you technically can replace on-disk files while services are running. In practice this can cause trouble if an application wants to read an updated file at the wrong time, and library dependencies can require restarting a lot of stuff. For ages people would install an update containing a security fix in glibc or libz or something, and keep on running the vulnerable version of the services that use them. At that point you might as well reboot. Modern Fedora has a very Windows-like mechanism where you reboot to update. You reboot, the system installs updates, then reboots again.
- viraptor 3y agoWhile Fedora did move towards that, it's not the only way. A lot of systems which require high reliability are built to reload correctly. At a generic system level, for example upgrading Nixos will pull new packages and put them next to the current ones, then reexec where possible. Nginx can replace its master process (SIGUSR2). Telephony software can often reexec and keep connecting open. Etc. Outside of desktops it's not that uncommon to do seamless live reloads of the whole system.
- justinclift 3y ago> Actually as of late, Linux has been moving ... That's a pretty broad generalisation. Which distro's are you meaning?
- magicalhippo 3y agoKDE Neon has done this. Before I had to reboot anyway because usually the desktop was full of random crashes if I updated without rebooting.
- ly3xqhl8g9 3y ago3 years is irresponsible? To quote Logan Roy, you, software developers, "are not serious people" [1]. Just out of curiosity looked for a list of longest running electrical devices [2]: 1840 - The Oxford Electric Bell 1871 – Souter Lighthouse in South Shields, UK 1896 – The Isle of Man’s Manx Electric Railway 1902 – The Centennial Bulb Apparently, "The Centennial Bulb has seen just two interruptions: for a week in 1937 when the Firehouse was refurbished, and in May 2013 when it was off for nine and a half hours due to a failed power supply." [1] https://www.youtube.com/watch?v=LZTaXjt2Ggk https://www.youtube.com/watch?v=LZTaXjt2Ggk [2] https://www.drax.com/electrification/4-of-the-longest-running-electrical-objects/ https://www.drax.com/electrification/4-of-the-longest-runnin...
- deleted 3y ago[deleted]
- dale_glass 3y agoThose are completely trivial complexity-wise compared to a modern server, and many don't have a real function, and mostly are artificially maintained as a curiosity. I mean, the centennial bulb barely glows, that's why it still works. The hotter the filament gets the faster it evaporates, so a light bulb that barely makes any light can stay working forever.
- ly3xqhl8g9 3y agoSure, was looking for electrical devices, a better example of what great engineering can achieve I suppose it's the Pons Fabricius [1], bridge built 2,085 years ago, still in use. The problem is, if we can't expect software to run essentially forever, to update without 'restarts', and so forth, how are we ever going to achieve neural chip implants, artificial organs, synthetic agents mining ore in outer space, and so on? Software is not a gear mechanism, a rack and pinion, there is absolutely no reason to restart an 'operating system' or to ever lose state, however we became accustomed and we commit these sort of crimes daily, restarts and refreshes. [1] https://en.wikipedia.org/wiki/Pons_Fabricius https://en.wikipedia.org/wiki/Pons_Fabricius
- cesarb 3y ago> A machine staying up for almost 3 years is irresponsible in this day and age. [...] but you only need to look at the ssh logs of a 5 minutes old machine to realize this is a terrible idea in modern times. You don't need to reboot a machine to update ssh. You only need to reboot the machine to update the kernel; for everything else, you just have to restart the corresponding user-space processes (and even PID1 can re-exec itself). Most kernel vulnerabilities are not remotely exploitable, so as long as you can trust your user-space processes (and keep them updated), it should be safe enough.
- kjs3 3y agoAs I recall, machines made by Tandem Computers, among other highly fault tolerant machines that have regrettably fallen out of fashion, didn't have to reboot even to replace the kernel. They didn't run Linux, tho.
- Neil44 3y agoIt seems a C6 state is an individual core sleeping. The intersection of people who don't reboot for 3 years and people who have sleep states enabled must be pretty small. It's an interesting bug though!
- vegardx 3y agoI had a very similar issue with some AMD-based servers (bulldozer, I think) about ten years ago. There was a bug where Xen-based virtual machines could set a C-state on cores it was assigned, but for whatever reason it wasn't able to wake them up. It was fun trying to figure out what the heck was going on.
- icybox 3y agoI have C-states already disabled because of old linux kernel bug where the kernel hang on Zen3 architecture. So not much to see here :)
- nicolaslem 3y agoDo you mean a bug in an old version of linux that is now fixed? Because I have been using Zen3 and Zen3+ on linux since their release and never had to mess with C-states.
- lUserAMD 3y agoEPYCs have many cores, and most applications (including those with long uptime requirements) use only a subset of the cores continuously. So it is totally normal for some of the cores to go to deep sleep C6 during phases of lower load. It will cause server operators headaches, when those cores don't come back eventually. Reboots help, disabling C6 in the (already running) OS also helps. Please note, that we are not talking about a core sleeping for three years. We are talking about a core going to deep sleep, when the system has been up for three years or longer.
- RobotToaster 3y agoI feel like some of the comments here are missing the point. Yes it's only likely to effect a small number of users, so did the intel fdiv bug, both are defective products. Back then intel were pressured into a recall, today we seem too willing to put up with being sold broken stuff.
- Waterluvian 3y agoIt feels a little bit different. One creates uncertainty in all floating point results, given you don’t know when it happens. The other requires you to reboot maybe every ~3 years and you know exactly when it happens. I’m not saying we should tolerate a defect, but it doesn’t feel nearly as problematic.
- gchamonlive 3y agoIt means anyone launching amd powered virtual machines on cloud providers can experience this now, at any point, and you don't know when it will happen, given this type of CPU could have been bought, booted or rebooted anytime in the past three years. Seems comparably problematic to me.
- Waterluvian 3y agoI’d prefer a bug that crashes a program than one that quietly inserts wrong data and keeps going.
- bee_rider 3y agoIt probably depends on your workload, which is a bigger deal. The fdiv bug was pretty bad, but at least fixable in software (at some cost). Anyway, recall is the right decision in either case (unless there’s a good enough workaround).
- __alexs 3y agoIf every CPU with an errata that needed software workarounds was recalled there would be no CPUs to use.
- deleted 3y ago[deleted]
- deleted 3y ago[deleted]
- mattpallissard 3y agoThis happened with a higher end Cisco switch (the model escapes me) we used in our core many moons ago. Stopped passing traffic completely after a number of days. At least Cisco told us about it themselves. We just fail-over rebooted until they fixed it.
- msla 3y agoPreviously: https://news.ycombinator.com/item?id=28340101 https://news.ycombinator.com/item?id=28340101 Watch Windows 95 crash live as it exceeds 49.7 days uptime [video]
- neilv 3y agoReminds me of the Intel Atom C2000 series brickings, circa 2017. https://www.anandtech.com/show/11110/semi-critical-intel-atom-c2000-flaw-discovered https://www.anandtech.com/show/11110/semi-critical-intel-ato... https://www.servethehome.com/intel-atom-c2000-series-bug-quiet/ https://www.servethehome.com/intel-atom-c2000-series-bug-qui...
- bushbaba 3y agoIn general it is good practice to have machines hard-restart every now and then. Otherwise you run into some weird edge-cases and rely too much on things being up and running 24x7x365
- tedunangst 3y agoThe good news is now I know why my server crashed last month, and it wasn't some other defect.
- SpaghettiCthulu 3y agoYou've had a Ryzen 7000 series CPU running for nearly 3 years already?
- eqvinox 3y agoMSR-poking Tool for Zen1 Ryzen CPUs to disable C6: https://github.com/r4m0n/ZenStates-Linux/blob/master/zenstates.py https://github.com/r4m0n/ZenStates-Linux/blob/master/zenstat... Not sure if this is applicable to EPYC CPUs, probably not. But I would expect that it's possible to disable C6 in some similar way on EPYC CPUs without rebooting the system. (If you are actually at risk of running into this issue, you likely don't want to reboot the system…)
- nh2 3y agoI filed a kernel bug 'System thrashes with "AMD-Vi: Completion-Wait loop timed out" after 247 days of uptime' for AMD Ryzen 7 3700X https://bugzilla.kernel.org/show_bug.cgi?id=217257 https://bugzilla.kernel.org/show_bug.cgi?id=217257 Might this be potentially related?