12 ms·
Debugging Hetzner: Uncovering failures with powerstat, sensors, and dmidecode
- V__ 2y ago> Looking back, waiting six months could have helped us avoid many issues. Early adopters usually find problems that get fixed later. This is really good advice and what I'm following for all systems which need to be stable. If there aren't any security issues, I either wait a few months or keep one or two versions behind.
- pwmtr 2y agoAuthor of the blog post here. Yeah, this is generally a good practice. The silver lining is that our suffering helped uncover the underlying issue faster. :) This isn’t part of the blog post, but we also considered getting the servers and keeping them idle, without actual customer workload, for about a month in the future. This would be more expensive, but it could help identify potential issues without impacting our users. In our case, the crashes started three weeks after we deployed our first AX162 server, so we need at least a month (or maybe even longer) as a buffer period.
- ThePowerOfFuet 2y ago>The silver lining is that our suffering helped uncover the underlying issue faster. Did you actually uncover the true root cause? Or did they finally uncap the power consumption without telling you, just as they neither confirmed nor denied having limited it?
- pwmtr 2y agoThe root cause was a problem with the motherboard, though the exact issue remains unknown to us. I suspect that a component on the motherboard may have been vulnerable to power limitations or fluctuations and that the newer-generation motherboards included additional protection against this. However, this is purely my speculation. I don't believe they simply lifted a power cap (if there was one in the first place). I genuinely think the fix came after the motherboard replacements. We had 2 batches of motherboard replacements and after that, the issue disappeared. If someone from Hetzner is here, maybe they can give extra information.
- oz3d 2y agohetzner is currently replacing motherboards of their dedicated servers [1] But I dont know if thats the same issue that was mentioned in the article. [1] https://status.hetzner.com/incident/7fae9cca-b38c-4154-8a27-14e6dfea5c1e https://status.hetzner.com/incident/7fae9cca-b38c-4154-8a27-...
- ubanholzer 2y agoThats the same issue, yes.
- axus 2y agoCustomers are the best QA. And they pay you too, instead of the reverse!
- rat9988 2y agoI'm pretty sure they pay for QA. QA cannot always catch every possible bug.
- knowitnone 2y agothese crashes should have been caught easily
- crishoj 2y agoWere you able to identify the manufacturer and model/revision of the failing motherboards? This would be extremely helpful when shopping for seconds hand servers.
- babuskov 2y agoI cannot find the link now, but it was mentioned that it was ASRock mobos.
- crishoj 2y agoThanks. This comment above does mention ASRock: https://news.ycombinator.com/item?id=43112594 https://news.ycombinator.com/item?id=43112594 On the other hand, dmidecode output in the article shows: Manufacturer: Dell Inc. Product Name: 0H3K7P
- InDubioProRubio 2y agoThis is a wildly successfully pattern in nature, the old using the young and inexperienced, as enthusiastic test units. In the wild for example in Forrest, old boars give safety squeaks to send the younglings ahead into a clearing they do not trust. The equivalent to that- would be to write a tech-blog entry that hypes up a technology that is not yet production ready.
- Tzela 2y agoJust for curiosity: do you have a source?
- esafak 2y agoGitHub is looking to add this feature to dependabot: https://github.com/dependabot/dependabot-core/issues/3651 https://github.com/dependabot/dependabot-core/issues/3651
- h1fra 2y agoIn theory, that works in practice nope. You get a random update with a possible bug inside that is only fixed by a new version that you won't get until later. The other strategy is to wait for a package to be fully stable (no update), and in that case, some packages that receive daily/weekly updates are never updated
- esafak 2y agoIt does help, because major version updates are more likely to cause breakage than minor ones, so you benefit if you wait for a few minor version updates. That is not to say minor versions can't introduce bugs. Windows is a well-known example; people used to wait for a service pack or two before upgrading.
- ajmurmann 2y agoWe could even wait for a patch version or the minor being out a certain amount of time. For a major I'd wait even longer and potentially for a second patch.
- Cthulhu_ 2y agoAnd then they went towards a more evergreen update strategy, causing some major outages when some releases caused issues. I mean evergreen releases make sense imo, as the overhead of maintaining older versions for a long time is huge, but you need to have canary releases, monitoring, and gradual rollout plans; for something like Windows, this should be done with a lot of care. Even a 1% release rate will affect hundreds of thousands if not millions of systems.
- deleted 2y ago[deleted]
- fdr 2y agoIt varies by system. As the legendary (to some) Kelly Johnson of the Skunk Works had as one of his main rules: > The inspection system as currently used by the Skunk Works, which has been approved by both the Air Force and the Navy, meets the intent of existing military requirements and should be used on new projects. Push more basic inspection responsibility back to the subcontractors and vendors. Don't duplicate so much inspection. But this will be the only and last time Ubicloud does not burn in a new model, or even tranches of purchases (I also work there...and am a founder).
- deleted 2y ago[deleted]
- vitus 2y ago> To increase the number of machines under power constraints, data center operators usually cap power use per machine. However, this can cause motherboards to degrade more quickly. Can anyone elaborate on this point? This is counter to my intuition (and in fact, what I saw upon a cursory search), which is that power capping should prolong the useful lifetime of various components. The only search results I found that claimed otherwise were indicating that if you're running into thermal throttling, then higher operating temperatures can cause components (e.g. capacitors) to degrade faster. But that's expressly not the case in the article, which looked at various temperature sensors.
- tecleandor 2y agoYep, that's weird, I've always read that high power/temp can degrade electronics way faster. Any EE can shed a light here?
- avian 2y agoAs an electronics engineer I have no idea what the author is talking about here and was about to post the same question.
- OptionOfT 2y agoThe only place I could find some answer that sheds some light was StackOverflow: https://electronics.stackexchange.com/a/65827 https://electronics.stackexchange.com/a/65827 > A mosfet needs a certain voltage at its gate to turn fully on. 8V is a typical value. A simple driver circuit could get this voltage directly from the power that also feeds the motor. When this voltage is too low to turn the mosfet fully on a dangerous situation (from the point of view of the moseft) can arise: when it is half-on, both the current through it and the voltage across it can be substantial, resulting in a dissipation that can kill it. Death by undervoltage.
- pwmtr 2y agoAt the time of our investigation, we found few articles supporting that power caps could potentially cause hardware degradation, though I don't have the exact sources at hand. I see the child comment shared one example, and after some searching, I found a few more sources [1], [2]. That said, I'm not an electronics engineer, so my understanding might not be entirely accurate. It’s possible that the degradation was caused by power fluctuations rather than the power cap itself, or perhaps another factor was at play. [1] https://electronics.stackexchange.com/questions/65837/can-electronics-be-damaged-by-under-currenting-it https://electronics.stackexchange.com/questions/65837/can-el... [2] https://superuser.com/questions/1202062/what-happens-when-hardware-tries-to-draw-more-power-than-power-supply-can-provid https://superuser.com/questions/1202062/what-happens-when-ha...
- TacticalCoder 2y ago[dead]
- chronid 2y agoWe will never know, but I wonder if it could be a power/signaling or VRM issue - the CPU non getting hot doesn't mean something else on the board has gone out of spec and into catastrophic failure. Motherboard issues around power/signaling are a pain to diagnose, they will emerge as all sort of problems apparently related to other components (ram failing to initialize and random restarts are very common in my experience) and you end up swapping everything before actually replacing the MB...
- jonatron 2y agoAt a previous company, devops would regularly find CPU fan failures on Hetzner. That's in addition to the usual expected HD/SSD failures. You've got to do your own monitoring, it's one of the reasons why unmanaged servers are cheaper than cloud instances.
- jeffbee 2y agoI regularly find broken thermal solutions in azure and when I worked at Google it was also a low-level but constant irritant. When I joined Dropbox I said to my team on my first day that I could find a machine in their fleet running at 400MHz, and I was right: a bogus redundant PSU controller was asserting PROCHOT. These things happen whenever you have a lot of machines.
- tryauuum 2y agoin my (limited) experience this only happened with GIGABYTE servers very weird behavior, I'd prefer my servers to crash instead of lowering frequency to 400MHz.
- dijit 2y agoI've seen it on nearly every brand, I have some Lenovo Servers in the basement that also down-clock if both PSU's aren't installed. I have alerts on PSU's and frequency for this reason. The servers are so cheap that overcommitting them by double is still significantly cheaper than using cloud hosting, which tends to have the same issue only monitoring it is harder. Though most people using cloud seem to be happy not to know and it's been a known thing that there's a 5x variation between instances of the same size on AWS.: https://www.brendangregg.com/Slides/AWSreInvent2017_performance_tuning_EC2.pdf https://www.brendangregg.com/Slides/AWSreInvent2017_performa...
- jeffbee 2y ago> I'd prefer my servers to crash instead of lowering frequency to 400MHz. 100% agreed. There is nothing worse than a slow server in your fleet. This behavior reeks of "pet" thinking.
- scottcha 2y agoI’d like to see what cpu governor is running on those systems before assuming a power cap is in place. Lots of defaults installs of Linux ship with the power save governor running which is going to limit your max frequencies and through that the max power you can hit.
- __m 2y agoschedutil on mine scheduled for mainboard replacement
- wink 2y ago> One of the providers we like is Hetzner because of their affordable and reliable servers. > In the days that followed, the crash frequency increased. I don't find the article conclusive whether they would still call them reliable.
- aduffy 2y agoTo their credit they actually fixed the problem. Good luck getting this level of support from any of the big 3 public cloud providers.
- frenchtoast8 2y agoFor example, AWS's Mac machines frequently run into hardware failures. My current job runs a measly 5 mac1.metal hosts for internal testing, and we experience hardware failures on these machines a few times a year. Doesn't sound like a lot, but these machines are almost always completely idle, and we almost never get host failures for Linux hosts. To make matters worse, sometimes a brand new instance needs replacement before it even comes up for the first time, which is annoying because you are billed a minimum of 24 hours for these instances. People have been complaining about this for years and seemingly nothing is being done about it. https://www.reddit.com/r/aws/comments/131v8md/beware_of_broken_macos_servers_mac1metal_on_aws/ https://www.reddit.com/r/aws/comments/131v8md/beware_of_brok...
- janc_ 2y agoThe main difference being that you talk with real humans who try to help you, not computer programs designed to give you an illusion…
- dumbledoren 2y agoRight. When you see the curt, cranky and blunt responses of Hetzner engineers in the support threads, you know you are talking to an actual engineer who can fix your stuff and not some American-style 'We are sorry you are experiencing this problem!' type of support rep who cant do anything about it.
- 2y ago
- vednig 2y agoas a CI/CD provider wouldn't it benefit if Ubicloud had their own servers?
- eitland 2y agoThey are in the early stages. I think the website said they recently raised 16 million euros (or dollars). Making investments into data centers and hardware could burn through that really quick in addition to needing more engineers. By using rented servers (and only renting them when a customer signs up) they avoid this problem.
- vednig 2y agounderstood, would love to know about it from founders tho, and what went through in their decision
- fdr 2y agoGP is more or less correct. Building and owning an institution that finances, racks, services, networks, and disposes of servers, both takes time and increases the commitment level. Hetzner is month to month, with a fixed overhead for fresh leasing of servers: the set-up fee. This is a lot to administer when also building a software institution, and a business. It was not certain at the outset, for example, that the GitHub Actions Runner product would be as popular as it became. In its earliest form, it was partially an engineering test for our virtual machines, and we went around asking friendly contacts that we knew would report abnormalities to use it. There's another universe where it only went as far as an engineering test, and our utilization and revenue pattern (that is, utility to other people) is different.
- immibis 2y agoDepends how many they need and how much control. Do they want to be a server company or an adapting-servers-to-run-your-CI/CD company or both? You can extract value from both parts of the equation, but theoretical economics tells us you can get the most value for the least effort by doing more of what you're best at and paying someone else to do what they're best at, rather than doing everything mediocrely yourself. Sometimes that other company isn't actually very good and you can increase value by insourcing their part of your operation. But you can't assume that is always the case. It wouldn't have solved this particular problem - I think we can safely guess that your chance of getting a batch of faulty motherboards is at least as high as Hetzner's chance.
- rikafurude21 2y agoSimilar thing happened to a AX102 I currently use, something related the network card which caused crashes. Thankfully hetzner support was helpful with replacement hardware. caused quite some grief but at least it was a good lesson in hardware troubleshooting. Worth it to me personally
- yread 2y agoYep same here. AX102 crashes with almost no load, nothing in the logs, won't come on. Hetzner looked at it multiple times and found either nothing or replaced cpu paste or a PSU connector. I migrated to AX162 and so far so good
- jaigupta 2y agoSame here. Hetzner found no issues with hardware in diagnostics, they insisted it is related to OS/Software side but on my request they changed hardware which fixed issue.
- rikafurude21 2y agoI Was given the choice between diagnostics and hardware replacement, and decided to do the diagnostics first, which turned out nothing. Decided to reinstall the os and when that didnt fix I was sure it had to be something related to hardware which diagnostics didnt catch. If you have a server mentioned in here and problems turn up, just get the hardware replacement immediately https://docs.hetzner.com/robot/dedicated-server/general-information/mainboard-replacement-for-several-dedicated-servers/ https://docs.hetzner.com/robot/dedicated-server/general-info...
- andai 2y ago> Hetzner didn’t confirm or deny the possibility of power limiting What are the consequences of power limiting? The article says it can cause hardware to degrade more quickly, why? Hetzner's lack of response here (and UbiCloud's measurements) seems to suggest they are indeed limiting power, since if they weren't doing it, they'd say so, right?
- radicality 2y agoRelated and perhaps useful: I’ve seen this in multiple cloud offerings already, where the cpu scaling governor is set to some eco-friendly value, in benefit to the cloud provider and in zero benefit to you and much reduced peak cpu perf. To check, run `cat /sys/devices/system/cpu/cpu/cpufreq/scaling_governor`. It should be `performance`. If it’s not, set it with `echo performance | sudo tee /sys/devices/system/cpu/cpu/cpufreq/scaling_governor`. If your workload is cpu hungry this will help. It will revert on startup, so can make it stick, with some cron/systemd or whichever. Of course if you are the one paying for power or it’s your own hardware, make your own judgement for the scaling governor. But if it’s a rented bare metal server, you do want `performance`.
- chpatrick 2y agoIs there any downside to ondemand? If your servers aren't running at 100% then there's no point wasting watts, even if you aren't paying for them, right?
- anarazel 2y agoIt performs terrible if you have an intermittent workload. Like e.g. a request response workload where request processing is cheap (so that the time to increase the frequency matters). I've seen cases it's a more than 2x request latency increase. It can be pretty annoying, because it means that systems can perform better under higher load and that you get drastically different latency depending on whether a request is scheduled on a core that just processed another request (already at high freq) or one that was idle. And because the frequency control isn't fun enough, this behavior also exists with cpu idle states. Even at high frequency Linux can enter idle states... I've debugged several cases where this set of issues has caused unintuitive behavior. E.g. a) switching to a more powerful servers drastically increased latency b) optimized code resulting in higher latency / lower throughout because that provided enough idle cycles for a deeper idle time between requests c) slightly increased IO latency leading to significantly worse overall performance, due to the IO getting long though to clock down
- nik736 2y agoMost other AX models (AX42, AX52 and AX102) also have serious reliability issues, where they will fail after some months. They are based on a faulty motherboard. Hetzner has to replace most, if not all, motherboards for servers built before a certain date over the next 12 months [0] [0] https://docs.hetzner.com/robot/dedicated-server/general-information/mainboard-replacement-for-several-dedicated-servers https://docs.hetzner.com/robot/dedicated-server/general-info...
- babuskov 2y agoI have two AX42's. One has been stable since I got it during the Eurocup discount period. The other got replaced 2 times so far, but it looks like the latest replacement is holding up. So, it's like 50% failure rate based on my small sample. I guess only Hetzner and ASRock know the real numbers.
- gtirloni 2y agoAnyone got experience with Ubicloud's OpenStack stack?
- jauntywundrkind 2y ago> To increase the number of machines under power constraints, data center operators usually cap power use per machine. However, this can cause motherboards to degrade more quickly. This was something I hadn't heard before, & a surprise to me.
- dangoodmanUT 2y agois there a provider that's like bare metal, but would detect these kinds of things mostly automatic? E.g. faulty or constantly crashing hardware.
- greggyb 2y agoManaged servers: https://www.hetzner.com/managed-server/ https://www.hetzner.com/managed-server/ There are also others, but Hetzner is under discussion here.
- Tijdreiziger 2y agoManaged servers are quite a different product, closer to ‘old-school’ shared webhosting. You don’t get root access, but you do get a preinstalled LAMP stack and a web UI for management.
- urbandw311er 2y agoWould anybody with data center experience be able to hazard a guess on what type of commercial resolution Hetzner would have reached with the Motherboard supplier here? Would we assume all mobos replaced free of charge plus compensation?
- wmf 2y agoWhen you buy name-brand servers you'll definitely get any faulty hardware replaced. Compensation would only happen if you negotiated for that and you'd have to pay extra. You're probably better off buying some kind of business interruption insurance instead of trying to get vendors to pay you for downtime (even if it is their fault). Hetzner is not a normal customer though. As part of their extreme cost optimization they probably buy the cheapest components available and they might even negotiate lower prices in exchange for no warranty. In that case they would have to buy replacement motherboards.
- babuskov 2y agoI think they probably got a batch of these really cheap in the first place, because those servers were offered without the setup fee initially. It was during the soccer World Cup in Germany.
- bayindirh 2y agoDell has this problem sometimes. I remember getting the first batch one of their older servers when they were new. We had to replace motherboards' I/O (rear) section because the servers lost some devices on that part (e.g.: Ethernet controllers, iDRAC, sometimes BIOS) for some time. After shaking out these problems, they ran for almost a decade. We recently retired them because we worn down everything on these servers. From RAID cards to power regulators. Rebooting a perfectly running server due to a configuration change and losing the RAID card forever because electron migration erode a trace inside the RAID processor is a sobering experience.
- merb 2y agoDell has tons of issues. A faulty mini board of the front led can actually stop the server from booting/running at all (even drac will be dead)
- bayindirh 2y agoInteresting. From my experience, Dell is generally one of the least problematic brands when compared in large numbers. Another surprising name is Huawei. Their servers just don't die.
- merb 2y agowell tbf their server pro support is actually good, but we still had a lot of minor issues, we barely had a dead one in Production. Most problems arises after unpacking. Like the one I told. Of course we had dead transrecievers and dead hard drives, but the pro support guys only want the support zip that you can create from the drac and they ask for some details and than they order replacements parts for you.
- bayindirh 2y agoAh, I understand what you go through now. In our case, they came with the parts that we said we gonna need, see that whatever device is dead with their eyes, and just replaced the problematic part. When the BIOS or the iDRAC is shot, there's no way they gonna get their support ZIP file. If they want they can connect that dead I/O board to a spare part and try. :)
- indulona 2y agoi am so glad my sign up process with hetzner failed when i was so dumb that i wanted to give them a chance even with the internet full of horrific stories of bad experiences from their customers. lucky me.
- cbozeman 2y agoHetzner is fine for what it is, you just need to know that it's all on you and only YOU. YOU do the monitoring. YOU do the troubleshooting. YOU etc., etc. If that doesn't appeal to you, or if you don't have the requisite knowledge, which I admit is fairly broad and encompassing, then it's not for you. For those of you that meet those checkboxes, they're a pretty amazing deal. Where else could I get a 4c/8t CPU with 32 GB of RAM and four (4) 6TB disks for $38 a month? I really don't know of many places with that much hardware for that little cost. And yes, it's an Intel i7-3770, but I don't care. It's still a hell of a lot of hardware for not much price.
- nobankai 2y ago[flagged]
- jaigupta 2y agoWe had been colocating servers from decades but there is too much "YOU", compared to that we find Hetzner doing a lot for us (hardware inventory, replacement, remote hands, networking etc). We are slowly moving away from colocating to renting at Hetzner. It is so much better.
- trod1234 2y agoIt would have been nice if they linked to the power metrics for the new servers. I think it would be amusing if it turns out they just raised the power limits for those servers not showing the problem up to base that was originally advertised.