9 ms·
“Unable to perform AVX2 instructions correctly under heavy load” is also a common “WTF Intel!?”–inducing phenomenon. I’m certain SREs who work at companies with
by bdd 7y ago
“Unable to perform AVX2 instructions correctly under heavy load” is also a common “WTF Intel!?”–inducing phenomenon. I’m certain SREs who work at companies with more than 1 million servers have a bunch of hair pulling stories.
Most (all?) Intel server CPUs in fact decrease clock speed when executing AVX2 (and some other) instructions to keep things a bit more sane. Vlad from Cloudflare wrote about this, more specific to AVX-512 back in 2017: https://blog.cloudflare.com/on-the-dangers-of-intels-frequency-scaling/ https://blog.cloudflare.com/on-the-dangers-of-intels-frequen...
Then there is PROCHOT signal. Which is supposed to protect the CPU from getting too hot but keeps getting raised in lopsided AVX2 loads not because CPU is too hot but voltage regulation gets whacked.
You may wonder: what is an example of AVX2 heavy load. RSA multiplication is a good candidate. AES constructions or modes (CBC with SHA, GCM) are implemented in AVX2-BMI2 as well.
- fivesixzero 7y agoI’m curious if this behavior defined by something in hardware, microcode, boot-time BIOS flags, or higher level kernel/hypervisor/application code.
- gameswithgo 7y agoon unlocked intel cpus you can change the avx multiplier in the bios.
- olliej 7y agoOh neat (I haven’t messed with over clocking in years) - is it just avx that you can tailor? (Beyond the old school bus multipliers)
- gameswithgo 7y agoThe newest AMD CPUs decrease clock speed based on parameters like heat, which AVX2 under heavy load will cause. So, AMD also decreases clock speed when executing AVX2, indirectly. Though in a very different fashion, continuously, rather than a distinct mode.
- TD-Linux 7y agoI do prefer Intel's solution of slowing down over AMD's solution of crashing. I guess the good news is if this does turn out to be a power delivery related bug, it's fixable with firmware.
- bdd 7y agoMaybe it wasn't very clear in my comment but "correctness issues" more-or-less literally mean the instruction's result is incorrect. Like you ask 2*2 and it gives you 0xfff1309f. So depending on how you rely on that result, you may crash. Extending on my examples, bad RSA multiplication means failed authentication. Bad AES construction means encrypting things to garbage, impossible to decrypted by the recipient. So, let's say if you rely on a result from such instruction to do memory references, you're definitely going to crash.
- olliej 7y agoSlowing down and producing the correct answers is clearly superior to crashing and/or incorrect answers. I still don’t understand what it is about avx2 that results in these kinds of issues - is it really just a matter of increased number of execution units running at once causing weird power and heat issues?
- wtallis 7y ago> is it really just a matter of increased number of execution units running at once causing weird power and heat issues? That's been the root cause of all the previous AVX weirdness I'm aware of. It's probably the case for these Threadrippers, too. Modern CPUs have very high power density and run at pretty low voltage, which results in insane current delivery requirements.
- paulmd 7y agoNot increased number of execution units, but increased number of transistors, yeah. A normal multiply does 1 number at a time (32 bit for simplicity). AVX2 can do up to 8 multiplications at a time. That's a huge amount more transistors firing all at once and that causes the voltage to start to droop. AVX-512 takes that even further and now it's 16 multiplications per unit, oh and Intel moved from 1x256-bit unit per core on Haswell/Broadwell to 2x512-bit units on Skylake-X, so it's potentially 32 multiplications at a time - 4x as much as AVX2. Basically to prep for that much power being drawn all at once, the chip has to switch to a higher-voltage mode to account for the voltage droop caused by all those transistors switching at once in one place. It takes time for the regulator (on-chip Fully Integrated Voltage Regulator or motherboard VRM) to wind up the voltage far enough. https://travisdowns.github.io/blog/2020/01/17/avxfreq1.html https://travisdowns.github.io/blog/2020/01/17/avxfreq1.html https://travisdowns.github.io/assets/avxfreq1/fig-volts256-1.svg https://travisdowns.github.io/assets/avxfreq1/fig-volts256-1... At this level behavior is intensively analog and thermals/voltage both significantly affect transistor current draw and switching time, which feeds back into thermals and power consumption/voltage droop. This gets even more problematic on 7nm/10nm type nodes and especially in GPUs where you are doing a huge amount of vector arithmetic all the time. Essentially it is no longer possible to design processors that are 100% stable under all potential execution conditions, or even under normal operating conditions, so you have to have power watchdog circuitry that realizes when it's getting close to brownout/missing its timing conditions and slows itself down to stay stable. That's why AMD introduced clock stretching in a big way with Zen2 (despite the fact that it's nominally been around since Steamroller). NVIDIA piloted this with Pascal, AMD piloted it with Vega and brought it to CPU with Zen2. You simply cannot design the processor to be 100% stable at competitive clocks anymore. You have to have power management that's smart enough to withstand small transient power conditioning faults. https://semiengineering.com/managing-voltage-drop-at-107nm/ https://semiengineering.com/managing-voltage-drop-at-107nm/ https://semiengineering.com/power-delivery-affecting-performance-at-7nm/ https://semiengineering.com/power-delivery-affecting-perform... https://www.realworldtech.com/steamroller-clocking/ https://www.realworldtech.com/steamroller-clocking/
- jedbrown 7y agoThis is a deep dive on frequency scaling and IPC throttling related to AVX512 instructions. The consequences are quite large, surprisingly complicated, and persist for an eternity, which is why you really have to coax the compiler if you want it to issue these instructions. https://travisdowns.github.io/blog/2020/01/17/avxfreq1.html https://travisdowns.github.io/blog/2020/01/17/avxfreq1.html
- sarosh 7y agoDiscussion here: https://news.ycombinator.com/item?id=22077974 https://news.ycombinator.com/item?id=22077974
- paulmd 7y agoPrime95 has been a bad test for years now. Literally since the introduction of AVX2, it has been known that overclocking and running Prime95 will cause rapid electromigration and processor degradation on Haswell-E if you attempt to overclock. It's fine for like a few minutes but you should not attempt to do 24 hour Prime95 runs like people used to do, that was in the days before AVX was a thing. https://rog.asus.com/articles/usa/rog-overclocking-guide-core-for-5960x-5930k-5820k/ https://rog.asus.com/articles/usa/rog-overclocking-guide-cor... It also does not test all parts of the processor equally, it really just is slamming the cache and the AVX units. You can be "Prime95 stable" and still crash in other things. It's just not a good test anymore in multiple respects. Same thing for Furmark on GPUs. And nowadays GPUs are actually smart enough to realize they're running a power virus and the power management pulls back the power. So again, it doesn't demonstrate anything. Zen2 is supposed to work the same way, if it realizes you're redlining the chip then it should be clocking down somewhat.
- jrockway 7y agoNo sequence of instructions should cause the machine to reboot.
- wtallis 7y agoThat's an excusable outcome if you're overclocking or otherwise tweaking critical operating parameters beyond their safe default ranges.
- marmaduke 7y agoExcept the 'reboot' command? I'm always curious how it works.
- deleted 7y ago[deleted]
- monocasa 7y agoNormally you'd ask the chipset so that it can reset all the peripherals at the same time. On Intel you can also intentionally "triple fault" the CPU and that'll cause it to reset. The way that works is you tell the CPU that it doesn't have interrupt handlers anymore, then cause a software interrupt. Then the CPU will first try to call your interrupt handler (but that'll fail because you told it there are no handlers), then errors during an interrupt cause a "double fault" interrupt that's supposed to handle those errors (but once again, there's no handlers), so now you're in a "triple fault" condition and the only way to recover is to reset, so that's what the CPU does.
- retrovm 7y agoYou know I think that cloudflare blog post is a great example of the engineering approach behind boringssl. It's optimized for actual workloads, where you decrypt or authenticate a short message and then move on to different activities, and engaging AVX512 doesn't actually pay off in reality. OpenSSL is optimized to produce the biggest number from `openssl speed` so of course in that light it makes perfect sense to enable AVX512. But if you're trying to use these libraries in realistic workloads you will begin to appreciate the boringssl approach.
- jrockway 7y agoI ran into this issue on one of my builds. Aida64 has a benchmark (floating point photo or something?) that uses AVX instructions. Pressing the "run benchmark" button would instantly black-screen crash my machine with 100% certainty. I debugged this problem over a number of years... I replaced the RAM, I replaced the motherboard, I eventually replaced the CPU... still happened no matter what I did. Even if I underclocked the machine and kept the voltages the same, instant crash. It was maddening. Exasperated, I eventually busted out an oscilloscope and looked at the waveform on the 12V supply to the CPU. When starting the AVX benchmark, there was a huge brownout. That basically explained everything; my power supply essentially turned off when the CPU started drawing a ton of power. I replaced the power supply and got lucky -- it handled it fine and I could run the benchmark. I even got some overclock out of it. After this whole experience, I've never looked at computers the same way again. You can buy high-spec brand-name components, and it's all just a crapshoot. Maybe your computer won't crash in the middle of an important task. Maybe it will. There isn't much you can do but cross your fingers.
- wtallis 7y agoYou can do a little bit to help by ensuring you shop for the right kind of high-spec: workstation rather than gaming, so that the extra money goes toward useful engineering and QA rather than RGB LEDs. But even then, it often ends up that the best you can hope for is a long warranty with a quick and easy replacement process.
- CoolGuySteve 7y agoI've had enough failed server and workstation components to believe this is snake oil. It seems like a lot of "Enterprise" hardware (particularly hard drives) is mostly the same chipset/components with a longer warranty and a more boring PCB color. Or more annoyingly, differentiated by feature lockout in some firmware. Gaming hardware has larger economies of scale which leads to more people running the hardware and submitting RMAs. Gamers are also more likely to overvolt their hardware, leading to most consumer components being overspec'ed for heat and power. As such, the second revision of a gaming board imo is more refined than a workstation/server board that almost never gets a new revision. On a more practical level, consumer CPUs usually have higher clockrates and newer architectures and gaming NVidia GPUs are so much cheaper than the enterprise product lines that you can buy multiple. My workflow is faster with the cheap stuff, sometimes twice as fast.
- jotm 7y agoIntel's TurboBoost is such a marketing mess. It used to be more sane when they stated base clocks - the guaranteed frequency you pay for. Anything above was just extra performance when your CPU is not running hot/at power limits. Nowadays it's just "up to XX GHz", and people expect it to run at those clocks. AVX really pushes the CPU, an Intel Core/Xeon will hit the power draw limit pretty fast. They don't decrease clock speed as much as fall back to base clocks. The ones you paid for. Anything above that is just a bonus. That said, you should never trigger PROCHOT even under full stress load with AVX! If you do, you need better cooling. It's a last resort throttling feature for when your processor is hot enough to boil water :/