8 ms·
I found a bug in Intel Skylake processors
- fnsa 9y agoMr pthread!
- raverbashing 9y agoGCC generates code that's smaller but it isn't optimal, because of the potential for partial register stalls (and just overall register renaming issues) It's of course not wrong, but using AH when you're dealing with RAX is a weird anachronism Clang does the obvious, correct thing.
- barrkel 9y agoAH is no more of an anachronism in 64-bit code than it was in 32-bit code; AL, AH, AX, EAX, RAX; all bit slices of the same register. The addition of RAX doesn't change the fact that AH is still a bit weird (it's the only one that doesn't include the low bits of the total register).
- wereHamster 9y agoIt's not obvious. If it was then GCC would not do it. Please don't use the word 'obvious' when it is to you but not to the majority of population.
- raverbashing 9y agoIt is obvious, though I meant "obvious and correct" not "obviously correct"
- zrth 9y agoWho can clarify? Google says: Qbvious - easily perceived or understood; clear, self-evident, or apparent.
- lmm 9y agoThe confusion is between "the assembly clang generates is the obviously correct assembly for this operation" (true) and "the obviously correct behaviour for a compiler is to generate the assembly that clang does" (disputable, as while the assembly GCC generates is less clear, there may be advantages to doing it that way in e.g. performance). "Clang does the obvious, correct thing." is ambiguous between these two meanings.
- qb45 9y agoI disagree with "ambiguous", you sneakily changed "obvious" to "obviously" in the second sentence. obvious, correct != obviously correct
- shawnz 9y agoSounds like by "obvious" the parent means that Clang uses the most trivial approach.
- acdha 9y agoHave you looked at the actual GCC codebase? It's very easy to say something is obvious when you're looking at a problem which someone else has nicely isolated; it's much harder to dive into a complex codebase which has a very wide support matrix and say it's worth the effort to change working code instead of so many other things. More bluntly, before now wouldn't most people have said it was “obvious” that Intel would support their own documented features?
- raverbashing 9y agoThe code produced by Clang is a direct translation of the C code. That's the obvious part of it Most people here are not familiar with x86 assembly and its caveats it seems. Reading the Intel and AMD optimization manuals might be a good start (and yes, the bug is not in GCC it's on the Intel processor)
- acdha 9y ago> Most people here are not familiar with x86 assembly and its caveats it seems. Reading the Intel and AMD optimization manuals might be a good start. Please don't make unsupported assertions that everyone but you is speaking out ignorance. It doesn't add anything to the conversation, especially when dealing with older codebases unless you can prove that this is and never has been the correct way to write that code. Otherwise it's just another way to say “CPU optimizations change over time and an open-source project doesn't have a team of experts tracking microbenchmarks to decide when to switch”.
- raverbashing 9y agoShould be easy to prove me wrong if it's as easy as you say, I don't see why you're so annoyed by it > “CPU optimizations change over time and an open-source project doesn't have a team of experts tracking microbenchmarks to decide when to switch” Yes, that's what's happening, but the problem comes from the P6 architecture (though Netburst doesn't have those problems), this problem has been known for around 20 years
- wyldfire 9y ago
- tom_mellior 9y agoDo you know for a fact that Clang's code is faster here, i.e., have you measured it on actual hardware? Armchair performance estimation about "the potential" for something is often wrong...
- raverbashing 9y agoIt might not, but the GCC code has a potential issue in how it does things https://stackoverflow.com/a/41574531 https://stackoverflow.com/a/41574531 (curiously the question is about GCC not doing it, apparently not always)
- aij 9y agoIt's a performance trade off. It's not really surprising that GCC would make different performance trade offs when optimizing for different microarchitectures. (let alone most likely across different versions of GCC) Why would you expect GCC to optimize for Pentium 4 (Netburst) in this day and age? (Especially given that the article is talking about Skylake.)
- raverbashing 9y agoI don't expect it to optimize to P4, are you talking about this? (it's one of the answers to that answer) > The quoted delay of 5 - 6 clocks is much better on later microarchitectures. For example from Sandy Bridge and Ivy Bridge, The Ivy Bridge inserts an extra μop only in the case where a high 8-bit register (AH, BH, CH, DH) has been modified And you see that even on later architectures the point of avoiding AH makes sense (which is the opposite of what that GCC code does)
- rbehrends 9y agoIn a micro benchmark, using the clang version appears to be a bit faster. However, note that code size can also become a relevant concern in larger code bases: Apple, for example, is said to generally build its OS code with -Os rather than -O2 or -O3 (and had -Oz as an even more aggressive custom code size optimization added to gcc back when they still were using it instead of clang).
- userbinator 9y ago"smaller but not optimal" is not really true. It depends on what you're optimising for. In my experience, optimising for size overall, and then speed in the really performance-critical parts (with some expected expansion), gives the best results. Even the non-performance-critical code will have a noticeable effect if its larger size causes more cache misses. Making use of the "partial registers" (I see them more as separate smaller registers that can be grouped together) effectively can avoid many more instructions.
- raverbashing 9y agoI don't disagree with you, but then use EAX, not AH (which would also produce a smaller code as you don't need to use a 64-bit constant) Optimizing for size is good, but what GCC did there made sense in the 32-bit days, but not that much today Some code snippets use AH/AL as 2 separate registers hence the processor might rename them to different internal registers. But then when reading EAX the processor needs to update EAX accordingly as well.
- qb45 9y agoTFA says clang actually uses an encoding of AND which operates on RAX but takes only 32b of constant so it is pretty much the solution you propose. And BTW, operating directly on EAX itself fucks up the upper half of RAX.
- pcwalton 9y agoAH/BH/CH/DH are so seldom useful that it's not really worth adding support for using them to the compiler. They only exist for 4 of the 16 registers, and they only let you address one random byte of the 8 bytes in that register. You get a bit slower code and a more complicated register allocator in order to save, what, a few bytes in the entire program?
- tvgggghh 9y ago> Clang does the obvious, correct thing. Can we all stop shitting on GCC all the time? Thanks.
- pcarolan 9y agoHow does this bug affect current hardware in production? Is it worth waiting for the fix before buying the new MBPs, for example?
- gfiorav 9y agoI'd say no, you likely won't run into trouble because of this
- j_jochem 9y agoThis is subjective and very anecdotal, so take with a heap of salt: My 2016 Skylake MBP used to crash very regularly when waking it up from suspend (sometimes multiple times a day). When I first heard about this issue a week ago, I used XCode Instruments to disable Hypterthreading. I have not observed a single crash since.
- striking 9y agoEDIT: my comment was wrong. Thanks. There is a microcode update. Install it and you'll be fine. On macOS it should be installed automatically, according to https://support.apple.com/en-in/HT201518 https://support.apple.com/en-in/HT201518 (although that page doesn't say whether it's available or not)
- j_jochem 9y agoThe article states that Kaby Lake is affected, too. Maybe except not the Kaby Lake versions used in MBPs?
- willvarfar 9y agoA comp.arch poster said: > The errata refers to the problem showing up on short loops of less than 64 instructions that use AH, BH, CH or DH. > Looking at the Skylake microarch, the instruction decode queue is 128 uOps thread, 2*64 uOps when threaded. The Loop Stream Detector "can stream the same sequence of µOPs directly from the IDQ continuously without any additional fetching, decoding, or utilizing additional caches or resources." ... "capable of detecting loops up to 64 µOPs per thread". https://en.wikichip.org/wiki/intel/microarchitectures/skylake#.C2.B5OP-Fusion_.26_LSD https://en.wikichip.org/wiki/intel/microarchitectures/skylak... > So maybe the microcode update just shuts off the loopback detector. https://groups.google.com/d/msg/comp.arch/UkO4Z2FT18c/7YlC0aH7AQAJ https://groups.google.com/d/msg/comp.arch/UkO4Z2FT18c/7YlC0a... So if the bug is in the loop-detector, and the patch possibly disables it rather than fixes it, then does anyone have any before-and-after performance stats?
- Tuna-Fish 9y agoIIRC the Skylake loop buffer is not any faster than the uop cache, instead the reason for it's existence is to save power by not touching the cache. So you'd have to test power consumption instead?
- CalChris 9y ago>> Looking at the Skylake microarch, the instruction decode queue is 128 uOps thread, 2x64 uOps when threaded. No. Skylake does not have 128 μops with HT disabled. Skylake indeed was a big jump from Broadwell where the loopback buffer has 56 entries, 28 per hyperthread or 56 with HT off. Skylake has 64 μops per thread, HT on or off. 64 μops is a lot.
- gfiorav 9y agoAmazing report
- ihnorton 9y ago> I worked from the executable provided by SIOU, first interactively under GDB (but it nearly drove me crazy, as I had to wait sometimes one hour to trigger the crash again), then using a little OCaml script that ran the program 1000 times and saved the core dumps produced at every crash. rr can often be a time-saver in situations by providing deterministic replays up to the point of a crash, whereas coredump analysis is a single retrospective snapshot. http://rr-project.org/ http://rr-project.org/
- tdullien 9y agoAs much as I love RR, I am not sure it would've helped here, as the bug requires multiple threads to concurrently run? Also, RR is based on achieving deterministic replay IIRC, so I am not sure it'd be the first choice for a nondeterministic hardware bug?
- KenoFischer 9y agoYou can do multiple concurrent rr recordings. A non-deterministic cpu bug would generally cause a divergence in the recording vs replay, so rr would be a decent way to go about this. The way I'd have probably used rr when faced with this is to bisect the recording to find which code is responsible.
- nosefouratyou 9y agoThere's also UndoDB if you can pay for it. Not sure what the differences are between it and RR, but I know with UndoDB you can create a binary for the client that has the recorder built in, so it can automatically record failing states.
- joshuata 9y agorr wouldn't help in this case. From the docs, it "emulates a single-core machine. So, parallel programs incur the slowdown of running on a single core." The skylake bug only occurred under heavily threaded loads.
- jacquesm 9y agoHow will emulation trigger a hardware bug?
- etatoby 9y agoI will be surely downvoted for this, but I would like to remind everyone how this bug is just one of the many consequences of Microsoft's evil policy of encouraging the sale and distribution of proprietary software in executable form. There is no other reason why a 64bit multi-core CPU developed in 2015, that makes heavy use of pipelining and other advanced and complicated code execution strategies, would need to support instructions that address the second-to-last byte of a register (eg. %ah) while keeping the rest of the register 'unchanged', which of course means making a complete mess of the code execution path. The only reason this crap still exists is to keep Windows users' ability to run random EXE and DLL files from the 90s, if not random COM files from the 80s, at the expense of CPU cost, stability, and correctness for everyone else (such as the OCaml developers and users who ran into this bug.)
- acdha 9y agoDid you miss the part about where the bug was found using the current versions of GCC to build the current versions of OCaml? It's lazy to the point of dishonesty to act as if Microsoft is the only one with decades of accumulated code.
- jstimpfle 9y agoThere is still a point that proprietary binaries are probably the biggest force keeping decades of cruft in processors. (Not judging here)
- etatoby 9y agoPrecisely. I would go as far as saying proprietary binaries for Microsoft systems are the single force making Intel processors keep decades of cruft, considering the immense cost that Intel must bear to keep those old instructions (barely) running on modern processors.
- striking 9y agoThose "decades of cruft" are usually emulated in microcode and not actual silicon. Intel knows very few people use the BCD instructions, but it costs them almost nothing to keep them in while running them slightly slower than most operations. Why mess up a stable API when you don't have to?
- dingo_bat 9y agoWhy is it so exciting when Intel has a bug? I had fun reading :)
- agumonkey 9y agoPhysicists tell if you sum Fabrice Bellard and Xavier Leroy, the universe enters an undefined behavior void.
- deleted 9y ago[deleted]
- timeu 9y agoRelated to this: https://tech.ahrefs.com/skylake-bug-a-detective-story-ab1ad2beddcd https://tech.ahrefs.com/skylake-bug-a-detective-story-ab1ad2... Also a pretty good read and recently discussed here: https://news.ycombinator.com/item?id=14661473 https://news.ycombinator.com/item?id=14661473
- hsnewman 9y agoCan this be exploited for malicious code?
- userbinator 9y agoIt's "unpredictable" what happens, so I think the best you're going to do is a DoS. I.e. if you could get the JS JIT in a browser to generate code like this and execute it repeatedly, you could crash a machine just by visiting a site.
- zurn 9y agoIn software systems you can nudge many similar situations to give you control over what badness happens when you drive the system off the rails of the invariants. No reason why things anogous to nop sleds, heap spraying etc would not be applicable here.
- qb45 9y agoIt also is "unpredictable" what happens when you overrun a stack frame. There is zero guarantee that the infamous Sufficiently Sophisticated Attacker couldn't predict it. Hardware is largely deterministic, even when it doesn't behave in the documented way. I wouldn't interpret this as literally unpredictable, it's just a generic slogan they always use in their errata. And they aren't going to say anything more for obvious reasons. Patch this damn microcode.
- hacktothefuture 9y agoI've worked so far away from the metal for such a long time but I still find these types of articles so interesting even though I only understand a small fraction of the info. Its amazing to think the levels of abstraction which are in place from the code at this level which make my work possible.
- libeclipse 9y agoReminds me of this: https://news.ycombinator.com/item?id=14279124 https://news.ycombinator.com/item?id=14279124
- Sean1708 9y agoI'm very upset that we didn't get the story behind #10.
- justin66 9y ago> SIOU's application was single-threaded and made no network I/O, only file I/O, so its execution should have been perfectly deterministic Really?
- piemonkey 9y agoOne amusing thing about this epic tale is Serious Industrial OCaml User disregarded direct, relatively easy to implement, very sound advice from Xavier Leroy about how to debug their system! I would like to think that, were I in a similar situation being advised by an expert of that calibre, I would at least humor his suggestions. Why seek the expert if not for his advice? It brings to mind people disregarding doctors who give them inconvenient medical advice.
- justin66 9y agoI know nothing of OCaml culture or why the author is deemed worthy of having his name italicized, but the doctor comparison is upsetting. If the guy who wrote this: I was tired of this problem, didn't know how to report those things (Intel doesn't have a public issue tracker like the rest of us), and suspected it was a problem with the specific machines at SIOU (e.g. a batch of flaky chips that got put in the wrong speed bin by accident). were a doctor, he'd be guilty of malpractice. This bug went unreported eight months longer than it needed to. Am I misreading all this somehow?
- piemonkey 9y agoHis name is italicized because he is the primary author of OCaml (and a plethora of other great tools, like CompCert, the first fully-verified compiler). Overall, an exceedingly competent and productive programmer and scientist. The doctor metaphor isn't perfect; what I was going for is, when you are seeking out an expert's advice and you ignore it, why do you go to see the expert in the first place?
- balls187 9y ago> That would not be the first time that GCC treats undefined behaviors in the least possibly helpful way, Oh compilers. Like VC++6.0 initializing uninitialized memory to 0xCDCDCDCD in DEBUG.
- tvgggghh 9y agoErr, that's on purpose? It's so you can tell your writes apart, which is helpful while debugging. They used to use, uh, more obvious patterns but the PC brigade called them on it so they settled on 0xcd.
- qb45 9y agoUm, what was it? You can paste in decimal to avoid triggering the PC brigade ;)
- detaro 9y agoWikipedia has a list that includes some used values: https://en.wikipedia.org/wiki/Hexspeak https://en.wikipedia.org/wiki/Hexspeak
- balls187 9y ago0xDEADBEEF how I miss you. Yea, I knew it was on purpose, but it had the unintended consequence of masking bugs that would only show up in release builds.
- insulanus 9y ago0xCD is the interrupt instruction on X86. If you tried to execute out-of-bounds memory, you get a trap right away (handily, the instruction is one byte long). http://www.mathemainzel.info/files/x86asmref.html#int http://www.mathemainzel.info/files/x86asmref.html#int Before Intel processors had execute protection, this was a good way to catch bugs in your buggy C bugs. I mean programs.
- balls187 9y ago
- newusertoday 9y agoI routinely see these bugs when the new hardware is still getting developed.
- civilitty 9y agoThey're completely unavoidable short of having all of the world's entire computing power with which to do formal verification (and even then, there's no guarantee). I've even seen, when developing standard cell libraries for a new fabrication process, bugs that occur because of unforeseen interactions between different semiconductor doping concentrations that occur when (due to pure statistics in fabrication) they overlap in the wrong way.
- Cellestro 9y agoI am even more convinced to learn OCaml, seeing the passion the creators have to solve the problems.
- SiempreZeus 9y agoI love reading about hardware bugs, and people their perseverance! Reminded me of a developer for Crash Bandicoot who had seemingly random crashes: http://www.gamasutra.com/blogs/DaveBaggett/20131031/203788/My_Hardest_Bug_Ever.php http://www.gamasutra.com/blogs/DaveBaggett/20131031/203788/M...