8 ms·
Speculating the entire x86-64 instruction set in seconds with one weird trick
- TrainedMonkey 6y agoThis just cinched to me that we need to sunset hardware x86 and run code that cannot be recompiled on emulators. x86 had a good run, but it becoming increasingly obvious that maintaining backward compatibility in modern high performance parts is incredibly expensive and bug ridden. At this point sunsetting x86 is not even a pipe dream, most people carry ARM powered computers in their pockets and Apple recently demonstrated that it can be quite successful in high performance devices.
- mhh__ 6y ago> but it becoming increasingly obvious that maintaining backward compatibility in modern high performance parts is incredibly expensive and bug ridden Moving everything to ARM will not be cheap. As for bugs that is entirely dependent on the company making the chip, which ones do you have in mind? (Also recall that M1 is vulnerable to Spectre too). I kind of hope X86s days are numbered as well, but I'm not looking forward to an ARM monoculture.
- deleted 6y ago[deleted]
- als0 6y ago> Moving everything to ARM will not be cheap I thought the same thing until I saw how well Apple's Rosetta 2 works. Now that we've seen what's possible, I'm hoping that the x64 emulation in Windows/ARM will rise to the challenge set by Apple.
- mhh__ 6y agoAnd to be able to use Rosetta 2 I only have to spend the combined value of all my computers again? Double that if I want the same amount of RAM and storage as I have now
- als0 6y agoYou don't have to use Rosetta 2 - that was an example of a good implementation. I did mention the announced x64 emulation on Windows, but that's only a preview release at the moment. If you're a Linux user the only thing I'm aware of is QEMU TCG but there might be faster projects out there.
- mhh__ 6y agoMy point was that it doesn't suggest cheap any time soon.
- monocasa 6y agoI imagine Apple has a patent on the cute optional TSO memory model that makes it work.
- duskwuff 6y agoI don't think they do. The memory model itself isn't patentable; the TSO memory model already existed on x86, as well as many other architectures before that. Having an option to enable/disable TSO might be patentable, but it'd be a stretch; there's plenty of precedent for allowing a CPU to select between operating modes at runtime. (For example, many PowerPC parts could be switched between little-endian and big-endian modes at runtime, and some developers even used this to assist in emulating x86!) What's more likely is simply that other ARM SoC vendors aren't implementing their own cores (e.g. they're using a standard Cortex-Ax design from ARM), so they can't add deeply integrated features like TSO themselves.
- monocasa 6y agoI mean, selectable, cheap, memory model changes are novel and I pretty much guarantee they have patents on some aspects of it. Depending on the patent there might be some cute ways to work around it, but it's going to take some work on your part. They might even have patents on other attempts to get to the same effect that they might have tried out. PowerPC's biendianness would be patentable too if it had been an absolutely ancient technique about as old as computers themselves. And there's people other than ARM and Apple making ARM cores. Samsung's Exynos M5 should be coming out before too long for instance.
- monocasa 6y agoWhile some of the specifics are obviously x86 specific, I'm not sure that the underlying issues here are. From undocumented instructions, to lack of documentation about speculation barriers most of the root issues you see here are equally applicable to the cores in an M1.
- als0 6y ago> to lack of documentation about speculation barriers Can you elaborate on this? I was having a look at the M1 instructions here [1] and it seems that they implement at least one of the barriers, CSDB, which is actually documented by ARM[2]. [1] https://dougallj.github.io/applecpu/firestorm-int.html https://dougallj.github.io/applecpu/firestorm-int.html [2] https://developer.arm.com/documentation/ddi0596/2020-12/Base-Instructions/CSDB--Consumption-of-Speculative-Data-Barrier- https://developer.arm.com/documentation/ddi0596/2020-12/Base...
- monocasa 6y agoThere's explicit speculation barriers that they document, but they don't tell you what _other_ instructions are speculation barriers simply due to microarchitectural compromises like what you're seeing in this article.
- als0 6y agoOK thanks for clarifying. If there are explicit instructions for blocking speculation then why are you concerned about implicit barriers?
- retrac 6y agoVME support was broken on Ryzen for a while before a microcode patch came out. The VME instructions are a relatively obscure, and originally Intel proprietary extension, to i386's virtual 8086 mode. Introduced in the 90s to speed up DOS virtual machines under OS/2 and NT, I think. I don't know much about how stuff is implemented these days, but virtual 8086 mode involves some page table muckery and similar. Surely implementing it creates a larger exposure area, from a security standpoint.
- muricula 6y agoMany years ago I worked with a greybeard AMD hardware designer. He told me that they commissioned a study about whether it made sense to ditch backwards compatibility, and realized that the parts of the CPU needed to support backwards compatibility they were willing to ditch contributed to less than 1% of die area, and of course were all already designed and battle tested. Unfortunately I don't have a better source for this than an anecdote.
- mhh__ 6y agoThat probably depends on whether you mean old instructions now cause a fault or moving x86 to a new format entirely that doesn't need a decoder from hell
- jeffbee 6y agoYou could lose the entire x86-specific part of the die in a corner of a 512x512 FMAC unit. Die cost due to x86 complexity seemed like something worth attacking in 1990 when RISC was gaining mindshare but over the years the complexity of that has not expanded very much while the other stuff on the die has gotten much larger.
- sesuximo 6y agoA lot of code doesn’t really care about speculative execution. It would be a shame to throw out decades of development and thousands of cpus just for one use case that didn’t totally fit it.
- matthewmacleod 6y agoThis is such a weird meme to me. In what way is the thing you described at all "increasingly obvious"? We see this kind of statement all the time, but almost never accompanied by any sort of actual rationale for it.
- woodruffw 6y agoThis is a really clever technique! I was impressed by sandsifter[1] when it originally came out, and this seems an awful lot faster and less prone to false negatives (since it's purely speculative and doesn't require sandsifter's `#PF` hack). At the risk of unwarranted self-promotion: the other side of this equation is fidelity in software instruction set decoders. x86's massive size and layers of historical complexity make it among the most difficult instruction formats to accurately decode; I've spent a good part of the last two years working on a fuzzer that's discovered thousands of bugs in various popular x86 decoders[2][3]. [1]: https://github.com/xoreaxeaxeax/sandsifter https://github.com/xoreaxeaxeax/sandsifter [2]: https://github.com/trailofbits/mishegos https://github.com/trailofbits/mishegos [3]: https://ww.easychair.org/publications/preprint_download/1LHr https://ww.easychair.org/publications/preprint_download/1LHr
- deleted 6y ago[deleted]
- blight 6y agobe sure to check out the data set extracted by this research over at https://haruspex.can.ac/ https://haruspex.can.ac/
- alpb 6y agoOff-topic: This was posted yesterday, but got no attention. (I tried to re-post yesterday, but got redirected to the existing post.) https://news.ycombinator.com/item?id=26576032 https://news.ycombinator.com/item?id=26576032 I wonder what makes HN disallow reposting of the same URL in a short periods of time but allowing it in long-term.
- tomcam 6y agoI think the first thing you described is called spamming
- cbhl 6y agoIf I recall correctly, this heuristic was found experimentally: Once upon a time, reposts were allowed, and so if something was topical and popular, sometimes you'd see multiple items on the front page pointing to the same URL. Sometimes the second or third posting would get more karma simply because of an editorialized title. This was undesirable, so they added the repost merging. Later, an emergent behavior was that posts that report on something found on another site and had been previously posted would find their way to the front page. Some subset of HN visitors would be newer than the original post, and simply upvote something they hadn't seen before. But also there would be times where the post would be interesting with the added value of hindsight, or would provide context to the topic-of-the-day. So at some point it was decided that reposts would be allowed in the long-term, since sometimes they had value. Thus we're at the state we have today.
- dan-robertson 6y agoSometimes the moderators encourage reposts of unnoticed articles they thought were good (by reaching out to the poster directly). Although I don’t know if that’s what happened here. It could have been a sufficiently different url.
- Iv 6y agotl;dr: he used a counter provided by intel that describes the total number of microcode instructions translated. He tried thousands of possible opcodes in "speculation mode" (this is the mode CPUs use to calculate both forks of a branch while waiting for the branch to be decided) and checked when an anomalous number of microcode instructions were translated. He found 13 likely candidates for unpublished ops, including the 2 that were recently found. Also a few unpublished quirks of some known instructions.
- tux3 6y agoA technical nitpick: speculation doesn't check both forks of a branch, it has to pick a side! The CPU tries very hard to guess which way branches go and that allows it to speculate much further than if it tried every possible combinations. What happens in this post is that the author writes a CALL instruction, but then manipulates the stack so that it doesn't actually return where the CPU expects it to. So the CPU will speculatively execute the instructions that follow the CALL linearly, even though they are never actually reached!
- Iv 6y agoOh I did not know that about speculation!
- alfiedotwtf 6y agoGood article! Now someone needs to do a follow-up to see what software uses these hidden opcodes :)