3 ms·
I think that a lot of discussion here is missing a key design aspect of modern speculative cores: the "late check" of permissions was likely not a sloppy shortc
by throwaway203937 9y ago
I think that a lot of discussion here is missing a key design aspect of modern speculative cores: the "late check" of permissions was likely not a sloppy shortcut or oversight, but rather due to a common design principle, or separation-of-concerns, in modern cores: exception-condition handling and speculation flush/restart are ordinarily handled when the excepting instruction reaches retirement, because at that stage, instructions are back in order again, and they are "real" (could not be flushed by any earlier mis-speculation). It seems completely natural w.r.t. modern design principles to handle page faults at retire.
In other words, as obvious as the vulnerability is in hindsight, the general design is "textbook computer
architecture" -- at least it would not raise any eyebrows. In fact the alternative, eager checking, would be the less obvious design, because almost every other exception is handled at retirement, and it would require careful thought and some additional complexity: do you cancel the instruction? (Potential problem: dependent instructions are already scheduled and will expect your result next cycle, because scheduling is deeply pipelined. Now you need some auxiliary scheduling information to let you cancel those too, which will be very complex and ripe for bugs.) Do you zero out the result? (Potential timing issue: this requires a few more gate delays between the data-cache word-select MUX and the results bus, a pipeline stage that's probably already timing-critical.) Clearly, AMD and ARM made other design decisions, but I would bet this is due to incidental aspects of their pipelines rather than an explicit prediction of a Meltdown-like vulnerability.
Overall: the issue looks very obvious in hindsight, but it arises due to an unexpected interaction of common design principles of modern cores, rather than a negligent oversight or shortcut.
(I worked on processor design for a while; throwaway for obvious reasons)
- wilun 9y agoIt's not about generating the exception. It's about not operating on privileged data. (And not loading it to registers seems a good way to not operate on them, btw)... Generating the exception can still obviously happen at retire. It's not even a problem to run an extra bunch of code with constant placeholder garbage data (if that would be too costly to interrupt it sooner), given you know all is gonna be cancelled by the eventual exception anyway. But you just don't fucking load the privileged data to begin with!!! Timing side channel attacks have been known for maybe a decade for heaven's sake. Even if Meltdown would not have been as simple as it is, as soon as you speculate on privileged data you are completely dead; because there will be tons of obvious side channels to retrieve said data. (measuring occupation of execution units through HT comes to mind, actually before meltdown was disclosed I thought this was the trick to leak the data that should not have been loaded, turns out it was even simpler...) Maybe the field needs better "textbooks", or to return to the classics about speculative execution. I've yet to see anybody blaming Intel for Spectre. Some other designers might have made the same mistake as Intel for Meltdown, but that does not make it less a mistake. Maybe AMD only had it correctly by some kind of strange accident, and it will probably be very hard to know. That would still be a silly mistake for Intel, even then.
- throwaway203937 9y ago> It's not about generating the exception. It's about not operating on privileged data. Yup. I think the subtle distinction I'm trying to get across is this: in the mind of a computer architect, up until Meltdown, there was a powerful and useful simplifying principle available: speculation unwinding (due to e.g. an exception) will clear away the results of any instructions after the excepting point, so it doesn't matter what we do on the "wrong path" (the instructions that will be cleared). If you set a bit in the ROB entry for an instruction that will trigger a page fault at retire, you know that it doesn't matter what data is returned, because the load will never commit to architectural state; it will be flushed. You can design the logic as "don't care" at that point. I'm not saying that this is the correct way of thinking now, in a post-Meltdown world. I'm simply saying that the blind spot can be understood from the point of view of that principle (which seemed reasonable to many people at the time). To the layman, "allow operation on privileged data" sounds careless and negligent. To a core architect, it's a (seemingly) correct design. Speculation unwind will reset your state anyway, so one might as well omit the (non-free) eager checking logic. It's a simpler, cleaner design, easier to verify, etc. The blind spot was that side-channels make this a leaky abstraction, and we do have to care about what happens during speculation that will be cleared. That is extremely non-intuitive to most computer architects (or at least, to me, and I did research on microarchitecture in academia then at a large chip company). If you like, just interpret my post as a report of the widespread mentality in the industry -- I'm not saying it's right, just that this is how it likely came about.