7 ms·
Solving Meltdown absolutely does not require a new type of processor. Performing out-of-order executions of privileged instructions is something you can just no
by furi 8y ago
Solving Meltdown absolutely does not require a new type of processor. Performing out-of-order executions of privileged instructions is something you can just not do, nobody except Intel is doing it right now in fact. Lots of Spectre variants likewise don't require entirely new processor varieties, such as the lazy floating point speculation vulnerability or the new L1 terminal fault vulnerabilities. The only one that possibly requires a new type of processor is the original Spectre, which is quite hard to completely stamp out all possibility of while retaining the concept of speculative execution.
- xenadu02 8y agoIndeed, AMD CPUs don’t speculate loads until the page protection is checked. That alone eliminates whole classes of vulnerabilities. Cache updates will need to be staged and not “retired” (aka committed and made visible to other cores) until the speculated instructions are retired. That will have some perf impact and cost some die space, but it is hardly fatal. Whether AVX survives in its current form is an open question. Those 512-bit vector units cost a lot of power. Intel gets away with powering them down when unused, leading to a large delay when the first vector instruction is encountered after an idle period. Powering them up continuously blows their TurboBoost strategy. If they can eliminate the penalty or make the units more power efficient it might not require many changes, otherwise AVX may need to shrink back down to narrower units, or maybe force them through a shutdown cycle on every context switch? Not sure what the solution is here.
- voidmain 8y agoOn the AVX thing, I think maybe you can eliminate it as a channel for speculative execution attacks by just not speculatively executing AVX-512 instructions when the units are powered down, which also sounds more efficient (it doesn't sound good that, apparently, if (expression_that_is_false) { do_avx_instruction() } can drop your clock speed for several milliseconds if branch prediction guesses wrong! The cost of powering up the AVX units is, I think, much greater than the cost of failing to speculate once.) In general, I think the problem is that we probably don't know about all possible side channels, and might not for many years. So the approach you suggest - eliminating side channels one by one so that you can't extract information from speculative execution that way - is inherently risky.
- phire 8y agoI don't know if Intel have done this, but you can design AVX-512 to run on AVX ALUs at half the IPC. You could feasabley design a CPU which has full AVX-512 units powered down and runs instructions at half speed until the full ALUs are powered up. One you have a CPU of that design, you can eliminate that AVX-512 sidechannel by not sending the signal to power up the full AVX-512 ALUs until a AVX-512 instruction is fully executed.
- mjevans 8y agoI'm not sure you could do that in a completely secure processor though; at least not the way you're describing. More ideally the delay would at least be /simulated/ even IF the units were already powered up because of prior instructions. It would be /per thread/ tracking of slow/fast/powered before paths. Edit: About 10 min after posting the above, I realize that this MIGHT be what your second paragraph is describing, though it isn't as unambiguous.
- sitkack 8y agoWould there still be a thermal side channel, if one had access to high enough resolution power monitoring?
- phire 8y agoNo. Simply providing two ways to executate with 256bit or 512bit ALUs is not enough to prevent side channels, you can still time the execution time to read the side channel. The key to closing the side channel is not allowing speculated instructions to trigger the activation of the 512bit ALUs. Hold off for a few cycles until execution of those instructions is confirmed.
- titzer 8y ago> ...by just not speculatively executing AVX-512 instructions... See this suggestion a lot. At any given time, a CPU is trying to execute 5-8 u-ops in different execution units every cycle. These 5-8 u-ops must come from instructions somewhere. To have something to do, CPUs need to load dozens, even hundreds of instructions from "the future" using branch prediction. As such, there literally could be hundreds of instructions in its reorder buffer at once. Of these, 5-10 of them might represent branches which cannot been executed due to dependencies but instead have been simply guessed at using branch prediction. TLDR; that pretty much means that the CPU is always speculating--perhaps 95% of all cycles. Typical CPU designs do not explicitly track branch dependencies in the reorder buffer. Instead, they either rely on anulling at commit time or clearing the reorder buffer when a mispredicted branch commits. That means the CPU literally has no way of knowing whether it is currently "speculating". To fix this, one would have to add control dependencies to AVX instructions so that they could never execute with branch instructions on which they depend (i.e. earlier on their proper architected path) in-flight. That would almost certainly annihilate performance, which, after all, is the point of AVX instructions.
- vmchale 8y ago> which is quite hard to completely stamp out all possibility of while retaining the concept of speculative execution So... why not do exactly that? Let compilers access cache and stop this nonsense where we pretend computer memory is flat in order to let C developers believe they're "close to the metal [sic]"
- gpderetta 8y agoThat's because compilers by are terrible at predicting the dynamic behaviour of a program, be it cache access patterns, branche prediction etc.
- lvs 8y agoThat's just precisely what the article says. The first version of Spectre is the problem.
- rurban 8y agoBut it didn't tell you that such processors do exist, you just need a 2nd c3 register. It's not rocket math.
- nickpsecurity 8y agoSince gruez mentioned it, I'll note Intel also pushed Itanium which had security benefits almost nobody talks about in x86 vs Itanium discussions. Secure64, co-founded by Itanium designer, uses them in SourceT OS which they claimed got positive analysis by Matasano Security. https://www.intel.com/content/dam/www/public/us/en/documents/white-papers/intel-itanium-secure-white-paper.pdf https://www.intel.com/content/dam/www/public/us/en/documents... They also claimed the Itanium-based solutions were immune to Spectre and Meltdown. I'm not a CPU expert. I'll let others review that. A lot of attacks are showing up. https://secure64.com/not-vulnerable-intel-itanium-secure64-sourcet/ https://secure64.com/not-vulnerable-intel-itanium-secure64-s... So, the processor that was about stronger reliability and security that the market and mainstream security ignored seems to mitigate some of the risks both are now griping about. Maybe folks who don't depend specifically on x86 might buy some Itaniums to signal they'll pay for security-enhanced processors. Be sure to tell the sales rep why so they can pass it up the chain. :) For embedded stuff, there's also Microsemi's CodeSEAL and Dover's CoreGuard which have advanced protections. The first has encrypted/authenticated RAM plus control-flow integrity at CPU level. The second has a flexible, metadata unit that enforces many types of security policies at CPU level on per-instruction basis.
- close04 8y agoItanium is dead, after 15 years on the death bed. To my knowledge there is no plan to launch any new Itanium CPU. The only reason the 9700 was launched in 2017 is to serve as a drop-in replacement for old CPUs, most likely for contractual obligations. There's simply no "killer feature" for Itanium. It's not coming back.
- nickpsecurity 8y agoThat serves my point that introducing a security-enhanced processor didn't work for Intel. i432 APX and i960 also failed at a loss of over a billion dollars so far when Intel tries to make better, safer CPU's. The customers wanted backward compatible with x86's warts, highest performance per dollar, and decreased energy. That's what Intel gave them at billions in profit. They're making patches for the security weaknesses and/or fixing them in newer CPU's. Vendors of insecure processors still dominate while the few processors with better security sell almost nothing. So, Intel should continue to make deliberately insecure processors and patch them since that's what people pay for. It's the market's fault for rarely buying anything better. Plus, voters not doing something about patent reform which would've led to more x86 competitors. The Chinese supplier is an interesting development where we might get a secure-ish x86 that way. "There's simply no "killer feature" for Itanium. It's not coming back." I just named its killer feature: security enhancements that make it hard to inject code into it. Most security audits of OS's find problems. The one of SourceT on Itanium didn't find a way to inject code into it since it used hardware to mitigate and contain most of that. Whereas, the protections on x86 are a source of vulnerabilities at the moment. Unfortunately, that was the only company that I'm aware of which saw that potential. The market's direction means they're porting the product to insecure hardware right now. Who knows what they'll come up with since it's impossible to secure software on insecure hardware. I'm so grateful to the markets for buying such things so little that they basically don't exist in the mainstream space. Note: Another example was Cell processor. Green Hills, who make INTEGRITY-178B OS for high security, did an assessment of the Cell processor for baking security into systems. It had some nice capabilities for that. Market wasn't going to buy it even to protect their systems, though. So, they're not pushing that.
- ohiovr 8y agoWhy can't Intel sell drop in replacements for the existing cpus out there? Can of worms?
- wolfgke 8y ago> Why can't Intel sell drop in replacements for the existing cpus out there? Can of worms? Even if this were financially suitable for Intel (it is not), a necessity for this is that Intel has a microarchitecture available that "really fixes these problems" - but there is none. I consider it as plausible that the old in-order Atoms are less prone to these problems - but do you really want such a much slower CPU?
- zaarn 8y agoOf course, simply solving Meltdown can be done on an existing CPU. However, VLIW-like architectures could give us back the performance of out-of-order execution without the drawbacks since the compiler has much more direct control over the caches and registers used in the CPU as well as branch prediction and lots of other internals. Of course, you also need much smarter compilers, since the Itanium failure, we've come a long way and I'm convinced that Rust and friends would be able to get the maximum out of a VLIW-like with some additional work.
- wolfgke 8y ago> However, VLIW-like architectures could give us back the performance of out-of-order execution without the drawbacks since the compiler has much more direct control over the caches and registers used in the CPU as well as branch prediction and lots of other internals. > Of course, you also need much smarter compilers, since the Itanium failure, we've come a long way and I'm convinced that Rust and friends would be able to get the maximum out of a VLIW-like with some additional work. - What kind of compiler magic do you have upon your sleeves that will enable that kind of magic, but circumvents the problem of non-existence of a "sufficiently smart compiler" that plagued Itanium? In particular: What kind of magic does Rust offer for this? - How do you intend to solve the problem (that also plagued Itanium) that mostly scientific code has the kind of parallelity that is very suitable to VLIW which lead to the problem that "lots of ordinary, existing code" did not benefit so much from the potential speed that Itanium offered? - How do you intend to solve the problem that putting deep microarchitecture details into the instruction set is usually a bad idea, because it "cements" these details (I just say: MIPS' delay slots), while having a good instruction set makes deep changes in the underlying microarchitecture easy.
- zaarn 8y agoCompilers have advanced a lot since Itanium as have languages, Rust has lot of information to properly track shortly lived objects and optimize that. Lots of code runs fairly well in CPU pipelines with multiple work units, if the work pipeline is too wide that sucks of course and is a bit of a waste but, say, having a Intel-sized pipeline would run most of the current code with the same efficiency. This isn't purely about parallelizing threads, this is about microcode instructions which can be parallelized very easily; pull a instruction from the list of the program and put it in the first free execution unit on no conflict, if a conflict occurs resolve it (WaW; drop it or assign new register, RaR; free reordering, WaR; assign new register, RaW, strictly no ordering before), repeat. A compiler should be easily able to do that. Of course having the microarchitecture and instruction set tightly married is a problem but considering we have a fairly alive ARM marketplace of varying instruction sets from ARM, I think it's not unsolvable from the compilation direction and you can abstract out some details so porting is easier or you can atleast emulate the other CPU with high efficiency.