18 ms·
New x86 micro-op vulnerability breaks all known Spectre defenses
- tester756 5y ago>"Intel's suggested defense against Spectre, which is called LFENCE, places sensitive code in a waiting area until the security checks are executed, and only then is the sensitive code allowed to execute," Venkat said. "But it turns out the walls of this waiting area have ears, which our attack exploits. We show how an attacker can smuggle secrets through the micro-op cache by using it as a covert channel." >"In the case of the previous Spectre attacks, developers have come up with a relatively easy way to prevent any sort of attack without a major performance penalty" for computing, Moody said. "The difference with this attack is you take a much greater performance penalty than those previous attacks." >"Patches that disable the micro-op cache or halt speculative execution on legacy hardware would effectively roll back critical performance innovations in most modern Intel and AMD processors, and this just isn't feasible," Ren, the lead student author, said.
- Randor 5y agoThe best part of the new "defense against Spectre" is that the LFENCE instruction has been around for ~20 years. It's not even not a defense against all variants.
- the8472 5y agolfence behavior varies. On AMD CPUs you need to set an MSR to make it serialize instruction dispatch.
- mhh__ 5y agoSo what? The defense relies on it serializing the instruction stream which is not necessarily true based on the semantics of the instruction (until it was retroactively documented to do so)
- londons_explore 5y agoThe solution will be "do not share the micro-op cache between different address spaces". Which for old hardware will translate to "flush the micro op cache every time the address space changes". I would guess that can be done with a microcode update and that the performance hit wont be too massive.
- deleted 5y ago[deleted]
- tachyonbeam 5y agoThe micro-op cache is very small, on the order of ~1.5K uops AFAIK. It can also be repopulated quite fast. So yes, the performance hit should be quite small. You should presumably also be able to reduce the performance hit if you reduce the frequency of context switches, which should get easier the more cores you have, if I'm not mistaken. That is, the OS can have its own dedicated core, and some programs can be more or less pinned to other cores where they are rarely interrupted.
- the8472 5y ago> You should presumably also be able to reduce the performance hit if you reduce the frequency of context switches, which should get easier the more cores you have, if I'm not mistaken. Context switches don't happen that often due to preemption unless your CPU is oversubscribed. Most context switches are due to syscalls, especially the ones used to wait for contended locks. Reducing those takes a lot more optimization work.
- sitkack 5y agoGiven the hockey stick number of cores coming at us, I see pinning and better temporal avoidance being solutions. High security code will be pinned to its own core, running in its own memory area. So much more scheduler work to do.
- deleted 5y ago[deleted]
- totallyabstract 5y agoThere are separate micro op caches per core however they are typically shared among hyperthreads. I wonder if this could be another good reason for cloud vendors to move away from 1vCPU = 1 hyperthread to 1vCPU = 1 core for x86 when sharing machines (not that there weren't enough good reasons already).
- tyingq 5y agoOr to roll out more ARM, where there isn't currently any hyperthreading.
- ljhsiung 5y agoEven putting aside security aspects aside, in general I've been seeing research pop up over the years criticizing SMT's performance claims of ~30%. Hell, even Amazon's Graviton CPUs don't have it (though I'm sure that's a product of being ARM derived rather than a design decision).
- tux3 5y agoARM vendors must be feeling pretty good about themselves yeah, but if you take AMD's cores... SMT might not be a huge win in every benchmark, but you just can't keep that wide backend fed from a single hyperthread (at least I can't!). So turning SMT off is at the least wasted potential for those cores, the way they've been designed
- spacemanmatt 5y agoIs ARM so much better? I can migrate my AWS hosts.
- tyingq 5y agoSeparate micro-op cache per core, and no hyperthreading, so ARM would seem better equipped to defend against this.
- spacemanmatt 5y agoGood to know. This new era of aggressive hardware flaw exploitation has me motivated to leverage my mobility and flexibility to evade. I don't think I have a better strategy.
- grishka 5y agoSo ARM CPUs do have microcode after all?
- mhh__ 5y agoMicrocode != Using micro-ops. Microcode has been around for half a century at least, much longer than micro-ops.
- grishka 5y agoUh. I thought that micro-ops were literally microcode instructions/operations?
- mhh__ 5y agoIn a modern processor they are however microcode as a general term is a catch-all term which basically means non-trivial configuration logic stored in somewhere not meant to be touched by people other than the vendor. I think IBM have millicode.
- 5y ago
- floatingatoll 5y agoI’d like to highlight this excellent post about x86 micro-ops “fusion” from three years ago, as it’s the reason I have any idea at all what micro-ops are: https://news.ycombinator.com/item?id=16304415 https://news.ycombinator.com/item?id=16304415
- PopePompus 5y agoI don't understand this at all; I didn't think the mico-op cache was visible to code written for the x86 ISA at all. Can anyone explain to an idiot (me) how something in micro-op cache can become visible to the outside world?
- tux3 5y agoI'm simplifying a bit (edit: quite a bit =]), but the way these attacks work is generally by exploiting the difference in timing between something being in cache, and something not being in cache. Or some resource being contended vs not contended. If something is in cache, and you also have access to that cache, accessing that thing will be fast and few CPU resources will be used. So you can tell that something is in cache. And you know you didn't put it there. So some other thread that you're sharing a CPU core with must have put it there. To exploit those attacks, you're going to intentionally watch the other thread as it, for example, (speculatively) takes a branch, and either puts something in cache or doesn't. Now you know whether the other thread (speculatively) took a branch or not! Just from measuring timings of the cache. From that, you work back to what the branch condition (that was still only speculatively executed) must have been, and if this branch is based on (speculatively loaded) data, you just leaked one or more bits of the data. Suddenly, things are not speculative anymore. You guessed data that wasn't yours, because speculatively using it had an effect on the cache, and you could measure that effect. Here, they use the micro-op cache (I haven't read the paper, so I don't know the details, but this is broad strokes). Any mechanism that you can use during speculation, and that you can extract timing information from is potentially a problem. And these are everywhere. That's why the Spectre problem is so hard to fix now that pandora's box is open.
- fnord77 5y agoso something say, sandboxed (like in a browser running webassembly) could get at non-sandboxed data? Or something in one VM getting at data from a different VM?
- DoomHotel 5y agoThere are examples of straight JavaScript exploits that allow a website to read memory from anywhere in the process its code is running in. https://cacm.acm.org/magazines/2020/7/245682-spectre-attacks/fulltext https://cacm.acm.org/magazines/2020/7/245682-spectre-attacks...
- darig 5y agoIt doesn't break the defense I used: Stop buying Intel x86 chips.
- iam-TJ 5y agoThe U of V Engineering Faculty release is at https://engineering.virginia.edu/news/2021/04/defenseless https://engineering.virginia.edu/news/2021/04/defenseless
- afturkrull 5y ago"In 2018, industry and academic researchers revealed a potentially devastating hardware flaw that made computers and other devices worldwide vulnerable to attack" Does this apply to all CPU archicetures? The article is a little vague. A fall-out from the monoculture. Hey, where did my post go?
- dataflow 5y agoQuestion: How relevant are these for the average person? I know these matter for things like shared hosting, but I've yet to hear of an actual exploit in the wild that ordinary people have been attacked by, even with Spectre defenses turned off. Should normal people be worried about this?
- 1e-9 5y agoYes. It could undermine your browser if you allow a malicious site to run JavaScript.
- dataflow 5y agoSpectre could too, but again, my point was that I didn't hear of actual attacks on people in the wild, at least not on any scale that seemed to make the news. Is there a reason to believe this will be different?
- 1e-9 5y agoIt can take months or even years for proof-of-concepts to become widespread in the wild, particularly by those sloppy enough to be easily detected.
- panny 5y ago>I didn't hear of actual attacks on people in the wild You never would. It's a passive attack. It's measuring response time to normal operations to discover secrets. https://mlq.me/download/netspectre.pdf https://mlq.me/download/netspectre.pdf "Software based side-channel attacks are particularly unsettling since they do not require physical access to the device."
- baybal2 5y agoThe first known Spectre-like concept was actually traced to Pentium 3 times in nineties. It took 2 decades for everybody to forget about it before the vulnerability dismissed as "not exploitable in the practice" came back with a vengeance.
- akersten 5y agoI've been saying this from the start: the well of issues is infinitely deep as soon as you decide that multiple tenants running on the same physical hardware inferring something about another is a vulnerability. I assert, but cannot rigorously prove, that it is not possible to design a CPU such that execution of arbitrary instructions has no observable side-effects, especially if the CPU is speculating. I don't know what that spells for cloud hosting providers - maybe they have to buy a lot more CPUs so every client can have their own, or commission a special "shared" SKU of CPU that doesn't have any speculative execution - but I know for me, if I have untrusted code running on my CPU, I've already lost. I could then care less about information leakage between threads. We're going to wind up undoing the last 20 years of performance gains in the name of 'security', and it scares me.
- MarkSweep 5y ago> if I have untrusted code running on my CPU, I've already lost Don’t forget about JavaScript, a common way for people to run untrusted code on their computers. Not all of micro-architectural data sample are exploitable in JavaScript, but some are.
- baybal2 5y agoChrome had 7 exploits caught in the wild within 7 weeks in 2020. I believe it is going towards JIT being disabled, or most severely limited.
- userbinator 5y agoIt sounds like a dream, but going back towards interpreted JS instead of JIT may finally stem the insanity of bloat that JS has evolved in an environment of increasingly fast implementations.
- otabdeveloper4 5y agoThe problem of Javascript bloat doesn't have a technical solution. Javascript bloat exists because of a social problem: the guy who fixes the corporate webpage's javascripts is called a "webdesigner", and "webdesigners" are the lowest rung on the corporate IT ladder, maybe only a bit above first-tier techsupport. If you want to make some sort of career you need to upgrade from "webdesigner" to "frontend developer", and that means cryptic, incomprehensible and pointless "frontend frameworks". It provides to value to business or users, but management puts up with it because it fixes the problem of employee churn. (Frontend positions are a big pain in the ass.)
- baybal2 5y agoYou cannot realistically make a CPU invulnerable to performance analysis And you don't need to. There is really very few uses for real multi-system vs multi-process shared systems. Take a look on that whole "cloud" thing. All people I knew who worked in cloud hosting tell that most system are ridiculously overprovisioned, effectively nullifying any economic justification for a shared system
- londons_explore 5y agoOne day, when margins shrink for cloud compute, we'll see less and less overprovisioning...
- lanstin 5y agoI usually end up over provisioning because I need something that is billed along with CPU; for example I have super good C or Go code to run proxies on, they use like 2% of the CPU when they max out the network connection. I add more so the bandwidth goes up.
- Uehreka 5y agoThis. I run into this all the time with WebRTC infrastructure. My SFUs run out of bandwidth long before they’re at 100% CPU. It’d be great if I could easily provision VMs based on bandwidth, but of course cloud providers are always real coy and say things like “this VM size class has Medium bandwidth, but this one has 25Gbps, no we won’t say which of those is bigger.”
- pm90 5y agoIts possible that there are technical reasons related to virtual networks that may be restricting what kind of configurations are possible on their infrastructure. I would expect them to disclose it as such, but cloud providers haven't been very open about sharing those details.
- anthk 5y agohttps://www.mail-archive.com/source-changes@openbsd.org/msg99141.html https://www.mail-archive.com/source-changes@openbsd.org/msg9... OpenBSD disabled HT by default.
- dTal 5y agoThat's less a case of "OpenBSD is prescient" and more "OpenBSD disables everything by default". Even a stopped clock...
- kjjjjjjjjjjjjjj 5y agoHere come more performance gimps. I bet intel is running all of their benchmarks with every single Spectre patch disabled.
- 1cvmask 5y agoThis quote from the article explains the danger quite well: "Intel's suggested defense against Spectre, which is called LFENCE, places sensitive code in a waiting area until the security checks are executed, and only then is the sensitive code allowed to execute," Venkat said. "But it turns out the walls of this waiting area have ears, which our attack exploits. We show how an attacker can smuggle secrets through the micro-op cache by using it as a covert channel."
- smasher164 5y agoMaybe EPIC [1] architectures need a revival. Rely on compilers to take advantage of explicit instruction-level parallelism, and keep the CPU dumb. [1] https://en.wikipedia.org/wiki/Explicitly_parallel_instruction_computing https://en.wikipedia.org/wiki/Explicitly_parallel_instructio...
- bonzini 5y agoThat failed for good reasons. Itanium processors ended up using out of order execution and speculation just like everyone else, because the compilers just don't have enough information compared to an out of order execution engine.
- Paianni 5y agoFrom Poulson onwards anyway. https://www.realworldtech.com/poulson/ https://www.realworldtech.com/poulson/
- pabs3 5y agoReminds me of the Mill ISA: https://millcomputing.com/ https://millcomputing.com/
- zokula 5y ago> Rely on compilers to take advantage of explicit instruction-level parallelism, and keep the CPU dumb. This very much is never going to be feasible for consumer and general purpose computing.
- smasher164 5y agoThe ML folks are pulling themselves out of that rut now. There’s lots of interesting work going on for the next generation of compilers.
- CalChris 5y agoThe paper: I See Dead µops: Leaking Secrets via Intel/AMD Micro-Op Caches http://www.cs.virginia.edu/venkat/papers/isca2021a.pdf http://www.cs.virginia.edu/venkat/papers/isca2021a.pdf
- Woodi 5y agoSimplest way around all of this is back to one-core MULTI-SOCKET systems for "civilian" computers like x86 is.
- Causality1 5y agoI expect this to be just like Spectre. The media sizes it as a tool to use fear to drive engagement, vendors partially cripple their hardware to guard against it, and literally nobody ever bothers trying to actually use it against innocent people.
- mhh__ 5y agoJust like y2k!
- deleted 5y ago[deleted]
- failwhaleshark 5y agoThe act of loading code into memory, be it a hypervisor or a guest OS, should've been gated by sanitation and validation callbacks. Building all of these macro- and micro-op runtime defenses and mitigations in the processor and slowing down the OSes for every possible runtime edge-case are a waste of speed that can be avoided by establishing trust of code pages. The morphing of data into code pages with JITs like JS should also be subject to similar restrictions.
- druud62 5y agoThe CPU needs to make the overheard signals look just like random noise. A cheap XOR-stream (compare 2FA like Google Authenticator, or the remote in your car keys) should cover that.
- vletal 5y agoWell, some of these attacks exploit the actual values present in the memory, not their stored representations. Therefore it would not matter how you encode them on the way, right?
- ineedasername 5y agoundocumented features in Intel and AMD processors Why is this at all a thing? Why would you ever leave something out there like that without documenting its existence?
- gravypod 5y agoI'm assuming these are instructions for self tests or verification. If so, removing the instructions after they are manufactured wouldn't be easy. You can do it in microcode at the cost of making all execution slightly slower (if instruction not in [a, b, c, d]) or by physically altering the die to remove those instructions. Either way, it doesn't sound fun. It's probably easier to leave them in.
- ineedasername 5y agoThere's no reason to remove them, that's not what I'm asking. By all means leave them in, but why leave them undocumented? Explain their existence, their parameters & capabilities. If not intended for use, explain that too. Then when something unexpected like Spectre comes along, the people that have to deal with it can say "Oh yeah, those testing instructions provide another vector of attack that our patch has to account for." Instead we're in this situation, and I'm pretty sure there's at least a half dozen nations that would have already devoted the resources needed to uncover undocumented instructions like this, meaning ample opportunity to have developed various exploits.
- amluto 5y agoMy response: https://lore.kernel.org/lkml/CALCETrXRvhqw0fibE6qom3sDJ+nOa_aEJQeuAjPofh=8h1Cujg@mail.gmail.com/ https://lore.kernel.org/lkml/CALCETrXRvhqw0fibE6qom3sDJ+nOa_... I don’t think any new mitigations are needed.
- ForOldHack 5y agoThis had to come. The only fix will be to add a BIOS setting for Speculative Access or no speculative access. Gamers all turn it on, with a machine patched, that runs nothing but their game. Everyone else, like browsing the web, off. Look for a encoded binary java script exploit that will own any speculative access system. Its coming too, just like this paper would eventually come.
- bruce343434 5y agoThis is not feasible. Not everyone has multiple computers. And not everyone wants to use different computers for different things, or take the effort to muck around in the (often mazelike) bios settings.
- juancn 5y agoThis may sound stupid, but the commonality in all these side channel attacks is that high precision time keeping is a non privileged operation. Maybe it’s time to make clocks a privileged op as a mitigation. Even making execution time non predictable on untrusted code, such as JavaScript? If precise time keeping is unavailable these become harder to do.
- lukastr0 5y agoIt is surprisingly difficult to make timekeeping unavailable. There are many methods, besides the official timer APIs, to get timing - as outlined in the paper "Fantastic Timers and where to find them" from TU Graz: https://www.researchgate.net/publication/322000263_Fantastic_Timers_and_Where_to_Find_Them_High-Resolution_Microarchitectural_Attacks_in_JavaScript https://www.researchgate.net/publication/322000263_Fantastic...
- vmception 5y agoOut of curiosity, is Apple's M1 processor seemingly faster because it is actually more similar to a normal CPU progression but all the other common CPU's - x86 - had retroactive performance hits due to patching Spectre. And therefore M1 seems so much more faster than it otherwise would?
- creshal 5y agoARM CPUs, including Apple's have been found vulnerable to some SPECTRE variants. As far as I understand, M1 already contains hardware mitigations for all known ones, but so do the latest Intel/AMD chips. (Or rather, contain fixes for all but these latest ones.)
- jlouis 5y agoI'm more inclined to believe M1 is fast because it's a modern design where every part of the architecture is under Apple's control. In particular, the design uses a memory hub with the memory chips very close to the CPU core. And it has a massive L2 cache on top.
- SG2000 5y agoA close reading of the paper “I see dead uOps” would seem to indicate that Intel’s static thread partitioning of their micro-op cache would confer some inherent protection against uOp cache information leakage between threads - as compared to AMD’s dynamic thread partitioning scheme which could theoretically allow threads to spy on each other using the described techniques. If true, wouldn’t this also imply that an Intel Skylake CPU mitigates against such attempted attacks by one user against another in a shared CPU/ISP/cloud environment, whereas an AMD CPU theoretically would not? If true, this would be a key point that the authors failed to mention in their concluding remarks. Anyone else read it this way? Or am I missing something?