16 ms·
LLVM patch to fix half of Spectre attack
- okneil 9y agoThe site is down for me. HN hug of death?
- arboroia 9y agoGoogle text cache: https://reviews.llvm.org/D41723 https://reviews.llvm.org/D41723 Wayback Machine: https://web.archive.org/web/20180104131631/https://reviews.llvm.org/D41723 https://web.archive.org/web/20180104131631/https://reviews.l... :)
- hultner 9y agoIt was a bit slow but eventually loaded for me.
- mayoralito 9y agoYeah, same thing happened to me... slow as hell but I guess it's common due the severity of the issue. All people wants to see this at the same time.
- XnoiVeX 9y agoYes. Give it about 5 minutes. It will load without images.
- ealexhudson 9y agoThis patch apparently implements this mitigation: https://support.google.com/faqs/answer/7625886 https://support.google.com/faqs/answer/7625886
- kough 9y agoThis is a really good writeup, thanks. I'm curious -- how often are google support faq articles deeply technical like this?
- JdeBP 9y agoAnd once one knows the technical background, one is better positioned to consider the response of Linus Torvalds to the idea that the entire Linux kernel be recompiled for all x86 CPUs with a compiler that implements this. * https://lkml.org/lkml/2018/1/3/797 https://lkml.org/lkml/2018/1/3/797 (https://news.ycombinator.com/item?id=16066968 https://news.ycombinator.com/item?id=16066968)
- deleted 9y ago[deleted]
- sempron64 9y agoIt's noted in the patch that one would have to recompile linked libraries, which seems impractical, unless a distro decides to build everything with this flag.
- jacquesm 9y agoNot just linked binaries, also the whole underlying OS, and, critically, the compiler itself. Otherwise you could replace the 'proofed' construct with one that is not proofed against the bug.
- JDevlieghere 9y agoWhy would you need to recompile the compiler? Both variants only provide read access.
- jacquesm 9y agoAh right, of course. Sorry, in the midst of doing a pile of stuff I should not be commenting on this without studying it further, I figured that the first level read access would allow you to dig up the secrets required to give you write access which would then allow you the free run of the whole system, but if you are still on the other side of a virtual machine then that won't do any good unless that virtual machine can be escaped as well.
- imtringued 9y agoAnd since this patch is opt in it isn't enough to secure cloud providers.
- badrequest 9y agoI, for one, am eternally grateful for the incredibly bright people who take the time to patch this sort of stuff.
- ben_jones 9y agoAnd the people who invented computers, programming languages, the internet, and all the learning resources, that allow me to get a paycheck writing extremely high level application code that feels like a coloring book in comparison. Truly the shoulders of giants.
- jacksmith21006 9y agoAlso to Google for finding and documenting it so well. Google security team really should be given an award.
- fooker 9y agoretpoline seems to be a novel concept. Can anyone ELI5? Also, any insight about performance impact here?
- sanxiyn 9y agoBy design, with retpoline indirect branches won't be able to take advantage of branch prediction. This is nontrivial, but can't be helped. Performance impact should be negligible otherwise.
- dingo_bat 9y agoConsidering that every single function call into a dynamically loaded library will be affected, that negligible "otherwise" won't be so negligible in the real world.
- revelation 9y agoretpoline is just a convoluted way of doing an indirect jump/call designed to make branch prediction entirely useless. It's a novel concept because doing this is completely opposite to making a program run faster. Here is an example of the most common programming patterns that end up causing indirect jumps/calls: https://godbolt.org/g/eThmnG https://godbolt.org/g/eThmnG Imagine every virtual function call in a C++ program being mispredicted and taking twice as long. (Instead of forcing us to recompile the world, maybe Intel should just disable branch prediction in microcode.)
- littlestymaar 9y ago> Imagine every virtual function call in a C++ program being mispredicted and taking twice as long. > (Instead of forcing us to recompile the world, maybe Intel should just disable branch prediction in microcode.) Wouldn't the performance impact be dramatic ? In this[1] example there's a 6 times slowdown between situation with and without correct branch prediction. [1]: https://stackoverflow.com/questions/11227809/why-is-it-faster-to-process-a-sorted-array-than-an-unsorted-array#11227902 https://stackoverflow.com/questions/11227809/why-is-it-faste...
- coldcode 9y agoI remember doing tricks like this in 6502 assembly and in other early processors. Amazing that to stop these attacks you have to come up with clever tricks again. Back in the 80's I would have never imagined this type of attack being something to worry about.
- FLUX-YOU 9y ago>early processors Early processors had speculative execution? I thought this had been added to Intel/AMD/ARM about 20 years ago?
- DiThi 9y agoI think it means they're tricks for better performance when you _don't_ have speculative execution.
- dzdt 9y agoI guess he means the retpoline. On the 6502 there is no indirect jump instruction, so you need such tricks just to achieve an indirect jump at all.
- pubby 9y agoThere's an indirect jump instruction. It's not very good though, and has a notorious bug with addresses ending in 0xFF.
- dzdt 9y agoGah, you're right. Guess my memory is fading. There is indirect JMP, but no indirect JSR or indirect branches. And the indirect JMP as you say is not very useful.
- gpderetta 9y agoSpeculative execution is as old as branch prediction, which is very, very old.
- tptacek 9y agoPage was down when I tried to read it, but it's archived here: http://archive.is/s831k http://archive.is/s831k. Its hard to get your head around how big a deal this is. This vulnerability is so bad they killed x86 indirect jump instructions. It's so bad compilers --- all of them --- have to know about this bug, and use an incantation that hacks ret like an exploit developer would. It's so bad that to restore the original performance of a predictable indirect jump you might have to change the way you write high-level language code. It's glorious.
- jncraton 9y agoAgreed. I haven't had this much fun thinking through the implications of a new exploit technique in a long time. It is truly beautiful.
- voidmain 9y agoAnd I fear there's little reason to think that the "three variants" from project zero's announcement are the full scope of the problem. They were just the variants that the few people in on this found time to develop exploits for. There can now be security bugs in things your program doesn't do; it seems like there is room for nearly unlimited creativity in finding them. From the spectre paper: "A minor variant of this could be to instead use an out-of-bounds read to a function pointer to gain control of execution in the mis-speculated path. We did not investigate this variant further."
- deleted 9y ago[deleted]
- pdpi 9y agoARM’s white paper details a variant 3a that affects some of their cores that are unaffected by var3 (and vice versa)
- geertj 9y ago> d I fear there's little reason to think that the "three variants" from project zero's announcement are the full scope of the problem. Agreed. This is an entirely new class of vulnerabilities, and we're just at the beginning.
- deleted 9y ago[deleted]
- dzdt 9y agoWhen using these patches on statically linked applications, especially C++ applications, you should expect to see a much more dramatic performance hit. For microbenchmarks that are switch, indirect-, or virtual-call heavy we have seen overheads ranging from 10% to 50%. Ouch! This is independent of other performance hurts, like from the kernel syscall overhead that was the hot topic yesterday. This is pretty crazy.
- jerf 9y agoThat's bad. A single 5% hit might not be the end of the world, but 5% here and 10% there and another 5% over there in the common case adds up badly enough. Doubly-pathological cases (indirect calling-heavy code calling lots of syscalls)... a 50% slowdown and a 30% slowdown combines to a 60% total slowdown. Yeowch. Will be intrigued to see how processor manufacturers respond to this. If they were even slightly relaxed about it prior to disclosure I expect there's going to be some very hurried attempts to engineer some solutions pronto. This is the sort of thing where it might even be worth throwing away all of your future roadmap plans and just getting a revision of the current chips out there ASAP, whatever that may do to the rest of your roadmap.
- ant6n 9y agoSounds like it could be great for processor manufacturers. In the age where CPUs don't get faster, there's finally a reason for customers to buy new CPUs again!
- CmdDot 9y agoNot really, once a program is compiled with -retpoline, new hardware won't bring back reliable branch prediction. I'd hope maybe, just maybe, this would be enough to put a focus on compilers producing code that ends up using processor-optimized paths chosen at runtime, to avoid "overheads ranging from 10% to 50%". Though, in this case, that would essentially mean making the entire executable region writable for some window of time, which is clearly too dangerous, so I guess the 0.1% speedups from compiling undefined behavior in new and interesting ways, will continue taking priority. I mean, it's a compiler flag right, obviously whoever's going to run a program on an unaffected platform will take the effort to recompile everything with the flag removed. Just the same way every serious application currently provides different executables for running on systems where SSE2, SSE4.1, or AVX2 is present.
- phkahler 9y agoRISC-V impact? With all the reports of these attacks, I have not seen mention of risc-v. Since they are in the process of finalizing a lot of specs including memory model and privileged instructions, I wonder if there will be last minute changes to mitigate these vulnerabilities.
- leoc 9y agoAt the risk of being a HN self-parody, I’ve also been wondering what this means for the Mill... https://millcomputing.com/docs/prediction/ https://millcomputing.com/docs/prediction/
- Tuna-Fish 9y agoThe details that this attack depends on are outside the architecture of the system, in the microarchitecture. A cpu of almost any architecture can be vulnerable or not depending on how it was implemented, thus Ryzen is immune to the worst variant while both Intel and the fastest Arm cpus are vulnerable. I'd presume that the slowest RISC-V designs are immune due to not speculating enough, while any high-performance implementation is vulnerable.
- ars 9y agoAs of right now every single CPU that does speculative execution. (I.e. runs both sides of a branch then throws away the one that didn't end up being valid.)
- bem94 9y agoThe problem (in my understanding) is not with the specification of the x86 ISA, but with the implementation of the speculative execution micro-architecture and probably the memory sub-system as well. That is why Intel is so badly affected by the problem, but not AMD, despite them both implementing the same instruction set. RISCV has already had to fix its memory consistency model, so it is not without problems. But it that is a spec bug, not an implementation bug. Whether there is an out of order, speculative execution RISCV core in the wild which suffers from this is as far as I know very unlikely. If there is, no doubt it's designers have had a busy time lately.
- contrarian_ 9y agoNote for a true fix to the BTB poisoning attack you would additionally have to disable SMT/HT. See here: https://news.ycombinator.com/item?id=16070304 https://news.ycombinator.com/item?id=16070304
- eptcyka 9y agoMill can't come soon enough.
- mike_hearn 9y agoWhat makes you think the Mill would be immune to these issues?
- eptcyka 9y agoMill has no speculative execution.
- phs2501 9y agoUh, yes it does. It has no /out of order/ execution but it certainly has speculative execution. They mention quite often in talks how they predict EBB exits and not each branch. They even go so far as to follow that EBB exit chain to speculatively load code from DRAM several calls ahead, which is much more speculation than current CPUs are capable of. You basically can't make a deeply-pipelined processor fast without speculative execution.
- marcosdumay 9y agoA simple model of access permissions that fit before L1 cache and can return a fault before loading anything.
- vfaronov 9y agoI have a hunch that the era of side-channel attacks is only now dawning, and that we should expect many more painful exploits and cumbersome mitigations in the coming years. What do people more knowledgeable in the field think about this?
- Klathmon 9y agoI'm not more knowledgeable than you, but I think I agree. side channels have always been some of the most insidious exploits. Many are basically un-solvable (timing attacks are always going to leak some information, and compression is basically completely at odds with secure information storage), many more are easily enough overlooked that it would be easy to maliciously include them without raising any eyebrows, and the "fixes" for them almost always murder performance. I think the only real fully-encompassing solution to this is a redesign in how we use computers. Either a massive step backwards on performance and turning off most automatic "optimizations" until they can be proven through a much more rigorous process (both in compilers, and in hardware), or a significant change in how computers are architected adding more hardware level isolation for processes and systems running on the machine (just daydreaming now, but something like a cluster of isolated micro-CPUs that run one application only).
- dingo_bat 9y agoHow about not running untrusted code? That seems to be much easier to do and won't kill performance. Kill js on the web. Run only apps signed by Microsoft. Develop ML based malware fingerprinting that can recognise timing attack patterns. Throwing OO execution away shouldn't be an option in the long term.
- Klathmon 9y agoIt's not just "untrusted code" any more. The "true secure" way is not running any untrusted code, not connecting to any untrusted networks, and not accepting or storing any untrusted data. At that point you can't run a computer. It's just not an answer to say "don't let bad things in" because bad things are always going to get in. And with side channel exploits getting more and more common, and with them being worse and worse, running any code on your machine is basically giving that code any and all of your data on the machine... Telling users (even highly technical users) to "never ever run any untrusted code ever, and if you mess up once and run untrusted code you have completely ruined the trust of the whole machine and need to start over from scratch while assuming all of your data has been compromised" is not only infeasible, it's impossible. If this is the case, we have lost the "security" game. It's easy to say "throwing oo execution away" is an overreaction, but if it's necessary to allow multiple programs to run on one machine without them all having what amounts to full access to one another's information, then it might be necessary. Already we know that we can't use compression with encryption, we can't use any kind of "exit early" with most kinds of encryption. It might just be that OO execution is fundamentally opposed to secure computing. At it's core, it's letting the processor do different things depending on what it can see is coming, it's almost a definition of an oracle! That's always going to be a very dangerous game to play. It could even just be that CPUs need a "secure computing" mode, or maybe even a secure co-processor that disables all of these optimizations. But at the very least, I think changes are going to be necessary, and a 15% perf reduction might be the least of our worries.
- leni536 9y agoIt has an interesting performance impact on calls to dynamic libraries. One alternative approach would be to avoid the indirect calls through not using '-fPIC --shared' when building shared libraries but '-mcmodel=large --shared'. This causes the relocations to happen at the direct calls and not through a GOT. The obvious drawback that it effectively disables sharing code in memory, it would still allow sharing code on disk though. So it would be a middle ground between the current state in dynamic and static linking. https://www.technovelty.org/c/position-independent-code-and-x86-64-libraries.html https://www.technovelty.org/c/position-independent-code-and-...
- andrewmcwatters 9y agoIn other news, Intel has found that by not using a computer at all, though performance overheads increase 100%, this counter-measure does secure any previously available attack vectors.
- peapicker 9y agoThis brings to mind Ken Thompson's "Reflections on Trusting Trust"[1] -- after all, all I have to do to write code with the exploit is be able to remove the patch and rebuild the compiler and build some executables. Trusting in a compiler you hope was used to build all the executables on your system isn't trustworthy enough to be the final solution. [1] https://www.win.tue.nl/~aeb/linux/hh/thompson/trust.html https://www.win.tue.nl/~aeb/linux/hh/thompson/trust.html
- pwg 9y agoEvery modern compiler usually has extensions that allow for bits of assembly to be inserted alongside the usual C or C++ code. Unless the compiler is also patched to either disallow inserted assembly, or to modify the inserted assembly (this being both hard and dangerous), someone who wants to exploit the bug will just add their own inserted assembly code that exploits the bug, and a patched compiler won't help one bit in that case.
- Pelam 9y agoMaybe some future architecture will allow software to tell CPU which regions it considers to be secret from the point of view of each other region. Something like that could allow the CPU to speculate agressively while preventing information leak exploits.
- jacobolus 9y agohttps://millcomputing.com/docs/ https://millcomputing.com/docs/ e.g. the most recent talk https://millcomputing.com/docs/threading/ https://millcomputing.com/docs/threading/
- Pelam 9y agoSomething like the portal calls and "turfs" described in there could help.
- pwg 9y agoThe CPU hardware already has that feature. It is the VM paging system and the permissions assigned thereto. The bug here is that the CPU is not aborting the speculation when fetches occur to addresses marked as "access denied". Instead the fetch happens and a line of normally inaccessible memory is put into cache by code that should not be able to get it read into the cache normally. One hardware fix would be to plug that hole. Speculative reads get blocked when they encounter permission denied errors from the paging system and do not change the cache state. That blocks the Meltdown attack, but not the Spectre attack.
- Pelam 9y agoI thought about that too... AFAIK currently paging system is not generally accessible to userland programs like browsers. They would need some way to setup different contexts for untrusted javascript code and the internal services that the javascript can call. Also maybe the context switching would need to be made faster, because you would need to do that whenever eg javascript calls browser interfaces.
- nathell 9y agoI can't help thinking of how the early-ITS approach to security (not only was there none, but looking at other users' work was a deliberate feature) was embraced by its users. I'm way too young to remember, but it rings a bell somewhere down my heart. There's a lot of prominence being given to all kinds of damage malicious users might inflict, and ways to prevent or mitigate, but little to the malice itself. Whence does it arise? What emotions drive those users? What unmet needs? Meanwhile, when these slowing-down patches for Sceptre and Meltdown arrive, I intend to not run them, to the possible extent. I intend to keep aside a VM with patches for critical stuff, like banking or others' data entrusted to me. But I don't want my machine to be slowed down just because someone, sometime, might invest effort in targeting these attacks at it. Given how transparent I want to be with my life, that's a risk I'm willing to take.
- fwip 9y agoMost attacks aren't targeted at specific people. Hackers don't want to read your emails, they want your credit-card information, digital account passwords, or to compromise your computer to use in their botnet. Sure, you might not have anything you want to hide in your life, but the drive-by javascript doesn't care about your secrets - it'll hack you anyway. Best-case scenario, you lose access to a bunch of accounts you used to use and need to create new identities from scratch. Worst-case, they clean you out financially, steal your identity, etc.
- teilo 9y agoIsn't it the case that the Itanium architecture would not be vulnerable to Spectre because it moves the onus of branch prediction from the CPU to the compiler?
- als0 9y agoAssuming the compiler knows what it's doing :)
- teilo 9y agoThat was always the problem with the Itanium compilers. They were crap because they couldn't benefit from the years of tuning traditional architectures enjoyed.
- acdha 9y agoAlso the compiler had to be absolutely brilliant to rewrite the serial branching code most programmers wrote to work with the EPIC model. They had some good results optimized math-heavy code but the general purpose code ended up with too many nops waiting on results.
- jzl 9y agoA new thing that's going to become a standard part of systems engineering: deciding whether any given system needs to run with or without these kinds of protections. Do you want the speed of speculative execution or do you want Meltdown/Spectre protection? In some cases lack of protection is fine. But figuring out the answer for any given system is often going to take expert-level security knowledge. Security is all about multiple layers of protection, and even a non-public facing machine might benefit from these layers depending on the context.
- crb002 9y agoCPUs should have a single instruction that wipes branch prediction caches. I would have it off by default, and add to the C/C++ spec this as a standard library macro or pragma. Easy peasy. You only need to wipe between syscalls that have side effects. Number crunching AVX heavy subroutines should never have to deal with safety once entered.
- ece 9y agoThis is what KPTI does, wipe caches, and if you did this often in user code, performance degradation would be all over the place. Also, heavy AVX routines that use encryption keys... would be great to attack.
- s4vi0r 9y agoSpectre relies on tricking the CPU into branch predicting its way into accessing protected memory, no? Is it not possible that we can keep most of the performance benefits of speculative execution by somehow having a built in "Hey, never ever speculate that I'll want to access this region of memory" sort of thing?
- senatorobama 9y agoUh, isn't this what AMD does?
- lorenzq 9y agoI read an ars technica article that this would be a possible solution but isn’t right now because the hardware to check access rights isn’t fast enough yet
- silimike 9y agoIf this were 15 years ago, I'd say the site was SlashDotted.
- jgowdy 9y agoThe problem I see with this concept is ROP mitigations like Intel’s control flow enforcement don’t seem compatible with intentionally using tweaked addresses with ret. The address they inject won’t match the shadow stack and the program will be terminated.
- DannyBee 9y agoThis is true, and so far, nobody has a better idea. (IE i would expect that unless someone comes up with one, that hardware CFE in its current form dies and won't happen for Intel until the processors are changed in a way that mitigation is not needed)
- crb002 9y agoThis was the fix I was going to suggest. Especially with AVX leakage. Right now many function calls don't safely wipe registers and the new side channel caches found in Spectre. There really needs to be two kinds of function calls. Maybe a C PRAGMA? The complier has parent function call wiping as a flag; the code has pragmas that over-ride the flag.
- userbinator 9y agoThis is horrible, really really horrible. And I'm not talking about the bug itself, but the mitigation --- which is basically "stop using indirect jump and call instructions and recompile all your software". The latter is beyond unrealistic. It also sets a very bad precedent: I understand people want to mitigate/fix as much as possible, but this is basically giving an implicit message to the hardware designers: "it doesn't matter if our instructions are broken, regardless of how widespread in use they already are --- they'll just fix it in the software."
- hn_throwaway_99 9y ago> it doesn't matter if our instructions are broken, regardless of how widespread in use they already are --- they'll just fix it in the software. What are any other options? It's hardware, that cannot be patched. Of course they will change chip designs going forward, but what else do you suggest folks do with the billions of chips that exhibit this problem?
- ychen306 9y agoGo ahead, smash your computer, wait a few months, and buy a new one.
- cws125 9y agoJust as a FYI, according to: * https://lkml.org/lkml/2018/1/4/432 https://lkml.org/lkml/2018/1/4/432 * http://xenbits.xen.org/gitweb/?p=people/andrewcoop/xen.git;a=blob;f=xen/arch/x86/spec_ctrl.c;h=79aedf774a390293dfd564ce978500085344e305;hb=refs/heads/sp2-mitigations-v6.5#l168 http://xenbits.xen.org/gitweb/?p=people/andrewcoop/xen.git;a... It appears that Skylake and later can actually predict retpolines? Some hardware features called IBRS, IBPB, STIBP (not a lot of details on this are out there) are supposedly coming in a microcode update.
- lousken 9y agowhat about performance impact after new CPU architecture arrives? how is that going to work?
- rntz 9y agoThis mitigates spectre variant #2, branch target injection. We also have a mitigation for meltdown, namely KPTI. Is there a known mitigation for spectre variant #1, bounds check bypass? Maybe I'm being naive, but would a simple modulo instruction work? Consider the example code from https://googleprojectzero.blogspot.com/2018/01/reading-privileged-memory-with-side.html https://googleprojectzero.blogspot.com/2018/01/reading-privi...: unsigned long untrusted_offset_from_caller = ...; if (untrusted_offset_from_caller < arr1->length) { unsigned char value = arr1->data[untrusted_offset_from_caller]; ... } If instead we did: unsigned char value = arr1->data[untrusted_offset_from_caller % arr1->length]; Would this produce a data dependency that prevents speculative execution from reading an out-of-bounds memory address? (Ignore for the moment that a sufficiently smart compiler might "optimize" out the modulo here.)
- AaronFriel 9y agoThis is brutal for all interpreted/JITed languages and all statically compiled languages with dynamic dispatch. I can hardly imagine worse news for performance oriented engineers. And what's worse is that dynamic libraries will probably need to be rebuilt with these mitigations in mind, so nearly everyone will pay the cost even if they don't need it. I feel bad for all of the engineers currently working on performance sensitive applications in these languages. There's a whole lot of Java, .NET, and JavaScript that's about to get slower[1]. Enterprise-y, abstract class heavy (i.e.: vtable using) C++ will get slower. Rust trait objects get slower. Haskell type classes that don't optimize out get slower. What a mess. [1] These mitigations will need to be implemented for interpreters, and JITs will want to switch to emitting "retpoline" code for dynamic dispatch. There's no world in which I don't expect the JVM, V8, and others to switch to these by default soon.
- strongholdmedia 9y agoAs Alex Ionescu has put it: > We built multi-tenant cloud computing on top of processors and chipsets that were designed and hyper-optimized for > single-tenant use. We crossed our fingers that it would be OK and it would all turn out great and we would all profit. > In 2018, reality has come back to bite us. This is the root of all the problems.