12 ms·
Protecting Google Cloud customers without impacting performance
- Jerry2 9y ago[censored]
- wmf 9y agoThat's not unusual; if a story gets one or two upvotes very soon after being submitted the HN ranking algorithm will put it on the front page. The rank then drops very quickly unless it gets more votes.
- deleted 9y ago[deleted]
- jaflo 9y ago"Retpoline ... modifies programs to ensure that execution cannot be influenced by an attacker. With Retpoline, we could protect our infrastructure at compile-time, with no source-code modifications" I am confused, doesn't this mean that Retpoline needs to sit in the compiler and won't protect from already-compiled binaries?
- itcmcgrath 9y agoYes. We deploy new code binaries all the time so it's not a big deal.
- hhw 9y agoThat was my question also. How could they guarantee that nobody else using a VM on the same physical machine isn't using binaries compiled elsewhere?
- puzzle 9y agoGoogle doesn't run VMs. Or, rather, even VMs that make up GCE are actually running in containers. From the Borg paper: https://pdos.csail.mit.edu/6.824/papers/borg.pdf https://pdos.csail.mit.edu/6.824/papers/borg.pdf VMs and security sandboxing techniques are used to run external software by Google’s AppEngine (GAE) [38] and Google Compute Engine (GCE). We run each hosted VM in a KVM process [54] that runs as a Borg task. There might be a few machines here and there running actual VMs on bare metal, but those are going to be very special cases. Using Infrastore, it's easy to find which teams are running which containers whose binaries were built before the new compiler defaults. https://github.com/kubernetes/kubernetes/issues/44095 https://github.com/kubernetes/kubernetes/issues/44095
- cthalupa 9y agoI don't understand how you think that answers his question, as being in a container wouldn't magically add any more protections against these exploits than being in just a KVM VM. I'm also not sure that "Borg task" means it's in a container, vs. just being something scheduled by Borg, but it isn't particularly important either way. If the question is 'How do they guarantee their customers aren't using bad binaries', the answer is they don't, and has nothing to do with VMs vs VMs in containers, etc.
- puzzle 9y agoI should have been clearer: their external customers can and will use bad binaries inside their VMs, but everything else, especially anything at the boundaries between VMs, will not: it will have the appropriate mitigations in place. Running a VM in a Borg task (container) does not add extra security in this case, but it does help with auditing and catching the snowflakes and stragglers that invariably end up taking a disproportionate amount of time with changes like this. That kind of auditing can be easily done organization-wide, instead of having every team reinvent the wheel their own way.
- foota 9y agoIt's my understanding that Spectre needs to share some address space with the victim, which isn't the case on containers, other than the host address space?
- Tomte 9y agoWhat does "retpoline" stand for? ret for the ret instruction, I suppose, but the rest?
- niftich 9y agoThe name “retpoline” is a portmanteau of “return” and “trampoline.” It is a trampoline construct constructed using return operations which also figuratively ensures that any associated speculative execution will “bounce” endlessly. (If it brings you any amusement: imagine speculative execution as an overly energetic 7-year old that we must now build a warehouse of trampolines around.) [1] [1] https://support.google.com/faqs/answer/7625886 https://support.google.com/faqs/answer/7625886
- UncleMeat 9y ago"Trampoline". The term actually shows up with some frequency among compiler people.
- Tomte 9y agoI actually know that word in the compiler meaning, I think I just never got the connection, because I was pronouncing it in my head as ret-poh-line, not ret-poh-leen. Thanks!
- clon 9y agoGood for them, giving the credit to Paul Turner. Unusual for an enterprise to allow cracks in the corporate "We".
- azinman2 9y agoGoogle is better than most in this regard. For one, it allows its employees to publish papers with their names on it. It also frequently lists names in press releases of key members.
- hueving 9y agoEvery company that publishes research allows employees to put names on it. It's not really an indicator of a forward thinking company so much as one that wants prestige in the academic world.
- discoursism 9y agoPaul is a boss. This is just one of his many amazing accomplishments. He's kinda the guy you go to if things are slow and you need some help to make them really ridiculously fast.
- make3 9y agoThey usually do for researchers bringing new ideas, like on ai research papers
- kyrra 9y agoDirect link to a description of the fix: https://support.google.com/faqs/answer/7625886 https://support.google.com/faqs/answer/7625886 Someone on SO trying to explain it another way: https://stackoverflow.com/a/48099456 https://stackoverflow.com/a/48099456 Some interesting discussion about how this patch isn't a 100% fix for Skylake processors (at least that's my understanding): https://www.mail-archive.com/linux-kernel@vger.kernel.org/msg1577670.html https://www.mail-archive.com/linux-kernel@vger.kernel.org/ms...
- x0x0 9y agoI don't understand how this fix can possibly be so performant. The speculative executions must be doing something positive, right? It seems like it should never hurt (ignoring security, obv). So how can disabling it not hurt? Thanks for the stackover flow link!
- bradleyjkemp 9y agoSounds like they're just forcing the CPU to use a different (less vulnerable) type of predictor. It's the branch target predictor that's vulnerable so instead of using a branch instruction the compiler now adds a mini stack frame (with the return address as the place we want to jump to) and then "returns". The CPU now happily predicts the target using the non-vulnerable stack predictor and does do speculative execution.
- tzar 9y ago> By December, all Google Cloud Platform (GCP) services had protections in place for all known variants of the vulnerability. Could other major cloud providers boast this? It seems like Google's brand is benefiting tremendously from Project Zero in all of this. On the other hand, it feels nervous-makingly like a clear step towards running one's own mainstream hardware being too hard for the little guys.
- kardianos 9y agoI agree it is impressive. But if you are running your own hardware, you probably aren't sharing CPU time with strangers as GCP is.
- tzar 9y agoThat's a fair point, but somehow the nervous feeling isn't limited to sharing CPU time with strangers! The amount of platform complexity abstracted by the GCP services is staggering. This is the job of taking a (hopefully) understandable piece of computer code that represents some real-world logic and installing it into reality, in a sense. It's clear that it's a very messy world out there for a program.
- boulos 9y agoDisclosure: I work on Google Cloud. Even if you aren't sharing with "strangers", applications may be vulnerable to these attacks. For example, the JavaScript attack clearly applies to people's individual computers. If you take untrusted code and execute it, you might be vulnerable to these intra-"instance" information leaks. It all comes down to your threat model though. Some people are rightly worried about insider risk. If a malicious employee can go run a binary on your shared computing infrastructure to get root credentials out of a machine, that's actually a real problem. Then again, there are lots of ways for a rogue employee to do bad things, so this is "just" another one. But don't take this as "only applicable to shared cloud environments", because it's not.
- qaq 9y agoSure but when you are worried about insider risk (and even if you are not) you most likely have strong access controls, extensive logging, endpoint security solution that can mitigate or at least alert on this activity.
- deepnotderp 9y agoWait... doesn't Reptoline have some irritating performance penalties?
- DannyBee 9y agoWe use automatic feedback directed optimization[1], which lowers the cost quite significantly in practice [1] https://research.google.com/pubs/pub45290.html https://research.google.com/pubs/pub45290.html Performance gains are even higher now than they were then.
- boulos 9y agoDisclosure: I work on Google Cloud. Not particularly (if you read Paul's post, the branch to the retpoline predicts perfectly for obvious reasons), and especially not compared to the brute force flushes as an alternative. Edit: I phrased that backwards. The return predicts, so that the whole thing is about as bad as an unpredicted indirect call: > This has the particularly nice property that the RSB entry and on-stack target installed by (1) is both valid and used. This allows the return to be correctly predicted so that our simulated indirect jump is the only introduced overhead.
- cthalupa 9y agoEdit: My post was prior to the parent edit, and now is largely unnecessary. Keeping for posterity, I suppose! I must be misunderstanding Paul's post. Isn't it specifically preventing any sort of prediction? "Naturally, protecting an indirect branch means that no prediction can occur. This is intentional, as we are “isolating” the prediction above to prevent its abuse." Of course, you can then go and manually add direct branch hints, as is noted in the post, but unless I'm misunderstanding things, there's not an obvious reason why these branches predict perfectly. Not that it means performance is impacted in a significant way, since that same section also says "Microbenchmarking on Intel x86 architectures shows that our converted sequences are within cycles of an native indirect branch (with branch prediction hardware explicitly disabled)." (which also confuses this issue - how is it predicting perfectly if prediction hardware is disabled?)
- yRetsyM 9y agoI've been impressed with how long it took for this information to be "leaked"/"declassified"/released given the sheer number of people who knew. The fact that hundreds maybe thousands of people knew and worked on this ahead of the press reports/rumors and subsequent information release speaks to how seriously everyone involved took this.
- cwzwarich 9y agoI don't think it was a complete success. An AMD engineer mentioned the nature of the vulnerability on lkml before the end of the embargo, even though AMD was a party to the embargo. Of course, it was probably unwise to discuss the code changes implementing mitigations in public anyways, but I don't have first-hand knowledge of how these things work in the Linux world.
- mjevans 9y agoWasn't that also a reply to some part of this which had already surfaced to the public (but the impact maybe not known to the public).
- cwzwarich 9y agoThere was an email thread discussing page table isolation mitigations for some bug, but Tom Lendacky mentioned it was about speculation. This let many others finally connect the dots, despite the earlier discussion taking place in the open for a while. Until the connection with speculation was revealed, it could've just been another KASLR info leak defense.
- masterleep 9y agoWe can safely assume that black hats have moles on these lists.
- throwaway2048 9y agoyeah the idea you can keep something like this away from malicious parties across dozens of huge companies is a ridiculous farce.
- bogomipz 9y agoA bit off topic but the amount of real estate taken up by the header, side nav and "related articles" footer on this is just obnoxious. Obnoxious to the point of making reading this a really rotten experience. I fear this is the medium.com effect of content on the web now. Simply having content for content's sake is now seen as a missed "growth hacking" opportunity.
- benjaminjackman 9y agoAlso the low contrast, hard (for me atleast) to read font.
- aroman 9y agoSad that this was collapsed-by-default — I realize it's not relevant to the article, but it likely never will be. If we complain enough hopefully someone can actually address it.
- stefanha 9y agoAccording to this page retpoline is "insufficient on Skylake and newer CPUs, where even ret may predict from the indirect branch predictor as a fallback; those need IBRS". https://github.com/marcan/speculation-bugs/blob/master/README.md#bti-linuxgccllvm-retpolines https://github.com/marcan/speculation-bugs/blob/master/READM... What gives?
- tw04 9y agoThe way I read that, Skylake and newer don't see a performance hit from ibrs. So it's retpoline on older CPUs = no performance hit. Ibrs on newer CPUs + microcode update = no hit. >Requires microcode update on current CPUs. Perf hit vs. retpolines on older CPUs
- stefanha 9y agoYou're missing the point. It says "insufficient on Skylake and newer CPUs". The Google Cloud blog post does not mention that retpolines are insufficient on newer CPUs. That omission is dangerous because readers will believe they are protected by recompiling with retpolines when in fact they aren't. Perhaps the blog post can be updated to clarify the limits of retpolines so people don't get the wrong impression and end up with vulnerable systems.
- boulos 9y agoThat's incorrect (in our opinion). See my comment at https://news.ycombinator.com/item?id=16130871 https://news.ycombinator.com/item?id=16130871
- stefanha 9y agoOkay, the link I posted is just a writeup summarizing publicly available information, some of which may be misleading or incorrect. On the other hand, "opinion" isn't enough. Everything depends on the microarchitecture and only Intel can give assurance on whether retpoline actually works in Skylake or not. I hope they will release information about it.
- Abishek_Muthian 9y agoIt's good to see Google crediting Retpoline to Paul Turner. As Senior Staff Engineer, Technical Infrastructure I wonder whether he was actually tasked with working on mitigation for these vulnerabilities or he came up with this in his free time.
- boulos 9y agoAs the article notes, lots of people worked basically round the clock on this as a top priority. That isn't to say they did nothing else, but Paul et al. definitely didn't just happen to do this in their spare time :). The relevant quote from the article: > For months, hundreds of engineers across Google and other companies worked continuously to understand these new vulnerabilities and find mitigations for them.
- sandGorgon 9y agoDoes anybody know if Retpoline will make it to the compiler that the Linux Kernel is compiled with ? It doesn't specifically mention in this paper, so I'm not able to figure out what was actually compiled using Retpoline - userland or the Linux kernel itself ?
- boulos 9y agoDisclosure: I work on Google Cloud. Patches for both LLVM (the infrastructure behind clang) and gcc are available. You choose what you compile your kernel and applications with, and others are actively looking at retpoline and retpoline-inspired techniques for other code generators (e.g., various JIT compilers). That's why Paul and the folks made this public.
- sandGorgon 9y agothanks for your reply. in Google Cloud's specific case, what was compiled using Repoline ? Was it the kernel or userland ? Because you are kind of claiming that the patches that the linux kernel used to fix this (with a 10% drop in performance) are no longer needed. I am kind of wondering if this was submitted to the core linux kernel to be mainlined ? If yes, why was this not used there. If no, then it means it is being used in some Google Cloud-specific way that is not mainlineable. EDIT - found this comment which seems to suggest that retpoline is not bulletproof and the kernel's performance-killing patches are still needed https://github.com/marcan/speculation-bugs/blob/master/README.md#bti-linuxgccllvm-retpolines https://github.com/marcan/speculation-bugs/blob/master/READM... http://lkml.iu.edu/hypermail/linux/kernel/1801.0/03137.html http://lkml.iu.edu/hypermail/linux/kernel/1801.0/03137.html Does it mean that Google Cloud is doing this only on non-Skylake CPU instances ? It is a very interesting stand - it will mean that it will suddenly be more cost effective to use Google's OLDER machines than the newer Skylakes... because the newer machines have a performance degradation that the older machines will not suffer.
- boulos 9y agoDisclaimer: I'm not Paul :). As mentioned elsewhere, we recompile everything at Google all the time. I'm not sure which things we've rebuilt with retpoline enabled. As Paul mentions in the article, the point is for software you believe needs to be protected, which may not be everything we build. That thread on retpoline on Skylake has a lot of confusion. For some folks, they aren't 100% certain it works (it relies on understanding internal details of Intel CPUs) and they argue that IBRS on Skylake is cheap enough so "why not just always use IBRS and not bother?". That's the gist of this comment: > personally I am comfortable with retpoline on Skylake, but I would like to have IBRS as an opt in for the paranoid. I want to highlight that Paul and the team have had a lot more time to think about this issue than the folks just joining the discussion. Could our folks be missing something? Sure, and that's the point of public discussion and code review. We hope that over the coming weeks and months it's decided one way or another, but we believe retpoline to be correct and a good optimization (especially for older hardware).
- cm2187 9y agoThe post seems to suggest not all CPUs are affected by Variant 2. Is it Haswell and earlier only?
- edf13 9y agoA very good reason why people are going to start moving to GCP over AWS.... Project Zero is a big win here.
- thinkMOAR 9y agoInteresting google has time, money etc for this. But actually showing search results on page 4 of google search or youtube, when it said there were 22 million results for my search seems too hard for them.
- bfrog 9y agoHas this at all caused google to reconsider the heterogeneous nature of the cloud in terms of hardware? It seems like Google the company is constantly fixing/redoing various intel problems such as ME and now this. Google is part of openpower after all, it would be interesting to see another architecture being pushed.