10 ms·
Project Zero: Exploiting the DRAM rowhammer bug to gain kernel privileges
- j_baker 12y agoYou know, this makes me wonder. If a car manufacturer or a toy company made a product that was found to be unsafe, there would be a recall. If hardware manufacturers make a product that is insecure, will there be a recall? Unfortunately, I suspect that this is a case where the law hasn't caught up with technology.
- Morgawr 12y agoA few years ago I built a home PC for myself and bought an i5 sandy bridge processor with an appropriate motherboard. A few months later it was found out that a huge batch of the SATA controllers shipped on those types of motherboards were faulty[0]. Back then, Intel made a statement recalling all faulty motherboards and shipping out new ones, I just contacted my retailer where I purchased my board, sent it for RMA and got a new one (different model, but that's another story). All of this for free. [0] http://www.pcadvisor.co.uk/news/pc-components/3259061/intel-sandy-bridge-recall-what-you-need-to-know-updated/ http://www.pcadvisor.co.uk/news/pc-components/3259061/intel-...
- EvanAnderson 12y agoIntel has a good history of recalls and replacements of their motherboards and processors. The Pentium FDIV bug comes to mind immediately, as does the recall of motherboards with the faulty 820-series memory translation hub.
- pacificmint 12y agoActually, Intels behavior with the FDIV bug was originally anything but good. They downplayed the bug and refused to recall them. Then they started offering replacements if you could prove that the bug affected you. It wasn't until the whole thing turned into a giant PR disaster that they started a generous exchange program. That whole affair is basically the reason that Intel is much more forthcoming with errata these days.
- joosters 12y agoIn the EU, products have to be fit for purpose. You could then argue that if you bought (for example) a server for hosting virtual machines, then the RAM was not fit for purpose because the flaw made it incapable of isolating separate VMs. Good luck trying that though!
- yuhong 12y agoOn the other hand, servers tend to use ECC memory.
- davidw 12y agoYes. This happened 20 years ago: http://en.wikipedia.org/wiki/Pentium_FDIV_bug http://en.wikipedia.org/wiki/Pentium_FDIV_bug
- jerf 12y agoOf the five vendors that they mentioned, the only one that did not have vulnerable memory was "DRAM vendor D", which also only had one entry on the table. Given the nature of the problem here, odds strike me as near-1 that "DRAM vendor D" has shipped RAM with this problem. For that matter, the "no"s on that table really only prove that the exact stick they tested with the exact memory locations they tested did not exhibit detectable bit flips. It doesn't prove that those sticks are "safe", let alone that the product line they come from is safe. So, basically, what's vulnerable? To a first approximation, everything. What would happen if we tried to recall every bit of DRAM produced in the past X years (where X is also unknown)? Well... you'd bankrupt the industry is what you'd do. That's not a very useful outcome. In fact this sort of thing happens all the time. New safety tech is developed for cars all the time, but you can't go back and sue the auto companies for not including it before it was invented or the need for it was discovered [1]. This seems more like that problem than an actual problem of negligence or "defects" being produced. [1]: Well... more or less. I know of cases where this was successfully done, though they tend to get overturned on appeal. Run with me here.
- pdpi 12y agoCan you get killed as a result of privilege escalation? The law hasn't caught up in part because the potential consequences aren't nearly as dire.
- datenwolf 12y agoModern medical technology relies heavily on computers and software. Take an infusion pump for example. Controlled by a microcontroller and using software. Or insulin pumps; and some vendors are actually considering to add Bluetooth to insulin pumps, so that patients using such a pump can check its status on their smartphone (or on the upcomming smart watches). Also you can adjust the infusion rate of an insulin pump to accommodate for ingested sugar. Overdosing on insulin can send a person into shock and kill.
- TeMPOraL 12y agoIt's an interesting attack vector, recently covered by Person of Interest episode, in which an abusive husband got killed by having his insulin pump wirelessly hacked and making him overdose the drug. While fiction, I'm pretty sure this kind of thing will happen (after all, no one writes bug-free software, and even if, you can always steal the keys...) - and initially will be very hard to detect because of its uncommon nature.
- tedunangst 12y agoIf somebody is running their ramhammer exploit on your insulin pump, it's probably a bit late.
- AnthonyMouse 12y ago> Modern medical technology relies heavily on computers and software. Which is why medical devices should all have ECC memory. And for that matter physical separation between any processor that might run attacker-controlled code and the processor responsible for That Which Must Not Fail. Product defects like this are foreseeable. If bad memory can cause a medical device to kill someone, the party at fault is the one who made a medical device without sufficient redundancy and error correction that bad memory could cause it to kill someone.
- userbinator 12y agoIt's not just insecure, this is memory that doesn't work 100% like memory should. I use MemTest86+ on every stick of DRAM I buy - if there's even a single error, it goes back as defective. The fact that this memory seems to work for most access patterns doesn't excuse the fact that it is completely broken for others, because good memory should be able to store any data and maintain its integrity for any access pattern. Unfortunately even MemTest86+ is not exhaustive, as I found out while troubleshooting a very strange issue: a specific file in a specific archive would unpack with corrupted bits (and an "archive damaged" message) on a coworker's computer, but on half a dozen other machines would be fine. A hash of the file matched, so HDD-based corruption was ruled out. His machine passed an overnight run of MemTest86+ perfectly and AFAIK unpacking no other archives would yield corruption. He reported never getting any crashes - but yet, that one file in that archive would fail to unpack correctly. It would always corrupt in the same strange way. On a whim, I decided to swap the RAM out and the problem went away. Even the "bad" stick seemed to work fine in other machines with the same model of CPU and mobo running the same OS and unpacking the same archive, but with his extremely specific combination of hardware and software, would always fail. That experience taught me that bad RAM can be extremely difficult to troubleshoot. This isn't like other storage technologies e.g. SSDs where their finite lifespan and sensitivity to access patterns is well-documented. It's a case of claiming to sell memory while giving consumers a close approximation of one that completely breaks in some situations. I think it needs to be treated like the FDIV bug.
- 0x0 12y agoDoes anyone know if Macbooks are known to be affected?
- joosters 12y agoDownload the tool and try it yourself? It supports Mac OS and there is a mailing list to report affected machines (nothing seems to be posted there yet).
- aselzer 12y agoSeems like my Macbook Air 2014 is not affected (with a high probability) here's the test: https://github.com/google/rowhammer-test https://github.com/google/rowhammer-test
- jacquesm 12y agoHow long did you test it? The tests they did ran fairly long, possibly you'd have to run this for days to really be able to state that a particular machine/ram combination is not vulnerable.
- tux3 12y agoThanks for the link. I haven't seen anything after 375 iterations (600s). So I may still be exploitable, but that means you'd have to keep something running at 100% CPU for > 600s and somehow have me not notice the laptop fans going crazy.
- Dylan16807 12y agoAn exploit tool could always run slower and hide from that. Also consider that it might work better when your laptop is in lower power mode because of reduced voltages.
- thomasdullien 12y agoYou may wish to try both single- and double-sided hammering. If you hit the right row size it is significantly more effective: https://github.com/google/rowhammer-test/blob/master/double_sided_rowhammer.cc https://github.com/google/rowhammer-test/blob/master/double_...
- zokier 12y agoSurprised that the mitigations section did not mention ECC RAM. Wouldn't it be effective mitigation?
- stefantalpalaru 12y ago"We also tested some desktop machines, but did not see any bit flips on those. That could be because they were all relatively high-end machines with ECC memory. The ECC could be hiding bit flips."
- lelf 12y agoNot necessary, see the original paper. For example, SECDED (single error-correction, double error- detection) can correct only a single-bit error within a 64-bit word. If a word contains two victims, however, SECDED cannot correct the resulting double-bit error. And for three or more victims, SECDED cannot even detect the multi-bit er- ror, leading to silent data corruption. Edit: link http://users.ece.cmu.edu/~yoonguk/papers/kim-isca14.pdf http://users.ece.cmu.edu/~yoonguk/papers/kim-isca14.pdf
- acveilleux 12y agoTechnically, SECDED cannot reliably detect errors involving more then 3 bits since they might generate a valid code, they might not however and in that case they might be detected as single or double bit error or possible something else.
- yuhong 12y agoAlso the typical reaction to an uncorrectable ECC error is to halt the system with a NMI.
- makomk 12y agoYeah, ECC is going to make exploiting this reliably a lot harder - you'd need to flip three or more bits in the right combination, without first hitting a combination of bits that'd be detected as an uncorrectable error. Google's report suggests they haven't even been able to cause uncorrectable two-bit errors yet, let alone undetectable three-bit ones.
- ymra 12y agoWould reducing the speed memory is clocked at prevent this?
- TheLoneWolfling 12y agoAnything that reduces the number of times memory can be accessed between refreshes can mitigate this, reducing RAM clock (probably) included
- thrownaway2424 12y agoYes, it would, as would overvoltage, and reducing the refresh interval. The latter reduces memory subsystem performance, however.
- rasz_pl 12y agorow access counters in memory controller would solve this problem - too many accesses between refresh cycles -> force refresh cycle for that particular row/potentially affected rows
- jacquesm 12y agoLaptops are particularly at risk for stuff like this: components are more densely packed and may use smaller process sizes and have less powerful supplies which may be a factor in keeping bits in adjacent rows stable. That may be the reason why the desktops mentioned are less sensitive, they'll use full size memory modules and will have beefy power supplies. It'd be interesting to repeat the experiments with the laptops running off their internal battery.
- Aissen 12y agoAlso, lower refresh rates on DRAMs means less power consumption (so it's an easy fix in BIOS, independent of OS, clearly attractive to laptop makers), but also more exposition to this issue.
- Aissen 12y agoVery little information on time scales. In one case they speak about 5 minutes vs 40 minutes (both might be acceptable for an exploit). Also no information about how long it took to bitflip in their per-hardware table. And why name no hardware vendor ? I'm guessing they expect people to use the tool they provided and draw their own conclusions, but I don't understand why they'd treat them differently from software vendors.
- jacquesm 12y agoAt a guess to avoid labeling laptop manufacturers and getting sued if it turns out that something else was at fault? The DRAM itself might be the culprit (probably is), laptops of a certain brand might come with RAM from different manufacturers.
- Aissen 12y agoI understood the litigation risk. In an integrated system it's always someone else's fault (DRAM, BIOS, CPU, laptop vendor). IMHO the last integrator (the one selling you the goods) is always the culprit. Why would they fear hardware manufacturers' litigation more than software vendors' ? Especially at such a big company like Google ?
- acveilleux 12y agoThey also don't want to say "DellappLenoHP" laptops could not be attacked and turn out to be wrong. Or maybe they're right but only with factory 2GB modules used between May '11 and July '13. Way too many variables to make any claims that is ethically defensible.
- userbinator 12y agoThey could specify the detailed system configuration with the CPU, chipset, and DRAM part numbers (including date codes) so others can compare. It's much better than leaving things in the dark completely.
- ChuckMcM 12y agoOnce again, I pine for ECC memory on my Laptop. I know you can get ECC SODIMMS, I got 16GB worth for a Supermicro ITX motherboard. And while the paper talks about multi-bit errors getting through ECC (which is certainly possible with enough flips) single flips causing alerts and double flips causing halts would really get your attention that something bad was happening. As opposed to silently sitting there while my memory is shredded.
- thrownaway2424 12y agoI don't think laptops have SODIMM memory these days.
- ChuckMcM 12y agoSadly true, the 'thin is in' crowd is more often than not soldering in the memory.
- jacquesm 12y agoI bought one two weeks ago that has two SODIMM slots.
- sspiff 12y agoDepends on the laptop. I have bought two laptops in the past 14 months, a $2000 ThinkPad and a $600 Acer. Both came with 4GB soldered on and a single free SODIMM. (On a different note: the ThinkPad maxes out at 8GB and the Acer at 12GB, whereas previous generations went up to 16GB at least. Intel intentionally nerfed Haswell and newer core i's, presumably to push their Xeons on more people)
- yuhong 12y agoIntel intentionally nerfed Haswell and newer core i's, presumably to push their Xeons on more people Really? I think the latest Haswell can go to 16GB just fine if there are two SO-DIMM slots.
- 12y ago
- kmowery 12y agoThe starting research that enabled this security work appeared last year at ISCA, but didn't fully discuss the security implications: https://www.ece.cmu.edu/~safari/pubs/kim-isca14.pdf https://www.ece.cmu.edu/~safari/pubs/kim-isca14.pdf
- userbinator 12y agoI noticed the security implications of "memory that doesn't always behave like memory" when that paper came out a few months ago and was discussed briefly on HN: https://news.ycombinator.com/item?id=8713411 https://news.ycombinator.com/item?id=8713411
- upofadown 12y agoIs a memory error actually an exploit? If so then are the unwanted changes that occur with no deliberate action an example of the computer cracking itself? Philosophical...
- jdmichal 12y agoErrors can be used as part or all of an exploit. Exploiting a system requires that ethereal value of "intent", and I don't think anyone would (currently) argue that computers can have intent. Without that intent, it's just an error.
- sharkbot 12y agoI think there is a useful distinction between a fault/error and an exploit. A fault is a break from the "desired" or "expected" semantics of a system, while an exploit is an algorithm to predictably utilize a fault (or faults) to access unexpected behaviours in that system. I.e., a buffer overflow is a fault in a program (breaking the expectation that a buffer's contents will remain within a certain bound), while an exploit targeting that overflow will likely allow running arbitrary code in a program not designed to do so. So, I'd put it, the memory error can be leveraged in an exploit.
- Dylan16807 12y agoEverything is a memory error on some level. Back to grounding in reality, a way to reliably[1] break security measures is an exploit. Cosmic ray bit flips are anything but reliable. [1]The threshold of reliability being somewhere below "instant and always" and somewhere above "one in a million if you give it a day to try".
- sharkbot 12y agoThere was an older paper discussing using various methods of fault injection (heat, voltage changes, etc) to attack Java smart cards, essentially destroying the type system guarantees and thus opening up an attack surface: "The Sorcerer’s Apprentice Guide to Fault Attacks", https://eprint.iacr.org/2004/100.pdf https://eprint.iacr.org/2004/100.pdf
- bri3d 12y agoFault injection is also how older Dish Network and DirecTV smart cards were hacked - there used to be a cottage industry selling "voltage glitchers" to reprogram Dish Network smart cards with the keys for additional programming tiers.
- makomk 12y agoI believe some pay TV smartcard hacks also made use of clock glitching, basically sending a shorter-than-usual clock pulse that means some of the internal signals don't make it to their destinations on time. The pay TV hacking industry had some pretty clever tricks a decade or two ago.
- Scoundreller 12y agoThey were quite cool. From memory, I think one card had some internal startup check that checked to see if its EPROM got marked by the "Black Sunday" countermeasure and then hung itself. The hackers, having a ROM dump and having knowledge of how many clock cycles each instruction took the CPU, knew that it was at ~clock cycle 525 or so that this internal check happened. Knowing that the instruction was a "Branch if equals to" (I think), and that instruction took 12 cycles, they figured out which of those 12 caused that branch to happen, figured out the precise time to glitch (whether via voltage or a single rapid clock cycle), and caused the CPU to skip changing the instruction pointer and then continue through its ROM code as if the check had passed. Within a month or two, hundreds of thousands of receivers had a man-in-the-middle device just to glitch reprogrammed cards every time they were started up. Apparently the north american provider had tested the same countermeasure in their south american division, so the north americans had advance notice of what they had to do to get back in action. I recall, for another system, a small memory chip was required for a pre-existing man-in-the-middle card, and overnight every electronics supplier went out-of-stock overnight. Digikey sold out of 50k units overnight.
- yuhong 12y agomemtest86 etc should add tests for this if they didn't already, as this is the best place for such tests.
- otakucode 12y agoIf they did so... the fallout would be interesting. Does anyone know what proportion of modern memory has this flaw? Would it result in tens of thousands of customers returning stick after stick of DRAM until they were able to get a reliable one?
- tenfingers 12y agomemtest86 has this feature in beta, and it's already generating some heat. I would be personally more interested in this test on memtest86+ though.
- AceJohnny2 12y agoWhat's the difference between memtest86 and memtest86+? OK, from WP [1]: "Memtest86 was developed by Chris Brady. After Memtest86 remained at v3.0 (2002 release) for two years, the Memtest86+ fork was created by Samuel Demeulemeester to add support for newer CPUs and chipsets. As of November 2013 the latest version of Memtest86+ is 5.01." And the original has become a commercial program by PassMark. So I think at this point if anyone is talking about memtest86, they're likely referring to the still open-source '+' version. [1] http://en.wikipedia.org/wiki/Memtest86 http://en.wikipedia.org/wiki/Memtest86
- hannob 12y agoThere is a github repo with a rowhammer test based on memtest86+: https://github.com/CMU-SAFARI/rowhammer https://github.com/CMU-SAFARI/rowhammer
- Kenji 12y agoNow someone has to come up with a JavaScript version of this exploit and the disaster is complete.
- yuhong 12y agoMore difficult since you can't execute CLFLUSH from there.
- Kenji 12y agoMaybe some JS commands trigger a CLFLUSH internally. I don't know, but it'd be "funny" if that exploit worked in JS.
- makomk 12y agoMaybe, though as they say it'd potentially be possible to cause a cache spill and attack it that way. I was looking at the associativity of various CPU caches with a vague eye to trying this in JavaScript a few days back and in theory it shouldn't take many reads to evict a cache line, so long as they're from the right addresses.
- deleted 12y ago[deleted]
- randomdevlpr 12y agoMy first gen Toshiba Chromebook ran the test 130 minutes without an error.
- p1mrx 12y agoOn my desktop (DH87RL / i7-4770 / 2x8GB Crucial DDR3L-1600), rowhammer_test reported errors after ~20 iterations (less than a minute). I went into the BIOS and tried lowering the tREFI value from 6300 to 3150 (not sure what the units are). So far, it's gone 1000 iterations with no problems detected. Edit: Actually, the units are probably multiples of the cycle time, just like CAS latency. So, for DDR3-1600, that would mean 6300x1.25ns=7.8μs, and 3150x1.25ns=3.9μs http://en.wikipedia.org/wiki/CAS_latency http://en.wikipedia.org/wiki/CAS_latency
- thomasdullien 12y agoSingle-sided or double-sided hammering?
- p1mrx 12y agoI used rowhammer_test.cc, which I think is single-sided.
- throwaway41597 12y agoI tried and it reported one error under a second. I had to reboot because gcc started to make bash crash, it seems. Then I saw the README (duh!): Be careful not to run this test on machines that contain important data. On machines that are susceptible to the rowhammer problem, this test could cause bit flips that crash the machine, or worse, cause bit flips in data that gets written back to disc. **Warning #2:** If you find that a computer is susceptible to the rowhammer problem, you may want to avoid using it as a multi-user system. Bit flips caused by row hammering breach the CPU's memory protection. On a machine that is susceptible to the rowhammer problem, one process can corrupt pages used by other processes or by the kernel. (Mine is Kingston Hyper X 2x8GB DDR3 1600MHz)