4 ms·
Invalidating the caches is kind of a cringe inducing approach on this (actual) problem. Especially in HPC radiation related single event upsets have become a re
by datenwolf 8y ago
Invalidating the caches is kind of a cringe inducing approach on this (actual) problem. Especially in HPC radiation related single event upsets have become a real problem. If you do the math, all the silicon area devoted to memory (DRAM, caches, registers) adds up, and what you've got is essentially particle detector.
Compared to the effective volume of a purpose designed one (ATLAS, CMS, Super Kaminokade, etc.) rather small, but a particle detector nevertheless.
A couple of months / years ago, there was an article (also linked here on HN, IIRC) that did a few back of the envelope calculations regarding expected event rates. IIRC it was something on the order of 1 event per day per 10^12 transistors. (EDIT: not the one I thought of but blows the same horn: http://energysfe.ufsc.br/slides/Paolo-Rech-260917.pdf http://energysfe.ufsc.br/slides/Paolo-Rech-260917.pdf )
Also radiation hardened software has been researched (and still is). Essentially the idea is to not only have redundant, error correcting memory, but also redundant, error correcting computation. NASA has some publications on that. e.g. https://ti.arc.nasa.gov/m/pub-archive/1075h/1075%20(Mehlitz).pdf https://ti.arc.nasa.gov/m/pub-archive/1075h/1075%20(Mehlitz)...
- Nokinside 8y agoCould physicists and astronomers use all this distributed particle detection for research? Like a app in a phone or desktop that sends reports back with location and time information.
- ajb 8y agoMost bit flip events are apparently due to alpha particles from radioactive decay in the package (source - worked for a company which used 'low alpha compound' in some of our ICs to put off having to implement HW error correction) , which would be a big confounding signal. I imaging physics experiments work hard to avoid this, or at least correct for it.
- mikeash 8y agoThus the demand for low-background steel, mostly recovered from pre-atomic-age ships: https://en.wikipedia.org/wiki/Low-background_steel https://en.wikipedia.org/wiki/Low-background_steel
- evanb 8y agoYou may be interested to learn of arXiv:1510.07655, "Detecting particles with cell phones: the Distributed Electronic Cosmic-ray Observatory" by Vandenbroucke et al. I don't know if any novel results have come out of this kind of thing. https://arxiv.org/abs/1510.07665 https://arxiv.org/abs/1510.07665
- privong 8y agoThere was also this one, a bit earlier, "Observing Ultra-High Energy Cosmic Rays with Smartphones". https://arxiv.org/abs/1410.2895 https://arxiv.org/abs/1410.2895 I (thought I) signed up to be informed of beta releases, but never heard anything. I just checked their website[0] and it mentions a beta app, but that seems to just go to a signup page. [0] https://crayfis.io https://crayfis.io
- Create 8y agoTo me, this seems like AstroTurf PR for the quoted manufacturer's Movidius™ Myriad™ 2, just tested at [hope to be trendy HEP institution]. Pity the photo with friendly smiles from the press package was left out, but the curious will find it with a brief search.
- biggerfisch 8y agoProbably not, due to the low fidelity of any data from this. "real" particle detectors can trace the decay chain, momentum changes, precise energy levels, etc. Processor cache flips etc can only say "a particle with energy >= X was here", which isn't super useful given the relative frequency of occurrence.
- bumby 8y agoNASA calls out preparing for this specifically in their software engineering requirements. My understanding is that this was added specifically to address bit flips from radiation effects. See section 3.7.2.f https://nodis3.gsfc.nasa.gov/npg_img/N_PR_7150_002B_/N_PR_7150_002B_.pdf https://nodis3.gsfc.nasa.gov/npg_img/N_PR_7150_002B_/N_PR_71...
- Pica_soO 8y agoThe reason this is researched is also, that it would allow for cheaper manufacturing, as nearly 100 % non radioactive waver raw material is hard to come by. So, the incentive is not just space and science resistant hardware, but computers, who survive there own flawed material with the occasional decay.
- skunkworker 8y agoAn article I read awhile ago addressed an interesting correlation between transistor process size, the physical size of the dram module and the expected failure rates. "As transistor sizes have shrunk, they have required less and less electrical charge to represent a logical bit. So the likelihood that one bit will "flip" from 0 to 1 (or 1 to 0) when struck by an energetic particle has been increasing. This has been partially offset by the fact that as the transistors have gotten smaller they have become smaller targets so the rate at which they are struck has decreased. More significantly, the current generation of 16-nanometer circuits have a 3D architecture that replaced the previous 2D architecture and has proven to be significantly less susceptible to SEUs. Although this improvement has been offset by the increase in the number of transistors in each chip, the failure rate at the chip level has also dropped slightly. However, the increase in the total number of transistors being used in new electronic systems has meant that the SEU failure rate at the device level has continued to rise." [1] https://phys.org/news/2017-02-particles-outer-space-wreaking-low-grade.html https://phys.org/news/2017-02-particles-outer-space-wreaking...
- oh_sigh 8y agoHow much shielding would it take to prevent most of these energetic particles from reaching the important bits?
- posix_me_less 8y agoIt depends on what you mean by "most". Even shielding 95% may not be sufficient, if the remaining 5% is too strong. It is a tradeoff - the thicker the plating, the better protection, but beyond some limit the plating gets too heavy, which is very costly especially when it is to be put in orbit.
- ChuckMcM 8y agoAnd vendors who put ECC on the data paths but not on the cache itself, on the assumption that check it while it is going in, its a cache so it is short lived, and you're window is small enough to meet your reliability goals. That goes out the window when you sleep though because you don't know how long you've been waiting to restart. And the P(bitflip) is a function of time. If you sleep too long you are non-spec compliant for silent data corruption, and since there isn't a way to know how long you sleep, the only "safe" option is to invalidate the cache and reload it. Sad but an understandable approach. The downside is your wake from sleep is slower by the amount of time it takes to warm up the cache.
- nonbel 8y ago>"1 event per day per 10^12 transistors" Why is this in such an inconvenient form? If I have x gb of ram how many unavoidable memtest errors should I expect per hour of testing? It seems like that could be used to tell us the minimum amount of time to run the test.
- Filligree 8y agoDRAM uses one capacitor and one transistor per bit, iirc, so about one bit-flip per two years per GB.
- nonbel 8y agoThanks, I saw that same value after a quick search but wasn't sure about it. That seems a bit high. So if you have 128 gb, you would expect 128/104 ~ 2.5 bitflips per week, about every other day?
- PhasmaFelis 8y ago128/104 would be ~1.23 bitflips per week.
- nonbel 8y agoHa, yea. Here is the calculation I did (carelessly): > 128/54 [1] 2.37037 Still, one a week seems high.