13 ms·
What about ECC memory? Intels desktop processors fail hard on that feature.
by std_throwaway 10y ago
What about ECC memory? Intels desktop processors fail hard on that feature.
- krzyk 10y agoWhy do you need that in desktop env? What's the usecase when it can be beneficial to have expensive ECC RAM instead of the ordinary one? I was always thinking that ECC can prevent blue screens/kernel faults, but I haven't seen those in years on my laptop without ECC.
- eliaspro 10y agoBesides many obvious reasons (preventing real bitrot), it's also a security factor acting against attacks like rowhammer.
- mrb 10y agoThe symptoms of memory corruption are much more varied (and can be much more stubtle) than BSODs. For example they can cause filesystem corruption, any random app crashes (SIGSEGV...), infinite loops, etc.
- grndzro 10y agoThat is why it is important after building/tweaking a system to do stability tests. If it runs full out without errors for 24h then you are fine.
- deaddodo 10y agoThat's.....not how non-ECC works. Even a little. RAM bitflips randomly, period. It's just how it works. A cosmic ray can hit the memory chip just right and flip it, there's no way to predict or control that no matter how "stable" your machine is. ECC still does the same, it just has a parity bit on each line to confirm against and flip it back, as needed.
- justincormack 10y agoIt is how faulty RAM works, and it is sensible to do that.
- Rebelgecko 10y agoNot just faulty RAM. As far as I know, all modern RAM is going to have bitflips every once in a while (it even happens with ECC RAM)
- deaddodo 10y agoI urge you to educate yourself. All RAM works this way, it's the very reason ECC RAM exists and the exact issue it solves.
- lisivka 10y agoI have no idea, why you downvoted. Stability test is important step of cluster building. If probability of something broken in new server is 10% then probability of something broken in new cluster of 10 servers is 1-(1-0.1)^10 = 65% .
- masklinn 10y agoBecause they objection has nothing to do with the original comment. Bitflips are not a RAM stability issue.
- lisivka 10y agoBitflips are not a RAM stability issue, they happen randomly due to radiation, but radiation is not random, especially in my area (I live not far from Chornobyl).
- masklinn 10y agoSure, but you won't solve excessive ambient radiation by running a stress test on your RAM.
- lisivka 10y agoA small radioactive («hot») dust particle will not change average radiation level a lot, but may cause problems with memory/cpu. Simple cleaning, by blowing dust out, fixes it. Saw that dozen of times, but years ago.
- EdHominem 10y agoDid you use a geiger counter to verify the dust was hot? Would something large enough to cause problems with the ram be large enough to detect?
- flamedoge 10y agoer.. you should probably move. Bitflips aren't the only things radiation does. Bitflips in DNA for example.
- kogepathic 10y ago> What's the usecase when it can be beneficial to have expensive ECC RAM instead of the ordinary one? Anyone using ZFS will (or should!) care about ECC support. [0] Lots of people build their own NAS/SAN boxes, so ECC support on a desktop CPU at a reasonable price point would be very appreciated. Currently you need to buy specific model CPUs (Celeron or Xeon, IIRC) to get ECC support from Intel. [1] [0] https://serverfault.com/questions/454736/non-ecc-memory-with-zfs-a-stupid-idea https://serverfault.com/questions/454736/non-ecc-memory-with... [1] https://ark.intel.com/search/advanced?ECCMemory=true&MarketSegment=DT https://ark.intel.com/search/advanced?ECCMemory=true&MarketS...
- snuxoll 10y agoThis is why I ultimately ended up buying an HP ML10 for my FreeNAS box, $300 for a Xeon E3-1220, board and iLO ended up being cheaper than buying a barebones + the CPU (which is just shy of $200 retail). I would like to have something that supports more memory, but to get anything with more sockets or support for larger DIMM's you start looking at the Broadwell-E or Xeon-E5 chips and those are considerably more expensive (my ThinkServer TD340 cost me $700 as an open box with a E5-2403v2 and 8GB of RAM installed). If Zen client chips have full ECC support and can handle 64GB of memory I'll be sold easily on an upgrade for my TrueNAS box, then I'll anxiously await some lower-cost (4/8c) dual-socket server CPU's and swap out my TD340 with some supermicro barebones build.
- lultimouomo 10y agoDo you know how much power does it use? If so, could you add some details (number of HDDs, typical load - light, medium, heavy)? I'm looking into building a home server and would like something decently beefy, not too power hungry and possibly with ECC.
- snuxoll 10y agoI know the ML10 only has a 300W PSU, but I haven't bothered to take a kill-a-watt as the ML10 itself would be outshined by the rest of my lab gear. The E3-1220 is usually idle and frequency scaled back unless something like a ZFS scrub is running, I've got ~45W of PCIe cards (SAS HBA, 10GBe NIC and 4x1GBe NIC), 4x8GB sticks of DDR3 UDIMM's probably uses 12W. I'd say without drives it probably draws no more than 150W on average, unless I'm doing something CPU intensive. My drives all sit in an external SAS enclosure, since I wanted more than 4 drive bays that the LFF expansion bracket provided. Anyway, the ML10 is pretty decent for light-medium work. It's got 4 really fast cores and 32GB of RAM is adequate for most home server use (this was why I got the TD340 though, I've got 72GB in it) - just beware that the Gen1 units don't include the drive bracket so if you want more than a single HDD you have to purchase one or get an external SAS enclosure. EDIT: the ML10 is basically silent too, even under load - I can't hear the fans unless I try (though my SAS enclosure makes up for this by being the loudest bit of kit in my lab).
- jpalomaki 10y agoIt is kind of funny that we even accept non-ecc memory and say that it is perfectly OK that sometimes a 0 written to memory turns to 1. This would be more understandable if we treated the memory as unreliable but in most cases we don't.
- dchest 10y agoECC (error-correcting codes) don't completely eliminate errors — they reduce probability of them happening — so even with ECC memory you're still accepting that it's perfectly OK that 0 is sometimes 1, just a lot less likely. No absolutes here, unfortunately.
- temac 10y agoThis EXTREMELY reduces the probability that a bit flips happens. And consumer PCs (or mobile devices) are the only devices where there is no ECC memory, RAM is the only place in such devices where there is no error detection and/or correction.
- vegabook 10y ago100%. Everywhere else including spinning disk, SSD, ethernet, TCP/IP there is error correction.
- vegabook 10y agoThis question comes up constantly. Anybody doing anything with more than 16GB, especially data science (think all the R and Python people, all scientists), anybody running a memory cached data store (all server side people on mongo, redis, etc), and anybody doing finance or engineering, wants ECC. Basically anybody doing anything where data persistence is important, and/or where even the slightest chance of silent corruption is catastrophic. And with Intel you had to spend well north of 1k to get it on an 8/16 system. Hoping Ryzen does it.
- semi-extrinsic 10y ago> Anybody doing anything with more than 16GB, especially data science (think all the R and Python people, all scientists) Not sure I agree. Most science datasets (both from simulations and experiments) are sufficiently noisy that if your scientific end results and conclusions change as a result of even thousands of bitflips in your 16 GB of data, you're Doing It Wrong and your article isn't worth the paper it's printed on. (There are probably exceptions, as always, but those working in those few specific subfields should be aware of it.)
- vegabook 10y agoAs soon as you are putting a noisy dataset through multiple iterative algos, though, if a bit error hits the control flow code or data structure delimiters, you face the possibility of massive silent, propogating corruption. Imagine your hash table getting a bit error in it (all R lists and Python dicts use them), or an incorrect branch in your code. Admittedly the risk is ultra low, but you have enough problems to worry about to have, in addition, a niggling sensation of RAM-risk as you work. More generally, it's not unreasonable to say that our entire computing paradigm rests on accurate RAM. Above 16GB, the risks just become too big for anybody doing serious work, and not just messing around with prototypes.
- semi-extrinsic 10y ago> Admittedly the risk is ultra low Exactly this. The size of your code is positively infinitesimal compared to your data. And unless you're writing your code and then running it exactly once, which is a) even more unlikely and b) bad practice, you'll catch any of those bit-flip errors in your code or data structures. This has been discussed a lot in the literature, especially for GPUs where ECC carries a performance penalty both on speed and available memory, e.g. in this paper where they've tested it on a GPU cluster: http://www.rosswalker.co.uk/papers/2014_03_ECC_AMBER_Paper_10.1002_cpe.3232.pdf http://www.rosswalker.co.uk/papers/2014_03_ECC_AMBER_Paper_1...
- fulafel 10y agoDesktop computers are sometimes used for actual work where data integrity is important. Only a small minority of main memory data corruptions lead to OS crashes, mostly the in-memory application or filesystem data just silently gets corrupted.
- std_throwaway 10y agoThat is like asking: "Why do you wear a seatbelt if you don't even drive a racecar?" Because I care about my data. Data corruption may kill your main storage and the first backup, too.
- abrookewood 10y agoThere's no need to down vote him. ECC is extremely uncommon on desktops & laptops. Yes, it's always going to be preferential to have it, but it typically comes at a price point that excludes it from these platforms.
- temac 10y agoThat's a problem, not a manifesto to continue to live with it... And if the prices difference are large, they are largely artificial. That means if ECC is more widely uses, prices difference will decrease to reasonable levels (at the marginal level, price of RAM with ECC should be ~= 9/8 price of RAM without ECC, and the price of supporting platforms should be only very slightly higher than the price of non-supporting platforms)
- abrookewood 10y agoFair enough - I'd like to see it happen, just don't think it will.
- aidenn0 10y agoECC RAM isn't that much more expensive despite being less common. It's a bigger premium to go from an i7 to Xeon than to equip the Xeon with ECC RAM. If ECC were manufactured in the volumes of non-ECC it would probably be almost exactly 12.5% more expensive (i.e. just the cost of the extra chip).
- temac 10y ago> It's a bigger premium to go from an i7 to Xeon than to equip the Xeon with ECC RAM. Not exactly. The prices between comparable i7 to Xeon are nearly the same.
- aidenn0 10y agoWow, I just priced skylake Xeon's vs i7 and prices are very comparable (in some cases the Xeon is even cheaper). Current Xeon v. i7 was significantly different last year when I purchased. I did check and it is still true that with laptops you always pay a premium for Xeon v Core on an equal performance basis (even within the same model).
- mtgx 10y agoI believe ECC RAM protects against RowHammer, doesn't it? That would be reason enough to get it.
- temac 10y agoIt mitigates. IIRC DDR4 also mitigates. With both at least I'm pretty sure RowHammer is not a security problem anymore, and maybe even the risk of crashing disappear completely (or have an extremely small probability)
- chatterbeak 10y agoAll of my desktops are Xeons with ECC. Maybe you or your customers don't care, or you just play games and browse the web. But if you're doing something critical, you want to be able to detect and repair a bit flip. Google estimated that 1 bit/GB/year flips. There's 128 GB in our standard desktop built. That would be 2 bit flips/week.
- krzyk 10y agoWhy the downvotes? It is a legitimate question, because ECC is very uncommon in desktops and laptops. Companies I worked for used those only on the servers where we run critical stuff, not a single developer laptop had ECC.
- tilt_error 10y agoSome links that may give more background to the area: Coding Horror To ECC or Not To ECC [1], What Every Programmer Should Know About Memory [2], Memory Errors in Modern Systems [3], and an analysis of memory errors in the entire fleet of servers at Facebook over the course of fourteen months [4]. [1] https://blog.codinghorror.com/to-ecc-or-not-to-ecc/ https://blog.codinghorror.com/to-ecc-or-not-to-ecc/ [2] https://people.freebsd.org/~lstewart/articles/cpumemory.pdf https://people.freebsd.org/~lstewart/articles/cpumemory.pdf [3] https://www.cs.virginia.edu/~gurumurthi/papers/asplos15.pdf https://www.cs.virginia.edu/~gurumurthi/papers/asplos15.pdf [4] https://users.ece.cmu.edu/~omutlu/pub/memory-errors-at-facebook_dsn15.pdf https://users.ece.cmu.edu/~omutlu/pub/memory-errors-at-faceb...