5 ms·
Bit worrying not to see a single mention of ECC memory when discussing protection from bitrot, especially with filesystems which depend on correctly functioning
by Freaky 13y ago
Bit worrying not to see a single mention of ECC memory when discussing protection from bitrot, especially with filesystems which depend on correctly functioning memory to provide the protections people expect from them.
People sneer at me for being a stickler for it, but between 50GB of memory and 26TB of ZFS-protected storage, I see ECC corrections about as often as I see disk checksum errors - maybe half a dozen of each in the past year or two. Frankly I think it's idiotic it's not more common and better supported.
- pedrocr 13y agoI agree and it's one of the reasons it's much easier to go with AMD for cheap home servers/NAS (HP Microserver line is great for example) as Intel tends to reserve ECC for very expensive parts. What doesn't seem easy is to get ECC in laptops. I don't think even traditionally business-focused lines have it (e.g., Thinkpads).
- stassats 13y agoNowadays you can get ECC with intel for quite cheap: http://ark.intel.com/search/advanced/?s=t&MarketSegment=DT&ECCMemory=true http://ark.intel.com/search/advanced/?s=t&MarketSegment=DT&E...
- pedrocr 13y agoThat's good to know. It seems the issue is actually motherboard support. Newegg shows a single intel motherboard with ECC and no AMD ones. I suppose the low-end server hardware (e.g., HP Microservers) just tend to be AMD so it's easiest to get a cheap AMD server with ECC than an Intel one. I had assumed it was an intel issue but apparently not. It would be nice if the intel ultrabook standard started including ECC as well. One can dream...
- Freaky 13y agoOn LGA1150 you need a C22x chipset for ECC support: http://www.intel.com/content/www/us/en/chipsets/server-chipsets/server-chipset-c222-c224-c226.html http://www.intel.com/content/www/us/en/chipsets/server-chips... Prices for these start about 3x higher than other boards, their choice is extremely limited, and availability even more so :/
- asdfs 13y agoUnfortunately while the CPUs themselves are quite reasonable, motherboards tend to be far more expensive and less available. I wonder how much of a markup Intel has on the server chipsets vs. desktop ones; it'd be interesting to know whether Intel or motherboard manufacturers are making bank on the boards.
- ChuckMcM 13y agoThe article was, sadly, missing a lot of stuff. When it said "(most arrays don't check parity by default on every read)" I knew the author was not up to the task of writing this article. FWIW, a RAID system has to calculate parity every time it reads a stripe so that it will know how to change it if something in the stripe changes. There are of course file system errors that are invisible to RAID, if your FS writes corrupted data (as it might with ECC failures) the RAID subsystem will happily compute the correct parity for the stripe. Granted I spent nearly 10 years immersed in storage systems (5 at NetApp, 4 at Google) but still there is some easily checked stuff missing from this article and so its point (which is hardening in the filesystems is good) is lost.
- dspillett 13y ago> a RAID system has to calculate parity every time it reads a stripe so that it will know how to change it if something in the stripe changes Nope. On a non-degraded array with parity (R5, 6, ...) it will only read the block it needs to from the drive it is on. For an array that mirrors without parity (R1, ...) it will read the block from one of the drives. There is no need to bother another drive with the read operation unless it needs to check parity (because it has been told to on each ready as some controllers can be so told), you actually reduce the performance benefit of the striping if you do as you may be moving the heads of the other drive(s) "unnecessarily" and potentially away from another block they were about to be asked to read. With an un-degraded array unless explicitly told to check parity on read the controller will not touch the parity blocks until a write happens at which point it will read the other relevant data blocks in order to regenerate the parity block. This is why the RAID 5 write penalty exists but there is no read penalty (in fact there is a read bonus due to striping over multiple devices). Parity blocks will, unless checking parity on every read, only ever get read if the array is in a degraded state, in which case you can only derive some data blocks by reading the other blocks in that stripe (data and parity) and working out from them what the missing block should be.
- justinsb 13y agoWouldn't you get higher read throughput (2x for R1, less for R5) by reading the parity disks as well?
- binarycrusader 13y agoAnd we can lay blame squarely on Intel and others here to a certain extent. Intel could easily offer ECC support in its consumer line of desktop (none) and laptop processors (only offers it on 3 of them: http://ark.intel.com/search/advanced/?s=t&FamilyText=4th%20Generation%20Intel%C2%AE%20Core%E2%84%A2%20i7%20Processors&ECCMemory=true http://ark.intel.com/search/advanced/?s=t&FamilyText=4th%20G...), but doesn't. To me, ECC support should be like SSL/TLS for modern applications; don't do it without it!