4 ms·
Am I the only one who thinks RAID controllers are a placebo and wouldn't trust anything but ZFS? What was even more interesting is that our SSDs are connected
by t_tsonev 11y ago
Am I the only one who thinks RAID controllers are a placebo and wouldn't trust anything but ZFS?
What was even more interesting is that our SSDs are connected to the Dell H700/H710 RAID controller which has a battery backup unit (BBU) which should make our drives power failure resilient. RAID controller with BBU in case of a power failure can hold the cached data until the power comes back, so that it can flush it to the drives when the drives come back online.
- jcrawfordor 11y agoI'm not sure that you could say that RAID controllers are a placebo - I've been through multiple hard disk failures that either hardware or software RAID has enabled a graceful recovery from (with no data loss). As for whether or not it's superior to ZFS, though, that's a tricky question. High-end hardware RAID gets you a lot of neat features that you don't find elsewhere, but it's expensive and the RAID controller itself tends to become more and more of a point of failure (always keep spares, especially if they go out of manufacture!)
- coldtea 11y ago>Am I the only one who thinks RAID controllers are a placebo and wouldn't trust anything but ZFS? No, there are others that cargo-cult believe in ZFS too -- a hyped single vendor OS, without first-tier-support on Linux, and with its own issues, compared to a industry wide standard, used for 3 decades in the most demanding data-centers protocol and its implementations.
- andor 11y agoa hyped single vendor Who's that single vendor? FreeBSD, OpenIndiana, Joyent, Oracle? compared to a industry wide standard Good luck exchanging your fried RAID controller for a "comparable" model
- aexaey 11y ago> Good luck exchanging your fried RAID controller for a "comparable" model If you are running Linux/*BSD/SmartOS, the best RAID controller is a JBOD controller, i.e. one that that exposes all connected disks as-is to the host OS, and than it would be in-kernel soft-RAID implementing the actual logic. This approach gives you a more predictable system with no vendor-specific idiosyncrasies or extra cache level to worry about (BBU in OP's article); you end up running RAID code that is peer-reviewed, fully integrated into fsync() algorithm, and surprisingly, often gives you better performance too. "True hardware RAID controllers", on the other hand, are nothing more than an application-specific computer with its own (non-upgradeable and often outdated) CPU, RAM, I/O and hard-to-upgrade proprietary software. And if you buy into the view described above, than replacing a JBOD controller for RAID use is exactly the same thing as replacing JBOD controller for ZFS use, by definition.
- ploxiln 11y agoYou're both right and wrong. RAID controllers that you could afford for yourself, costing less than $1000, are just another point of failure. I've seen them fail more often than the disks, causing corruption as they went. However I've also seen the very expensive datacenter RAID controllers keep a whole bunch of servers up for years as their (15k rpm spinning rust) disks were failing and being swapped. SSD problems with power failure are such that ZFS won't save you. ZFS's checksumming will at least know what's corrupted and what isn't. But the extent of the corruption will be practically unlimited if the OS can't trust the drive to honor write barriers / flushes - any byte since the last powerup could still be only in volatile cache on the SSD and never make it to persistent NAND. Every filesystem depends on these occasional flushes/barriers to establish checkpoints where all previous writes are really written, including ZFS. Consider, it may have created an updated copy of the filesystem tree root node, made sure it was flushed, made another updated copy, made sure it was flushed, and then (only after the flushes) re-used the space that the original copy occupied for other data. If you can't trust that the flushes actually happened when the drive indicated they were completed, then it might be that only the last write, over the original root node copy, actually made it to the NAND.
- kokey 11y agoI currently deal with a few disk failures every week, all on RAID6 (on Dell H7xx controllers) and a few RAID-1 on HP and Adaptec controllers. Mostly SAS drives, but a few older model SATA SSD drives (that are dropping like flies after lots of writes over a year). This is across about 32000 disks so the chances for some disk failure every week is high. The RAID controllers work as advertised almost always, file system intact. Sometimes there's a performance degradation during a disk rebuild, but only on arrays where disk i/o is near max ability. There has been, I think two, catastrophic failures, for example when there's an issue with the cable or backplane, causing corruption, then having more than two disks go corrupt before getting to the bottom of it and invalidating the entire array. That said we have very few issues with power, and the batteries have died on many of the controllers, so I'm not sure how it will pan out on dodgy data centre power flapping.
- mgrennan 11y agoI worked in Dell's High complexity support for a couple of years. I saws RAID disk failures all day long. I BELIEVE SOFTWARE RAID IS YOUR FRIEND. I've seen hardware RAID system become corrupted because a miner version of the controller software was used after a failure. Except for multiple platter crashes, there has not been a software raid (linux) and Spinright there has never been a RAID I couldn't recover. So... Hardware Raid is NOT worth the extra speed.
- scurvy 11y agoYou and the author of the title are confusing cache layers and "battery protection." A RAID BBU will only protect the RAID controller's write cache; the thing that your write IO sits in until the controller flushes to the disk set. That's either 4-5 seconds depending on Dell/LSI model or until it gets pushed out for cache capacity. The above has absolutely zero to do with the drive write cache and power protection. The drive write cache is used to cache the write on the drive after the RAID controller. Spinning metal drives have caches. SSD's have caches. Whether you use the cache is up to you and your OS. The safest setting is off. You can usually get more performance by leaving it on. This is why Windows would throw warnings all over the place when enabling write caching without a connected UPS. I'm not sure exactly what the logic is on Linux for sd/sg devices. The default Dell RAID controller behavior with write caching on drives is based on drive interface. SAS = off. SATA = on. If you run non-power safe SATA drives with a Dell RAID controller, you must disable the write cache if you want data durability. Then again, you could also run your databases on power-safe, non-consumer SSD's. This whole article was basically a giant warning sign saying "run away they need serious help."