27 ms·
I manage a server fleet with local storage ranging from single-disk systems, to servers with 12+ internal disks and multiple attached JBODs, with disk tech incl
by theevilsharpie 6y ago
I manage a server fleet with local storage ranging from single-disk systems, to servers with 12+ internal disks and multiple attached JBODs, with disk tech including SATA and SAS HDDs of various RPMs, as well as SATA SSDs. Many of these machines are equipped with hardware RAID controllers of various brands. Hopefully that helps to demonstrate that I do actually have experience with enterprise hardware, and I'm not just some kid with a home lab parroting what I've read online.
So, here goes...
The management tooling for hardware RAID is a pain to deal with. Dell's OpenManage tools aren't that bad (other than being stupidly bloated), but every time I have to use LSI's MegaCLI utility, I end up wanting to shoot myself. (Seriously, do a Google search for "MegaCLI cheat sheet". It's _bad_.) For the cheaper LSI FusionMPT-based cards, I've completely given up trying to manage them from the OS, because I've never been able to get it to work reliably. And because hardware RAID controller manufacturers have changed ownership so many times over the past decade, even _obtaining_ the software can be a challenge. And that's assuming that it supports your OS, which isn't guaranteed.
Monitoring the health of hardware RAID arrays is challenging. Not only do you generally need to use a proprietary tool of varying quality to do so (with all the issues stated above), the information you get is typically not that helpful. Most RAID controllers will at least tell you if the array is degraded or not (i.e., one or more disks are missing), but the quantity and quality of that information can vary. Additionally, monitoring the health of the underlying physical disks that make up the array is also important, but getting this information from a hardware RAID controller is tedious (and sometimes impossible), which can leave you blind to the actual health of your storage. The disk health typically reported by RAID controllers (the greem/amber blinky light that indicates a failed or failing disk) is based entirely on the disks reported SMART health check status, which is so prone to false negatives that it may as well not even exist. Oh, and if you're using an OEM RAID controller with third-party disks, you often have to deal with a nag message informing you that the disks aren't "official," which at best is annoying, and at worst can hide legitimate disk issues.
Hardware RAID can have annoying limitations. Want to mix a SAS and SATA disk in the same array? Most controllers won't support that. Want to mix RAID levels on the same disks, like in the case of a distributed database where you might want some redundancy for the OS volume, but the actual database volume can be RAID 0 for performance/space efficiency? Most controllers don't support that, either. Want to create multiple distinct logical volumes from a single array (e.g., a 20 GB volume for OS, and a 10+ TB volume for data)? That's usually possible, but in the case of many LSI MegaRAID controllers (including their Dell PERC rebrands), that can prevent you from expanding the space on the array without a rebuild. Want things like RAID 6, more advanced nested RAID levels like RAID 50 or 60, or tiered storage? If the controller supports it at all, it's probably a premium-feature that costs extra to unlock.
Hardware RAID is slow. The onboard processors on most controllers have no problem keeping up with mechanical disks, but struggle with the performance potential of SSDs. The centralized cache on the controller can be useful in some situations, but the typical configuration is to use this cache while disabling the cache located on the disks themselves, which runs into scalability issues for obvious reasons. And lastly, you're ultimately limited by the throughput of the bus to which the controller's attached, which means that a hardware PCIe RAID controller can't possibly compete with NVMe disks that are directly attached to the processor via the PCIe bus.
Software RAID (at least on Linux) is fast, simple, easy to use, the UX is consistent no matter the underlying hardware, it's well integrated with many existing monitoring tools, and frankly, it's been more reliable than hardware RAID. All of our newer machines use software RAID with either NVMe disks or SAS HBAs, and as our older machines are being repurposed, they are being converted to software RAID where possible. It's made my life a lot easier, and nothing of value has been lost.
Hardware RAID is dead. The only reason I'd ever use hardware RAID again is if I had to use an OS which didn't have decent software RAID support.
- rconti 6y agoI'll agree with all of this, I've had many of the same pain points over, whatever, 15 years of doing the same. The UI on these controllers is terrible and often impenetrable. I remember whatever HP used for 'hardware raid' on DL380g6/g7s was prone to locking up+crashing under heavy load. But, again, this is very low end to middle end 'hardware raid'. It's single-digit disk stuff. It's a cheapo PCI card in a cheapo server.
- vardump 6y ago> But, again, this is very low end to middle end 'hardware raid'. It's single-digit disk stuff. It's a cheapo PCI card in a cheapo server. You can multiplex hundreds of disks across multiple SAS backplanes. And, yeah, those software RAID [0] systems are really cheap. :-) Software RAID is deployed in low end systems, sure. But also in very high end systems where compromises can't be made. And anywhere in between. [0]: For example, see: https://www.oracle.com/storage/nas/zs7-2/ https://www.oracle.com/storage/nas/zs7-2/