4 ms·
> RAID or storage replication in distributed storage <..> is not only useless, but actively undesirable I guess I'm different from most people, good news! When
by innot 5y ago
> RAID or storage replication in distributed storage <..> is not only useless, but actively undesirable
I guess I'm different from most people, good news!
When building my new "home server" half a year ago I made a raid-1 (based on ZFS) with 4 NVMEs. I rarely appear at that city, so I brought the fifth one and put it into an empty slot. Well, one of the 4 nvmes lasted for 3 months and stopped responding. One "zpool replace" and I'm back to normal, without any downtime, disassembly, even reboots. I think that's quite useful. When I'm there the next time I'll replace the dead one, of course.
- zxexz 5y agoWhat setup do you use to put 4 NVME in one box? I know it’s possible, I’ve just heard off so many different setups. I know there are some PCIE cards that allow for 4 NVME drives. But you have to match that with a motherboard/CPU combo with enough ones to not lose bandwidth.
- birdyrooster 5y agoThat’s exactly what they are doing. Anyone else is using proprietary controllers and ports for a server chassis
- pdpi 5y agoI've been looking into building some small/cheap storage, and this is one of the enclosures I've been looking at. https://www.owcdigital.com/products/express-4m2 https://www.owcdigital.com/products/express-4m2
- isotopp 5y agoFor distributed storage, we use this: https://www.slideshare.net/Storage-Forum/operation-unthinkable-software-defined-storage-bookingcom-peter-buschman#14 https://www.slideshare.net/Storage-Forum/operation-unthinkab... We then install SDS software, Cloudian for S3, Quobyte for File, and we used to use Datera for iSCSI. Lightbits maybe in the future, I don't know. These boxen get purchased with 4 NVME devices, but can grow to 24 NVME devices. Currently 11 TB Microns, going for 16 or more in the future. For local storage, multiple NVME hardly ever make sense.
- birdyrooster 5y agoDoes zpool not automatically promote the hot spare like mdadm?
- seized 5y agoIt can, if you set a disk as the hot spare for that pool. But a disk can only be a hot spare in one pool, so to have a "global" hot spare it has to be done manually. That may be what that poster was doing.
- sjagoe 5y agoAlso, if I understand it correctly, there are a few other caveats with hot spares: It will only activate when another drive completely fails, so you can't decide to replace a drive when it's close to failure (probably not an issue in this case, though, with the unresponsive drive). Second, with the hot spare activated, the pool is still degraded, and the original drive still needs to be replaced; then the hot spare is removed from the vdev, and goes back to being a hot spare. It's these reasons that I've decided to just keep a couple of cold spares ready that I can swap in to my system as needed, although I do have access to the NAS at any time. If I was remote like GP, I might decide to use a hot spare.
- SkyMarshal 5y agoI’ve recently converted all my home workstation and NAS hard drives over to OpenZFS, and it’s amazing. Anyone who says RAID is useless or undesirable just hasn’t used ZFS yet.
- amarshall 5y agoThe article’s author only said RAID was useless in a specific scenario, not generally, and the post you’re replying to omitted this crucial context.
- toast0 5y agoThis article is speaking of large scale multinode distributed systems. Hundreds of rack sized systems. In those systems, you often don't need explicit disk redundancy, because you have data redundancy across nodes with independent disks. This is a good insight, but you need to be sure the disks are independent.
- merb 5y agowell most often hba's and raid controllers are another thing which increases latency and makes maintenances costs go up quite a bit (more stuff to update) and also it's another part that can break. that's why it's not recommended when running ceph.
- Aea 5y agoI'm pretty sure discrete HBAs / Hardware RAID Controllers have effectively gone the way of the dodo. Software RAID (or ZFS) is the common, faster, cheaper, more reliable way of doing things.
- amarshall 5y agoDon’t lop HBAs and RAID controllers together. The former is just PCIe to SATA or SCSI or whatever (otherwise it is not just an HBA, but indeed a RAID controller). Such a thing is still useful and perhaps necessary for software RAID if there are insufficient ports on the motherboard.
- karmakaze 5y agoHardware caching raid controllers do have the advantage if power is lost, the cache can still be written out without the CPU/software to do it. This let's you safely run without write-thru cache fsync. This was a common spec for provisioned bare-metal MySQL servers I'd worked with.
- amarshall 5y agoSometimes. Other times they may make things worse by lying to the filesystem (and thereby also the application) about writes being completed, which may confound higher-level consistency models.
- amarshall 5y agoYou omitted the context from the rest of the sentence: > most database-like applications do their redundancy themselves, at the application level … If that’s not the case for your storage (doesn’t sound like it), then the author’s point doesn’t apply to your case anyway. In which case, yes, RAID may be useful.
- twotwotwo 5y ago> > RAID or storage replication in distributed storage <..> is not only useless, but actively undesirable > I guess I'm different from most people, good news! The earlier part of the sentence helps explain the difference: "That is, because most database-like applications do their redundancy themselves, at the application level..." Running one box I'd want RAID on it for sure. Work already runs a DB cluster because the app needs to stay up when an entire box goes away. Once you have 3+ hot copies of the data and a failover setup, RAID within each box on top of that can be extravagant. (If you do want greater reliability, it might be through more replicas, etc. instead of RAID.) There is a bit of universalization in how the blog post phrases it. As applied to databases, though, I get where they're coming from.
- linsomniac 5y agoWe are currently converting our SSD-based Ganeti clusters from LVM on RAID to ZFS, to prepare for our NVMe future, without RAID cards (1). Was hoping to get the second box in our dev ganeti cluster reinstalled this morning to do further testing, but the first box has been working great! 1: LSI has a NVMe RAID controller for U.2 chassis, preparing for a non-RAID future, just in case.
- goodpoint 5y agoCompare your solution with having 4 SBCs, with 1 NVME each, at different locations. The network client would handle replication and checksumming. The total cost might be similar but you have increased reliability over SBC/controller/uplink failure. Of course there are tradeoffs on performance and ease of management...
- Godel_unicode 5y agoYou think that building 4 systems in 4 locations is likely to have a similar cost to one system at one location? For small systems, the fixed costs are a significant portion of the overall system cost. This is doubly true for physical or self-hosted systems.
- isotopp 5y agoMy environment is not a home environment. It looks like this: https://blog.koehntopp.info/2021/03/24/a-lot-of-mysql.html https://blog.koehntopp.info/2021/03/24/a-lot-of-mysql.html