3 ms·
As a lay person I don't understand why they wouldn't, could you give some more information as to why?
by zdkl 6y ago
As a lay person I don't understand why they wouldn't, could you give some more information as to why?
- cm2187 6y agoFor these large scale infrastructures, they typically use JBOD, i.e. no RAID whatsoever and they achieve redundancy by maintaining multiple copies of the data over multiple disks spread around the datacentre. So it's not hardware RAID that requires a certain response time (I suspect the drives mentioned here drop because they timeout as it take a long time to random write on an SMR disk), more like a distributed software RAID. I think they likely also have their own file system so they are not dependent on a fixed block size, which allows them to save space if they have loads of small files.
- smueller1234 6y agoI think that's about half right: for large storage infrastructure, you would indeed eschew local RAID, but instead of storing multiple copies, you'd use Reed Solomon encoding to stripe a single copy across many disks/servers/failure domains with a configurable number of added parity stripes. Full copies are really expensive!
- cm2187 6y agoI believe azure uses full copies [1]. I assume the other cloud providers must do the same. [1] https://docs.microsoft.com/en-us/azure/storage/common/storage-redundancy https://docs.microsoft.com/en-us/azure/storage/common/storag...
- namibj 6y agoGoogle seems to use erasure coding for some (iirc. multi-AZ nearline, or so) storage classes, last I checked.
- Serow225 6y agohttps://www.backblaze.com/blog/reed-solomon/ https://www.backblaze.com/blog/reed-solomon/
- sp332 6y agoIt's a similar idea but they distribute the copies between machines, not just across disks in the same machine. https://www.backblaze.com/blog/reed-solomon/ https://www.backblaze.com/blog/reed-solomon/ So the drives in each box are not in a RAID configuration.
- klodolph 6y agoRAID only protects you from disk failures, not machine failures. If you want to protect against machine failures (whole server offline), which might be an availability concern, then you would naturally want to replicate the data at a higher level. Once you’re doing that, it makes sense to only replicate the data at a higher level, because for a given level of safety, it is more efficient to replicate the data once, in one layer, than to replicate the data multiple times, at multiple layers. RAID is only effective for workloads small enough that you care about a single machine.