3 ms·
Disclaimer: I'm the author of the blog post. :-) > some of the assumptions that Brian was tossing out are not good ones to make. We COMPLETELY welcome other
by brianwski 8y ago
Disclaimer: I'm the author of the blog post. :-)
> some of the assumptions that Brian was tossing out are not good ones to make.
We COMPLETELY welcome other analysis and listing other assumptions. Internally, we argued endlessly about why this or that wasn't totally accurate, and finally decided to publish the math WITH all of our assumptions exposed so you could be the judge. If Amazon wants to publish their assumptions for S3 for comparison, we're all ears.
> that Drive Savers will exist as a company 10 years from now
Absolutely true, this calculation is only good RIGHT NOW. For example, one of the things that came up internally was "well, when drives get more dense the rebuild time rises, so this calculation will no longer be accurate in two years". But at the same time, we have some additional tricks and optimizations to make which we have not done yet to cut the 6 day average drive rebuild time down to 3 days. Also, drives last us about 5 years, so your data will be migrated to totally new drives 5 years from now. Those drives will absolutely have a different drive failure rate (maybe higher, maybe lower) so the calculation will no longer have the same result 5 years from now.
- ChuckMcM 8y agoHey Brian! For the record I love that you are transparent about your assumptions, it is really helpful. I have been in very very similar debates, both at NetApp and at Google of all places. Peter Corbett, the guy who invented the dual parity scheme that NetApp uses, wrote a similar analysis as well for Fast '04 [1]. As someone who likes to geek out on failure proof systems and perfectly secure systems, neither of which are attainable but can be asymptotically approached, I think you are seeing the "there is always one level deeper" kinds of discussions. Personally I think of them as endorsements because if the exceptions get too extreme (say 'what if an asteroid hits?') then you know you've got all the bases covered. Its only a problem if the person analyzing the analysis finds something that you really did not even consider. Then it opens up an opportunity to look at the problem a whole new way. [1] https://www.usenix.org/legacy/publications/library/proceedings/fast04/tech/corbett/corbett_html/ https://www.usenix.org/legacy/publications/library/proceedin...
- brianwski 8y agoGood link! I've both forwarded it on, and will study it when I have some time later.
- jorangreef 8y agoHi Brian, thanks for open-sourcing Backblaze's JavaReedSolomon, which is really well-written. A few months ago I ran into an issue with Reed Solomon coding throughput not saturating the write throughput of 16 drives, and wrote a new Reed Solomon module based on Cauchy matrices: https://github.com/ronomon/reed-solomon https://github.com/ronomon/reed-solomon The Cauchy matrices remove the need for a table lookup to do the Galois multiply, replacing it with pure XOR. Together with other optimizations, this gives nearly 3x-5x more coding throughput for the same (17,3) parameters, assuming you're still using your open-sourced JavaReedSolomon in production. I don't know if Reed Solomon coding throughput is a factor in your rebuild times?
- kdkeyser 8y agoThere is also Intel's ISA-L: https://github.com/01org/isa-l https://github.com/01org/isa-l It contains optimized Galois Field multiplication, resulting in Reed Solomon (both Vandermonde as well as Cauchy) at multiple GB/s on a modern x86 CPU. Still, I doubt that the Reed Solomon coding speed is the limiting factor in their rebuild time. There is a mention of a 6-day duration, so even with a very slow Reed Solomon implementation ( ~ 100 MB/s) that should not be a bottleneck for a 10 TB drive rebuild (assuming a distributed rebuild approach, not a traditional RAID style rebuild).
- voidmain 8y ago> If Amazon wants to publish their assumptions for S3 for comparison, we're all ears. The 2010 S3 calculation is obviously also wrong! I totally feel your pain in terms of wanting to have a directly comparable answer, but the reasons you give that "it doesn't matter" (and others, including correlated faults, software bugs, model risk, and security breaches) are actually reasons why the stated durability number is wrong. Honestly IMO a probability of 1-10^-11 is the wrong answer to pretty much any question; model risk is going to dominate that for any problem more complex than 1+1=2. That said, although neither your system nor Amazon's should be expected to have anywhere near "eleven nines" of durability in reality, if as I understand it S3 is split across availability zones in a region and your product has all splits in a monolithic DC, I would expect S3 to come out ahead in a more careful analysis. (But note that S3 is not really a seamless multi region product, though there is an option to set up cross region replication.)
- pas 8y agoAs far as I know, if they claim multi AZ in a city (region), that really means at least separate buildings, but likely separate locations withing the region too. Though I'd welcome hard data on this very much.