4 ms·
TLDR: 22 petabytes, stored redundantly as 44 petabytes, and a warehouse of physical media (books, vinyl records) that grows by a 'shipping container' every two
by codeulike 8y ago
TLDR: 22 petabytes, stored redundantly as 44 petabytes, and a warehouse of physical media (books, vinyl records) that grows by a 'shipping container' every two weeks.
edit: Also this article mentions a bunch of things I didn't know about (e.g. playable 80s video game archive) so its worth reading.
- exikyut 8y agoActually approx 100 PB of raw capacity thanks to redundancy, I was told once. Unsure if this is on RAID or ZFS.
- ddorian43 8y agoI was expecting some kind of reed-solomon encoding.
- patrickg_zill 8y agoZfs does have Reed-Solomon in the double and triple parity modes of operation.
- romed 8y agoI guess I would expect something more like massively distributed erasure encodings rather than ZFS directly. For example something like Dropbox’s magic pocket but taken to extremes with, say, 100 data stripes and 20 parity stripes, all on different machines.
- klodolph 8y agoOne of the known problems with these large encodings is that the cost of reconstructing lost data increases as the size of the data increases. If you use Reed-Solomon (100,20), then you only have 20% overhead and have a vanishingly small probability of losing data, but if you lose a single 10TB disk, you need to do 1PB of I/O to rebuild it! Even forgetting I/O for a moment, you might be churning through a bunch of CPU time just to rebuild a single block of data. Of course, you don't need to rebuild immediately. You're effectively working with (100,19), and you can put off the reconstruction as long as you like, maybe you don't reconstruct until someone wants to read the data or until enough other disks fail, and you can prioritize the I/O as low as you like. But in practice, super large encodings become more and more expensive as the size increases.
- romed 8y agoYep, the degraded read cost increases a lot! But it only matters if you are losing disks at a rate similar to the time it takes you to do the 1PB of i/o, and also depends on how often you might need to serve from the data in question. I imagine that the Archive's data is tremendously cold, but I could be wrong. You can make a sensible tradeoff for any given use case. For example you might get away with a 12,8 orthogonal nested encoding, that "only" has 67% space overhead and you will be able to rebuild a single lost member without reading the entire stripe.
- Dylan16807 8y agoLet's see. At 1-3% annual failure rate, we expect to need to rebuild a couple drives in each array per year. To make the math simpler, let's have each server send and receive 8.5TB and do 1/120 of the parity math. Since we have plenty of redundancy, let's keep things low-priority up to 3 drive failures, and try to rebuild each drive in 90 days. For bonus points, if two drives are rebuilding at once the increase in bandwidth is negligible. 8.5TB in 90 days is less than 10mbps. That means we could build servers with 50 drives and a single gigabit connection and if they were rebuilding every array at the same time it wouldn't even use half that bandwidth. Real servers are going to have vastly faster connections and probably a lot fewer drives, so honestly that petabyte of I/O is not a big deal in context. In practice you could replace failed drives in a day, and the limiting factor is the speed of a single drive, not the network.
- klodolph 8y ago> ...8.5TB in 90 days is less than 10mbps. The arithmetic is correct but that's not the correct value. The I/O necessary for reconstructing one 8.5TB drive in a (100,20) Reed-Solomon group is 850 TB, which comes out to 875 Mbit/s, averaged over 90 days. It's not uncommon to see data centers with 10 Gbit/s connections, and sure, maybe it's a 10 Gbit/s per link on a Clos fabric but it's hard to claim that this bandwidth usage is trivial. The point is not that the rebuild is prohibitively expensive or impossible, the point is that as the group size increases, the cost of data reconstruction increases and the cost of the encoding overhead decreases. At some point the cost savings from reduced encoding overhead are smaller than the additional I/O costs incurred by reconstruction. So the ideal encoding size is not as large as possible, but some medium size which balances the cost of the encoding overhead with the cost of reconstruction. And consider that if any of this data is being served, you incur the 100x I/O penalty immediately. > Real servers are going to have vastly faster connections and probably a lot fewer drives, so honestly that petabyte of I/O is not a big deal in context. In practice you could replace failed drives in a day, and the limiting factor is the speed of a single drive, not the network. The Internet Archive has 24 disks per machine, I believe. https://en.wikipedia.org/wiki/PetaBox https://en.wikipedia.org/wiki/PetaBox > In practice you could replace failed drives in a day, and the limiting factor is the speed of a single drive, not the network. To rebuild an 8.5TB drive in 1 day requires 78 Gbit/s of bandwidth. Even inside a data center, oof.