5 ms·
The Archive currently has about 46 Petabytes of content ("bytes archived"), and over 120 PB of raw disk capacity; the difference is due to data replication, "cu
by bnewbold 8y ago
The Archive currently has about 46 Petabytes of content ("bytes archived"), and over 120 PB of raw disk capacity; the difference is due to data replication, "currently filling" storage, non-storage infrastructure, etc.
We save a lot on web content storage by de-duplicating "revists" when the page hasn't changed. This works out to save a whole lot for content like jQuery served from a common CDN URL; it doesn't work well when there is a page counter or any trivial changing content on a page.
If you are interested in the storage back-end, it's actually pretty simple: HTTP requests/responses are concatenated and compressed in WARC files (sort of like .tar.gz) that get stored on regular old ext4 filesystems. An index of "what URL captures are in what WARC files on what servers" is continuously generated in the form of, basically, a giant sorted (and shareded) .tsv file; replay requests on web.archive.org look up the URL and timestamp and get a reference to a machine, file, and file offset, and make an HTTP 1.1 range request for the content in question. There are a bunch of other details, like checking robots.txt status, but the core design is super simple, cheap, and (relatively) easy to operate at scale.
Apart from web crawl content (including, these days, "heavy" video content which is difficult to de-dupe), we have a large amount of live recorded TV, scanned books (raw photos), etc.
(I currently work at IA)
- jvz 8y agoHave you evaluated compression algorithms that support custom dictionaries, like zstd? You could generate a compression dictionary for each domain, or just for those above a certain size.
- Nition 8y agoWould be nice if you could store a diff, so a changing counter would only have to store the changed counter after the initial save.
- murukesh_s 8y agoWondering the same. How if used git protocol itself? Not sure how efficient is git, but if it is then it's a relatively easier change
- zaarn 8y agoYou can do this if you group all site visits into one common WARC and compress it (or dedup it otherwise). WARC itself does have a method of dedup if the response is the same (or mostly the same) but terrible if content changes.
- __mp 8y agoHow do you combat bit-rot/silent corruption?
- dylan604 8y ago> (including, these days, "heavy" video content which is difficult to de-dupe) Still waiting for the Shazam for video to know that 2 videos are the same even when they are of different codecs/framesizes/etc; just based on the visual imagery.
- nitrogen 8y agoIsn't that Youtube Content ID? Upload a video and see who sends you a takedown
- nojvek 8y ago120PB is still an insane amount of data. I imagine at least a million dollars an year to keep the lights on just for the infrastructure. Does IA have a big endowment to keep it going for a while?