3 ms·
I get the use case, but in most cases (and particularly this one) I'm sure it would be much better to implement that client-side. You may have seen in the WARC
by uniqueuid 2y ago
I get the use case, but in most cases (and particularly this one) I'm sure it would be much better to implement that client-side.
You may have seen in the WARC standard that they already do de-duplication based on hashes and use pointers after the first store. So this is exactly a case where FS-level dedup is not all that good.
- nikisweeting 2y agoWARC only does deduping within a single WARC, I'm talking about deduping across millions of WARCs.
- uniqueuid 2y agoThat's not true, you commonly have CDX index files which allow for de-duplication across arbitrarily large archives. The internet archive could not reasonably operate without this level of abstraction. [edit] Should add a link, this is a pretty good overview, but you can also look at implementations such as the new zeno crawler. https://support.archive-it.org/hc/en-us/articles/208001016-About-data-de-duplication https://support.archive-it.org/hc/en-us/articles/208001016-A...
- nikisweeting 2y agoAh cool, TIL, thanks for the link. I didn't realize that was possible. I know of the CDX index files produced by some tools but don't know anything about the details/that they could be used to dedup across WARCs, I've only been referencing the WARC file specs via IIPC's old standards docs.