3 ms·
It is highly unlikely that the duplicate portions of the file will have an offset that's a multiple of 2^16 which would be required for chunks to have matching
by SethTro 5y ago
It is highly unlikely that the duplicate portions of the file will have an offset that's a multiple of 2^16 which would be required for chunks to have matching hashes. On the client side you could theoretically run lbfs over your files but on the swarm side this isn't going to happen
- Dylan16807 5y ago> It is highly unlikely that the duplicate portions of the file will have an offset that's a multiple of 2^16 which would be required for chunks to have matching hashes. That's exactly what chunking based on a rolling hash solves. You set the average size of chunks and the content controls the exact boundaries.
- gkfasdfasdf 5y agoRight, exactly. Chunk boundaries are not determined by fixed size chunks, but rather when the rolling hash matches some prefix, which means chunk sizes will vary but by controlling the prefix can set the average size of the prefix. Besides the lbfs paper, another nice writeup here: https://moinakg.wordpress.com/2013/06/22/high-performance-content-defined-chunking/ https://moinakg.wordpress.com/2013/06/22/high-performance-co...
- brokenmachine 5y agoRabin fingerprinting can do this I believe. "the idea is to select blocks not based on a specific offset but rather by some property of the block contents" https://en.wikipedia.org/wiki/Rabin_fingerprint#Applications https://en.wikipedia.org/wiki/Rabin_fingerprint#Applications