3 ms·
I agree with the article in general except part of the final conclusion > The simple fact that image data is reduced to a small number of bits leads to collisi
by ris 5y ago
I agree with the article in general except part of the final conclusion
> The simple fact that image data is reduced to a small number of bits leads to collisions and therefore false positives
Our experience with regular hashes suggests this is not the underlying problem. SHA256 hashes have 256 bits and still there are no known collisions, even with people deliberately trying to find them. SHA-1 only has only 160 bits to play with and it's still hard enough to find collisions. MD5 is easier to find collisions but at 128 bits, still people don't come across them by chance.
I think the actual issue is that perceptual hashes tend to be used with this "nearest neighbour" comparison scheme which is clearly needed to compensate for the inexactness of the whole problem.
- dogma1138 5y agoThis isn’t due to the entropy of the hash but due to the entropy of the source data. These algos work by limiting the color space of the photo, usually to only black and white (not even grey scale) resizing it to a fraction of its original size and then chopping it into tiles using a fixed size grid. This increases the chances of collisions greatly because photos with a similar composition are likely to match on a sufficient number of tiles to flag the photo as a match. This is why the women image was matched to the butterfly image, if you turn the image to B&W resize it to something like 256x256 pixels and divide it into a grid of say 16 tiles all of a sudden a lot of these tiles can match.
- giantrobot 5y agoPerceptual hashes don't involve diffusion and confusion steps like cryptographic hashes. Perceptual hashes don't want decorrelation like cryptographic hashes. In fact they want similar but not identical images to end up with similar hash values.