6 ms·
No. You can design a system where the FPR is essentially zero. Even if shitty md5 is used.
by DSingularity 5y ago
No. You can design a system where the FPR is essentially zero. Even if shitty md5 is used.
- Conlectus 5y agoPerceptual hashes are not exact hashes, otherwise they would be useless for this task; you would just mirror the image or change 1 pixel. They are instead fuzzy classifiers, and thus have non-zero error rates.
- nomel 5y agoBut this trigger requires many of these false flags to exceed the threshold (most likely why it exists). I imagine the "one in a trillion" numbers they claim for false flagging rate are probably cemented in reality, and make it trivial for human review.
- skinkestek 5y agoI've explained the problem here: https://news.ycombinator.com/item?id=28099927 https://news.ycombinator.com/item?id=28099927
- DSingularity 5y agoHow does this relate to the allegation I was responding to? There is a difference between false positives and false negatives you know. False negative is where the criminal evades the system through measures you describe. False positive — what I was responding to — is where Apple incorrectly accuses an innocent person and thus violates their privacy during the subsequent review.
- Zababa 5y ago> No. You can design a system where the FPR is essentially zero. Even if shitty md5 is used. How can you do that, considering md5 can have collisions?
- schoen 5y agoA confusing thing that I think people haven't addressed clearly for the most part: For a hash (whether cryptographic or perceptual), there is a chance of random collisions and also a difficulty factor for adversarially-created intentional collisions. The random collision probability has to be estimated based on some model of the input and output space (with cryptographic hash functions, you would usually model them as pseudorandom functions and assume that the collision probability is the same one created by the birthday paradox calculation). Intentional collisions depend on insight about the structure of the hash function, and there are also different kinds of difficulty levels depending on the nature of the attack (preimage resistance, second-preimage resistance, and collision resistance). Gaining more insight about the structure of the hash function can act to reduce the work factor required for mounting these attacks. That should be true for perceptual hashes just as much as cryptographic hashes, but presumably all of the intentional attacks should start off easier because the perceptual hashes' threat models are weaker and there's much less mathematical research on how to achieve them. And in AI systems involving classifiers, it was generally easy for people to create adversarial examples given access to the model. Perceptual hashes for estimating similarity to specific known images aren't the exact same thing because it's less like "how much like a cat is this image?" and more like "how much like NCMEC corpus image 77 is this image?", but maybe some of the same techniques would still work. In the cryptographic hash analogy, I guess that would be like trying to break preimage resistance.
- Zababa 5y agoThanks for the detailed explanation. I understand why that works for perceptual hashes if you make them really precise, however I doubt it would work with md5, which is why I asked.
- DSingularity 5y agoThe discussion I thought we were having was about false positives and not adversarially induced false positives. For the former the random collisions have a probability of 1/(2^64). To mitigate adversarial false positives one idea is to use the combination of a cryptographically strong hash along with a randomly selected perturbation of the file. Prior to hashing, perturb the file and submit both the hash and the selected perturbation to apple. Apple selects the DB based on the perturbation and proceeds with matching and thresholding.