8 ms·
This seems naive, there are always false positives
by paul_f 5y ago
This seems naive, there are always false positives
- fossuser 5y agoIt’s not naive, it’s math.
- mpol 5y agoIt's not a math problem, it's a human problem. suppose you have a partner who is a 'petite' woman of 34. She enjoys posting nudies on a website, but without her face in that picture. Someone who collects child porn downloads it, because he enjoys that picture. A year later he gets caught by the police and all his pictures get marked as 'verified child porn'. Suddenly you get marked as owning child porn.
- fossuser 5y agoThat isn’t CSAM. https://www.nytimes.com/interactive/2019/09/28/us/child-sex-abuse.html https://www.nytimes.com/interactive/2019/09/28/us/child-sex-... Apple’s thing has some sort of threshold anyway so one image would not trigger it. I don’t buy your example - the CSAM images are not what you’re describing.
- farmerstan 5y agoWho determines whether something is CSAM? How do you or I know that every single one of those images actually is CSAM? How do we know if the FBI or CIA or CCP adds hashes of innocent pics in order to pin a crime against someone?
- DSingularity 5y agoNo. You can design a system where the FPR is essentially zero. Even if shitty md5 is used.
- Conlectus 5y agoPerceptual hashes are not exact hashes, otherwise they would be useless for this task; you would just mirror the image or change 1 pixel. They are instead fuzzy classifiers, and thus have non-zero error rates.
- nomel 5y agoBut this trigger requires many of these false flags to exceed the threshold (most likely why it exists). I imagine the "one in a trillion" numbers they claim for false flagging rate are probably cemented in reality, and make it trivial for human review.
- skinkestek 5y agoI've explained the problem here: https://news.ycombinator.com/item?id=28099927 https://news.ycombinator.com/item?id=28099927
- DSingularity 5y agoHow does this relate to the allegation I was responding to? There is a difference between false positives and false negatives you know. False negative is where the criminal evades the system through measures you describe. False positive — what I was responding to — is where Apple incorrectly accuses an innocent person and thus violates their privacy during the subsequent review.
- Zababa 5y ago> No. You can design a system where the FPR is essentially zero. Even if shitty md5 is used. How can you do that, considering md5 can have collisions?
- schoen 5y agoA confusing thing that I think people haven't addressed clearly for the most part: For a hash (whether cryptographic or perceptual), there is a chance of random collisions and also a difficulty factor for adversarially-created intentional collisions. The random collision probability has to be estimated based on some model of the input and output space (with cryptographic hash functions, you would usually model them as pseudorandom functions and assume that the collision probability is the same one created by the birthday paradox calculation). Intentional collisions depend on insight about the structure of the hash function, and there are also different kinds of difficulty levels depending on the nature of the attack (preimage resistance, second-preimage resistance, and collision resistance). Gaining more insight about the structure of the hash function can act to reduce the work factor required for mounting these attacks. That should be true for perceptual hashes just as much as cryptographic hashes, but presumably all of the intentional attacks should start off easier because the perceptual hashes' threat models are weaker and there's much less mathematical research on how to achieve them. And in AI systems involving classifiers, it was generally easy for people to create adversarial examples given access to the model. Perceptual hashes for estimating similarity to specific known images aren't the exact same thing because it's less like "how much like a cat is this image?" and more like "how much like NCMEC corpus image 77 is this image?", but maybe some of the same techniques would still work. In the cryptographic hash analogy, I guess that would be like trying to break preimage resistance.