4 ms·
While the precedent this sets is indeed concerning, the specific hypotheticals this article give are nonsensical. > an adversary could trick Apple’s algorithm
by shadowfacts 5y ago
While the precedent this sets is indeed concerning, the specific hypotheticals this article give are nonsensical.
> an adversary could trick Apple’s algorithm into erroneously matching an existing image
In which case the malicious, adversary-controlled images are sent to Apple. After which—the implication is—they can be re-obtained by... the adversary that created them. So what?
An adversary could conceivably lower the reporting threshold by getting the victim to save a bunch of false-positive images. Again, so what? Surely if the adversary has reason to believe there are some number of CSAM images on a user's phone, there are more direct ways of going after them.
> These kinds of false positives could happen if the matching database has been tampered with or expanded to include images that do not depict child abuse
An adversary would either have to:
A) carry out a supply-chain attack on Apple,
B) ship different iOS images in different countries, or
C) insert entries into the database on a specific users phone.
Options A and C are irrelevant: if the phone is compromised, the CSAM database being modified is the least of your concerns. And option B is independently verifiable (granted, Apple does not do enough to make third-party auditing of iOS easy, but it is possible).
- kemayo 5y ago> An adversary would either have to: More importantly, they'd need to suborn Apple's human-review process, because the people doing the review would need to know what not-CSAM they're looking for. Could Apple be coerced into (or willingly) do this? I have no idea. But it's a very different threat-model than the article suggests, that boils down to the "do you trust your OS vendor?" question.
- intricatedetail 5y agoHow a human looking at low res CSAM matching collision picture that looks "innocent" will be able to tell for sure it is a false positive? Can they know with certainty that it is a false positive and not a manipulated real image? It seems to me that they would have to report these anyway.
- kemayo 5y agoIs your argument "a human review step is fundamentally a rubber-stamp that won't reject a false-positive?" Because I don't personally think that's how it'd work out, but I'll acknowledge that I might just be an optimist. (I mean, assuming that you need ~30 matches to trigger the review phase of the process, I'd think it'd be weird to a reviewer looking for child porn if you got 30 pictures of apparently-random political documents or subversive memes.)
- intricatedetail 5y agoI am trying to say it's not possible to tell with certainty from a lowres picture that you are looking at false positive. For example low contrast CSAM imposed on a document could trigger NeuralHash match but the lowres image will look like a false positive.
- kemayo 5y agoFor your example, wouldn't that only work to make the original source image that's polluting the CSAM database look like CSAM in lowres? The actual document-image the oppressive government is looking for that'd trigger the match wouldn't have the CSAM included. That said, I do think it'd be nice to have a better demonstration of exactly what this "derivative" the reviewers would be looking at is. There's a lot of variations there, balancing false-positive privacy concerns, the mental health of the reviewers, potential downsampling issues, etc.
- simondotau 5y agoI agree, it would be useful if Apple could be clearer by what they mean by a derivative. I recall reading somewhere that it's a reduced resolution, grayscale copy of the image. I can't vouch for that, but that would be a plausible notion of what the "derivative" would be. Personally I would also be placing a hard watermark in the middle of the image, or maybe some hard slashes randomly through the image, so that "clean" images cannot leak out of human review. Let's imagine that the derivative is a 0.5 megapixel, grayscale, watermarked, HEIC-compressed copy of the original image. This would be plenty to determine with zero ambiguity that the image is actually "A1" classified, i.e. depicts a prepubescent minor ("A") engaged in a sex act ("1").
- bitwise-evan 5y ago> an adversary could trick Apple’s algorithm into erroneously matching an existing image This is a very real, possible attack. Apple ships its CSAM model on device so any attacker can have a copy of the model. Then the attacker creates an image that triggers CSAM but looks like a panda [1]. Now the attacker sends tons of triggering photos to the unsuspecting victim, who now gets questioned by the FBI. 1: https://medium.com/@ml.at.berkeley/tricking-neural-networks-create-your-own-adversarial-examples-a61eb7620fd8 https://medium.com/@ml.at.berkeley/tricking-neural-networks-...
- carom 5y agoSo the attacker creates an image then the user has to download it. Then the FBI digs in and see it was a crafted false positive, then begin to investigate who sent it and why. Then the user takes civil action against the person who sent it for harassment.
- simondotau 5y agoMore precisely, 30 carefully crafted false positives. All of which need to be imported into your iPhone's photo library to sit alongside pictures of your dog and your mum. And then they have to get past human review. Not impossible, but so far beyond implausible that it can be dismissed as ridiculous. And if this trick ever works, it could only be done once before Apple has the opportunity to plug holes in their NeuralHash algorithm and fix any deficiencies in the manual review process.
- shadowfacts 5y ago> Now the attacker sends tons of triggering photos to the unsuspecting victim, who now gets questioned by the FBI. That's glossing over the middle part where a human from Apple (before it even gets to law enforcement) actually look at the images and goes "oh, these are actually pandas" and realizes they were erroneously detected.