6 ms·
Now that the model is known, I wonder how hard it is to create "adversarial collisions": given an image and a target hash, perturb the image in a way that is ba
by y7 5y ago
Now that the model is known, I wonder how hard it is to create "adversarial collisions": given an image and a target hash, perturb the image in a way that is barely perceptible for a human, so that it matches the target hash.
- xucheng 5y agoSee https://news.ycombinator.com/item?id=28105849 https://news.ycombinator.com/item?id=28105849, which shows a POC to generate adversarial collisions for any neural network based perceptual hash scheme. The reason it works is because "(the network) is continuous(ly differentiable) and vulnerable to (gradient-)optimisation based attack".
- misterdata 5y agoApparently not that hard: https://github.com/AsuharietYgvar/AppleNeuralHash2ONNX/issues/1 https://github.com/AsuharietYgvar/AppleNeuralHash2ONNX/issue.... See also here: https://gist.github.com/unrealwill/c480371c3a4bf3abb29856c29197c0be https://gist.github.com/unrealwill/c480371c3a4bf3abb29856c29...
- Spivak 5y agoWhile super impressive I haven't seen the thing that would actually destroy the algorithm which is given a hash and random reference image produce a new image which has that hash and looks like the reference image.
- SV_BubbleTime 5y agoWhat you've seen is worse than that. All you need to do to cause trouble right now, would be to get a bad image, hash it yourself, make a collision and distribute that. Let's say for the time being that the list hashes themselves will be server-side. You won't ever get that list, but you don't need it in order to cause a collision. You would need your own supply of CSAM to hash yourself, which while distasteful is clearly also not impossible.
- dsign 5y agoSo, these adversarial collisions are the images that I need to send to my enemies so that they go to prison when they upload those images to iCloud? It seems trivially easy to exploit.
- xucheng 5y agoYou can technically hide an adversarial collision inside a complete legit normal image. It won’t be seen by human eyes but it will trigger a detection. In addition, you can do the complete opposite by perturbing a CSAM to output complete different hash to circumvent the detection. All of these vulnerabilities are well known for perceptual hash.
- SV_BubbleTime 5y agoSo right there seems to be an issue to me. It seems like if you were trading in CSAM, you would run CLEANER -all on anything and everything. Because you know someone has already written that as proof of concept here.
- nsizx 5y agoThe images are sent for manual verification before you're turned in, so no.
- sneak 5y agoYou can also just send your enemies CSAM, it is more effective at imprisoning them.
- eknkc 5y agoThey probably would not hold onto them and also now you have a paper trail of sending CSAM to people. But if you were to alter some innocent looking photos and send them to someone, they might store those.
- zepto 5y agoYes, but that won’t work since innocent looking photos won’t match the visual derivative.
- shuckles 5y agoIt might be useful to read the threat model document. Associated data from client neural hash matches are compared with the known CSAM database again on the server using a private perceptual hash before being forwarded to human reviewers, so all such an attack would do is expose non-private image derivatives to Apple. It would likely not put an account at risk for referral to NCMEC. In this sense, privacy is indeed preserved versus other server scanning solutions where an adversarial perceptual hash (PhotoDNA is effectively public as well, per an article shared on HN) would trigger human review of all the account’s data.
- nlitened 5y ago> before being forwarded to human reviewers Does that mean that Apple employs people who manually review images known to be child pornography 9-to-5? Is it legal?
- shuckles 5y agoYes, and so does every other major cloud service provider. The working conditions of these people are notoriously difficult and should be the subject of attention.
- deleted 5y ago[deleted]
- kemayo 5y agoI'd imagine the key is that it's "manually review images suspected to be child pornography". The point of the review process is presumably that there are possible hash-collisions / false-positives, so the reviewers are what cause that transition from suspected to known.
- y7 5y agoI assume these "non-private image derivatives" are downscaled versions of the original image. But for downscaling there also are adversarial techniques: perturb an image such that the downscaled version looks like a target image. See https://bdtechtalks.com/2020/08/03/machine-learning-adversarial-image-scaling/ https://bdtechtalks.com/2020/08/03/machine-learning-adversar...
- tvirosi 5y agoOr for criminals to generate perceptually similar illegal images that are no longer triggered as a 'bad' hash.
- visarga 5y agoSpy agencies could add CSAM images adversarially modified to match legit content they want to find. Then they need to have someone in Apple's team to intercept the reports. This way they can scan for any image.
- madeofpalk 5y agoIf this is the level that a "spy agency" is going to get involved, they would already skip all this BS and just upload the images directly to themselves.
- m-p-3 5y agoThat reminded me of the AT&T Room 641A https://en.wikipedia.org/wiki/Room_641A https://en.wikipedia.org/wiki/Room_641A