6 ms·
Apple's recent CSAM announcement does not use PhotoDNA. Apple is not using PhotoDNA's hashes. Apple is using a different system called NeuralHash. NCMEC operat
by hackerfactor1 5y ago
Apple's recent CSAM announcement does not use PhotoDNA. Apple is not using PhotoDNA's hashes. Apple is using a different system called NeuralHash.
NCMEC operates like a roach hotel: CSAM goes in, and it only goes out to law enforcement. I am 99% certain that Apple never received CSAM material from NCMEC for training. While Apple's people may have sat in NCMEC's building for meetings, I doubt they were ever shown any actual CSAM. I've been told that NCMEC used a tool by Apple to create hashes for NeuralHash. The hashes are based on some number of "the most common" CSAM content. (In the past, "the most common" was around 30,000 files.)
In effect, Apple may not know what is in their trained data set. However, researchers (not me) are already looking for ways to extract data from the libraries. I think it's only a matter of time before this becomes a much bigger problem.
- nomel 5y ago> However, researchers (not me) are already looking for ways to extract data from the libraries. I think it's only a matter of time before this becomes a much bigger problem. What's the use of the data, and what problems do you see?
- pornel 5y agoWhen the algorithm is known, adversaries can figure out how to minimally modify images to certainly avoid detection. Uncertainty about detection capabilities may have been a deterrent.
- nomel 5y agoOne could claim that knowing the algorithm exists, for all cloud photo services (scanning isn't just an Apple iCloud thing), means that these people just don't upload to cloud services and not worry about any of this. For the Apple implementation, avoiding detection is as simple as switching off iCloud sync in the settings. In the case where they've modified the images to enable iCloud sharing, that would be risky game since those tweaked images would most likely make their way back into the database, and any tweak to the algorithm would mean it's all detected. My naive assumption would be that the latest model would be required to interact with the iCloud service (seems like a good idea at least).
- judge2020 5y agoWell, according to this site: > Perhaps there is a reason that they don't want really technical people looking at PhotoDNA. Microsoft says that the "PhotoDNA hash is not reversible". That's not true. PhotoDNA hashes can be projected into a 26x26 grayscale image that is only a little blurry. 26x26 is larger than most desktop icons; it's enough detail to recognize people and objects. Reversing a PhotoDNA hash is no more complicated than solving a 26x26 Sudoku puzzle; a task well-suited for computers. > I have a whitepaper about PhotoDNA that I have privately circulated to NCMEC, ICMEC (NCMEC's international counterpart), a few ICACs, a few tech vendors, and Microsoft. The few who provided feedback were very concerned about PhotoDNA's limitations that the paper calls out. I have not made my whitepaper public because it describes how to reverse the algorithm (including pseudocode). If someone were to release code that reverses NCMEC hashes into pictures, then everyone in possession of NCMEC's PhotoDNA hashes would be in possession of child pornography. https://www.hackerfactor.com/blog/index.php?/archives/929-One-Bad-Apple.html https://www.hackerfactor.com/blog/index.php?/archives/929-On...
- deleted 5y ago[deleted]
- JimDabell 5y ago> 26x26 is larger than most desktop icons; it's enough detail to recognize people and objects. This doesn’t seem to be true. These are 26×26 icons, for instance: https://cdn.dribbble.com/users/975594/screenshots/6193663/icon___26x26__.png https://cdn.dribbble.com/users/975594/screenshots/6193663/ic...
- TonyTrapp 5y agoDepends on what you define as "Desktop Icons". Many icons in Windows are 16x16. The icons that literally sit on the Windows desktop are 32x32 or bigger, though.
- nomel 5y agoHere's a greyscale 26x26 of Obama: torso: https://imgur.com/a/Oz0pnyF https://imgur.com/a/Oz0pnyF closeup of face: https://imgur.com/a/Hr1OS2P https://imgur.com/a/Hr1OS2P I guess it has to be decided if wide access to blurry images like this is better or worse than finding full resolution images. This all assumes there's not some obfuscation layer, I suppose (is that possible?).
- hackerfactor1 5y agoAll perceptual hashes, including AI-based perceptual hashes, have a "projection" property. If you have the hashes, you can project them into some kind of image. Some hashes result in blurry blobs (pHash, wavelet hashes). Some result in silhouettes (aHash, dHash). And some show well-defined images (PhotoDNA). We don't know how Apple's solution works. But if it can project recognizable images, then it means people can use Apple's hash system to regenerate child porn. Without details and an actual code review, I'm not willing to accept Apple's assurance that an image projection isn't possible. We hear that same promise from Microsoft's PhotoDNA, and it turned out to be false. Imagine the problems if every iPhone and every Mac was in possession of child porn because Apple put it there... Yes, this is a lot of 'if's, but until it has been evaluated, it is in the very real realm of possible.
- JimDabell 5y ago> Apple is not using PhotoDNA's hashes. Do you have a source for this, or is it an assumption you are making? > Apple is using a different system called NeuralHash. No, one of the systems Apple is using is NeuralHash. They also use a second, undisclosed system on the visual derivative before human review. While it’s possible they have invented two separate perceptual hashing systems, it’s also quite possible that they would use PhotoDNA for this. Details from Apple: > Once Apple's iCloud Photos servers decrypt a set of positive match vouchers for an account that exceeded the match threshold, the visual derivatives of the positively matching images are referred for review by Apple. First, as an additional safeguard, the visual derivatives themselves are matched to the known CSAM database by a second, independent perceptual hash. This independent hash is chosen to reject the unlikely possibility that the match threshold was exceeded due to non-CSAM images that were adversarially perturbed to cause false NeuralHash matches against the on-device encrypted CSAM database. If the CSAM finding is confirmed by this independent hash, the visual derivatives are provided to Apple human reviewers for final confirmation. — https://www.apple.com/child-safety/pdf/Security_Threat_Model_Review_of_Apple_Child_Safety_Features.pdf https://www.apple.com/child-safety/pdf/Security_Threat_Model...
- cwizou 5y ago> They also use a second, undisclosed system on the visual derivative before human review. While it’s possible they have invented two separate perceptual hashing systems, it’s also quite possible that they would use PhotoDNA for this. Yep I probably didn't phrase my question well enough but that was exactly part of what I was wondering. They claim it's a second independent hash as you pointed out, so I would assume this is PhotoDNA server side, unless they made two separate hash systems (one which was not initially disclosed, if not implemented, at announcement time) ? I'm also wondering exactly how they trained their NeuralHash, but may have missed some parts of the discussion. Supposedly NCMEC is not sharing the database of the "original" content that one would use to train such a system. Edit : Guess that's a moot point now, they postponed it!