6 ms·
My initial thoughts were so they could scan for csam while pretending as if users have a choice to not have their privacy violated.
by 14 1y ago
My initial thoughts were so they could scan for csam while pretending as if users have a choice to not have their privacy violated.
- bayindirh 1y agoFrom my understanding, CSAM scanning is always considered a separate, always on and mandatory subsystem in any cloud storage system.
- odo1242 1y agoYes, any non E2EE cloud storage system has strict scanning for CSAM. And it's based on perceptual hashes, not AI (because AI systems can be tricked with normal-looking adversarial images pretty easily)
- heavyset_go 1y agoI built a similar photo ID system, not for this purpose or content, and the idea of platforms using perceptual hashes to potentially ruin people's lives is horrifying. Depending on the algorithm and parameters, you can easily get a scary amount of false positives, especially using algorithms that shrink images during hashing, which is a lot of them.
- dotnet00 1y agoI imagine you'd add more heuristics and various types of hashes? If the file is just sitting there, rarely accessed and unshared, or if the file only triggers on 2/10 hashes, it's probably a false alarm. If the file is on a public share, you can probably run an actual image comparison...
- heavyset_go 1y agoA lot of classic perceptual hash algorithms do "squinty" comparisons, where if an image kind of looks like one you've hashed against, you can get false positives. I'd imagine outside of egregious abuse and truly unique images, you could squint at a legal image and say it looks very much like another illegal image, and get a false positive. From what I'm reading about PhotoDNA, it's your standard phashing system from 15 years ago, which is terrifying. But yes, you can add heuristics, but you will still get false positives.
- JimDabell 1y agoI thought Apple’s approach was very promising. Unfortunately, instead of reading about how it actually worked, huge amounts of people just guessed incorrectly about how it worked and the conversation was dominated by uninformed outrage about things that weren’t happening.
- dmitrygr 1y ago> instead of reading about how it actually worked, huge amounts of people just guessed incorrectly about how it worked and the conversation was dominated by uninformed outrage I would not care if it worked 100% accurately. My outrage is informed by people like you who think it is OK in any form whatever.
- JimDabell 1y ago[flagged]
- dmitrygr 1y agoNo amount of my device spying on me is acceptable, no matter how cleverly implemented. The fact that your comment said anything positive about it at all without acknowledging that it is an insane idea and should never be put into practice is what I was referring to.
- WarOnPrivacy 1y ago> Unfortunately, instead of reading about how it actually worked, huge amounts of people just guessed incorrectly about how it worked Folks did read. They guessed that known hashes would be stored on devices and images would be scanned against that. Was this a wrong guess? > the conversation was dominated by uninformed outrage about things that weren’t happening. The thing that wasn't happening yet was mission creep beyond the original targets. Because expanding-beyond-originally-stated-parameters is thing that happens with far reaching monitoring systems. Because it happens with the type of regularity that is typically limited to physics. There were 2ndary concerns about how false positives would be handled. There were concerns about what the procedures were for any positive. Given Gov propensities to ruin lives now and ignore that harm (or craft a justification) later, the concerns seem valid. That's what I recall the concerned voices were on about. To me, they didn't seem outraged.
- odo1242 1y agoYeah, it’s not a great system due to the fact that perceptual hashes can and have been tricked in the past. It is better than machine learning though because you can make any image trigger an ML model without necessarily looking like a bad image. That is, perceptual hashes are much harder to adversarially fool.
- heavyset_go 1y agoI agree, and maybe I'm wrong, but I see a similarity between phash quantization and DCT and ML kernels. I think you could craft "invisible" adversarial images similarly for phash systems like you can ML ones and the results could be just as bad. They'd probably replicate better than adversarial ML images, too. I think the premise for either system is flawed and both are too error prone for critical applications.
- robotresearcher 1y agoPerceptual hashes? An embedding in a vector space by a learned encoder. Phew, not AI then… ?