11 ms·
Lots of people responding to this seem to not understand how perceptual hashing / PhotoDNA works. It's true that they're not cryptographic hashes, but the false
by fortenforge 5y ago
Lots of people responding to this seem to not understand how perceptual hashing / PhotoDNA works. It's true that they're not cryptographic hashes, but the false positive rate is vanishingly small. Apple claims it's 1 in a trillion [1], but suppose that you don't believe them. Google and Facebook and Microsoft are all using PhotoDNA (or equivalent perceptual hashing schemes) right now. Have you heard of some massive issue with false positives?
The fact of the matter is that unless you possess a photo that exists in the NCMEC database, your photos simply will not be flagged to Apple. Photos of your own kids won't trigger it, nude photos of adults won't trigger it; only photos of already known CSAM content will trigger (and that too, Apple requires a specific threshold of matches before a report is triggered).
[1] "The threshold is selected to provide an extremely low (1 in 1 trillion) probability of incorrectly flagging a given account." Page 4 of https://www.apple.com/child-safety/pdf/CSAM_Detection_Technical_Summary.pdf https://www.apple.com/child-safety/pdf/CSAM_Detection_Techni...
- deleted 5y ago[deleted]
- deleted 5y ago[deleted]
- shawnz 5y agoThe 1 trillion figure is only after factoring in that you would need multiple false positives to trigger the feature. It's not descriptive of the actual false positive rate of the hashing itself.
- FabHK 5y agoProbability of a false positive for a given image = p Probability of N false positives (assuming independence) = p^N Threshold N is chosen by Apple such that p^N < 10^-12, or N log p < -12 log 10, or N > -12 log(10)/log(p) [since log(p) < 0, since p < 1]. ETA: Suppose, just for the sake of the argument, that p = 10^-3 (one false positive in 1000, so really quite bad). Then log(p) = -3 log(10), so N > -12 log(10)/(-3 log(10)) = 12/3 = 4. Similarly, if p is one in a million (10^6), then N would be required to be > 12/6 = 2. In practice, I'd expect N to be larger than 4, in other words, Apple being very conservative here. ETA: The above doesn't take into account how many images M you have. The analysis gets more complicated, but N needs to be way larger than 4. I'll think about it some more.
- ec109685 5y agoIf the threshold is 4, then 4 photos needed to have incorrectly matched, meaning the accuracy needs to only be 1/10000 if the average user has 10k images. (1/10000)^4*10k is 1 trillion.
- deleted 5y ago[deleted]
- stickfigure 5y agoThe false positive rate for any given image is not 1 in a trillion. Perceptual hashing just does not work like that. It also suffers from the birthday paradox problem - as the database expands, and the total number of pictures expands, collisions become more likely. The parent poster does make the mistake of assuming that other pictures of kids will likely cause false positives. Anything could trigger a false positive - especially flesh tones. Like, say, the naughty pictures you've been taking of your (consenting) adult partner. I'm sure Apple's outsourced low-wage-country verification team will enjoy those.
- kristofferR 5y ago> Like, say, the naughty pictures you've been taking of your (consenting) adult partner. I'm sure Apple's outsourced low-wage-country verification team will enjoy those. I'm not sure they'll be able to after looking at CSAM all day...
- ajsnigrutin 5y agoDepends... This would be a great job for an actual pedophile.
- windexh8er 5y ago> The false positive rate for any given image is not 1 in a trillion. Perceptual hashing just does not work like that. It also suffers from the birthday paradox problem - as the database expands, and the total number of pictures expands, collisions become more likely. There was a good article [0] that was on HN a couple days ago that touches on the flat out lie regarding "one in a trillion" and how PhotoDNA sounds poorly thought out. [0] https://www.hackerfactor.com/blog/index.php?/archives/929-One-Bad-Apple.html https://www.hackerfactor.com/blog/index.php?/archives/929-On...
- jrek 5y agoApple is not using PhotoDNA, but in any event - this is a good article but they misconstrued the '1 in a trillion' quote, as is canvassed in the comments for the article itself. According to apple there is a 1 in a trillion chance of wrongly flagging an account, not a 1 in a trillion false positive rate for individual images, and we know that step to banning include: - multiple flagged images; and - human review The details for those processes hasn't been fully disclosed, and it isn't possible to say whether 1 in a trillion is a reasonable estimate or otherwise.
- ummonk 5y agoTo be clear, it's 1 in 1 trillion per account. 1 in 1 trillion per photo would potentially be a more realistic risk, since some people take tens of thousands of photos.
- samstave 5y agoexcept for the time that its 100% for an account... But the truly sad thing is that there are trillions of instances, photos, moments that will never even see the light of day. and plenty of human children get abused, raped, murdered every single day. Child abuse should be a capital crime.
- andrei_says_ 5y agoWho looks at the photos in that database? How do we know it is a trustworthy source? That it doesn’t contain photos of let’s say activists or other people of interest unrelated to its projected use?
- samstave 5y agoHow do we know that the people who look at the photos are not child molesters, pedophiles, psychopaths themselves? -- Seriously. The people who worked in the abuse department at facebook when I was there were an odd group of people with whom I would not trust my kids to be around unattended. Zuck had some weird fucking bodyguards as well... --- Cool - I'll expect that child abusers are downvoting this. When did you stop abusing your kids?
- _carbyau_ 5y agoDon't forget staff rotation. A database can live forever. You not only have to trust what is now, but what the future leaders of Apple may do, what all the future image viewers may do and who they are... This is a pandaora's box of trust. Once you open it, you have to trust in perpetuity.
- hsn915 5y agoI think most people don't upload to facebook pictures of their kids taking a bath? But they more than likely store such pictures on their phones/laptops.
- simondotau 5y agoSure, but none of those images will be a hash match to any material in NCMEC databases.
- hsn915 5y agoWhat if they decide next year to use AI to automatically detect potential violating images? I hope you won't throw the "slippery slope is a fallacy" fallacy at me.
- simondotau 5y agoSlippery slope isn't always a fallacy, but it is the way most people use it. Most people just define the slope however they want to suit their argument. Let me Does this change to the software code represent a slippery slope of motive or opportunity? Many here have said yes, but in my opinion, no. When software can update itself, every single update is an opportunity for the software to betray you. That risk is already high; the risk profile doesn't increase because a particular change feels slippery slopey to you. As soon as any closed-source software implements automatic software updates, you've always one malicious update away from the system betraying you. Interim steps are unnecessary. Whether it's Chrome, or Firefox, or Windows, or Android. Heck, even Ubuntu. Any of them could betray you at any time. Potentially trash their reputation in the process, but that's a mere technicality. Therefore the slippery slope is the wrong metaphor. The correct metaphor is trust. Does this change lower my trust in Apple? Me personally, no. If anything, Apple's transparency has increased my trust in them. It gives me confidence that Apple won't use the fear of bad PR as an excuse to conceal serious things like this.
- hsn915 5y agoThe reason that "slippery slope" is almost never a fallacy is that people with intentions to perform actions that are outside the overton window[0] of acceptable behavior will first do a thing that lies just on the edge of the overton window. It is currently considered completely unacceptable for companies to scan the data on the user's own disk. If someone wanted to start doing that, they would have to create a plausible and convincing excuse. Once people accept that excuse and time passes, they can slowly "expand" the territory of this excuse (push the overton window further). "Protecting the children" is exactly this kind of plausible and convincing excuse. It needs not be the case that Apple specifically wants to scan user data so they came up with this excuse. It's simply that, once scanning users data "for the good of humanity" becomes acceptable, _some_ malicious actor will push the overton window further and further. Imagine in 10 years from now, all operating systems will scan users data to detect potential child porn material. It might even become required by law, or just by "social pressure". Just like it is now almost required by "social pressure" that social media platforms censor discourse and information. It's then very easy to expand this capability to also detect illegally downloaded music, or adult videos, or whatever deemed unacceptable. [0]: https://en.wikipedia.org/wiki/Overton_window https://en.wikipedia.org/wiki/Overton_window
- farmerstan 5y agoHow do new hashes get added to this database? How do we know that all the hashes are of CSAM? Who is validating it and is there an audit trail? Or can bad actors inject their own hashes into the database and make innocent people get reported as pedophiles?
- fortenforge 5y agoApple has stated point-blank that the only source of CSAM content to generate the hash list will be NCMEC and other child safety organizations. While I fully admit that NCMEC could do a better job with transparency and auditing, they are currently being used by several other platforms right now (Facebook, Google, Microsoft) without issue. Could bad actors inject hashes of non-CSAM content into the database somehow? Well even if they could do this, Apple employes human reviewers who must visually confirm that the flagged photo contains actual CSAM before they report the image. If the image does not contain CSAM, Apple is under no legal obligation to report it. More information here: https://twitter.com/AlexMartin/status/1424703642913935374/photo/1 https://twitter.com/AlexMartin/status/1424703642913935374/ph...
- cdolan 5y agoWait a second. You’re telling me that you believe a low-wage worker reviewing the worst of the worst human depravity, is going to stand up on a soap box and defend another nameless and faceless denizen of the Earth, when the crux of the argument is basically this: “Yeah I know the person tripped the safety threshold for CSAM, but these images aren’t that bad!” It’s not reasonable to trust the human reviewer. They put themselves at risk by going against the automated system. I’d bet 95% of people in this circumstance would pass the buck to the FBI to make their determination, at which point “your” life is already ruined.
- simondotau 5y agoThis is a misunderstanding of the system. It's not a classifier, it's a hash. The only images that are going to be seen by this low-wage worker are going to be: A) Images which are a correct hash match to an image already known to NCMEC or other agencies which have already been assigned CSAM category A1 (A=prepubescent, B=sex act); B) Images which are a hash collision. According to Apple, the likelihood of a collision is 1 in 1 trillion per user account. This system isn't a child detector strapped to a porn detector, being backed up by a low-wage worker making legal or editorial judgement calls. It's searching for images already known to child safety organisations—and even then only the most unambiguously horrific classification within the set of known images, far far far beyond the point where any ambiguity could possibly reside.
- akersten 5y agoThis is all behind a huge, neon-flashing-lights asterisk of "for now." How long until they try to machine-learn based on that database? The door's open.
- fortenforge 5y agoApple previously stored photo backups in their cloud in cleartext. The door was always open. At some point if you are providing your personal images for Apple to store, you have to exercise a modicum of trust in the company. If you don't trust Apple, I suggest you don't use an iPhone.
- Hackbraten 5y agoI don’t even use iCloud but what you suggest is exactly what I’m going to do. I no longer trust Apple, and I’m going to get rid of my iPhone.
- cgio 5y agoI would argue that in a discussion about privacy, if trust is brought up it’s no longer a discussion about privacy.
- ajsnigrutin 5y agoApple never had the technology to do offline scans of stuff directly on your phone. Now they have it. How long will it take, until some three/four letter agency makes them scan other stuff too, since they clearly have the tech to do so, and it's just one "if()" removed, to check eg. who was the first person with wikileaked photos.
- sneak 5y agoTiananmen square tank man. The CCP already gets Apple to censor the Taiwanese flag emoji and store all Chinese iCloud user data on CCP-accessible servers in China. We know Apple presently actively and eagerly cooperates with CCP censors to be permitted to operate and sell in China. This is tailor-made for CCP abuse.
- itake 5y agoHow do we know the database itself does not have any false positives?
- b215826 5y agoBecause it has been vetted by trustworthy FBI agents who don't make any mistakes.
- paulie4542 5y agoI don’t want it done on my device. Might as well let the police in my home whenever they want to rummage through my papers.
- samstave 5y agoExplain to me how photos get into the NCMEC database to begin with
- simondotau 5y agoI've absolutely no knowledge of how they operate, but it occurs to me that there would be at least two very obvious avenues: 1) During the course of investigation an officer infiltrates a CSAM sharing ring and/or poses as a customer for CSAM. Material is shared with the officer as it would be to an actual consumer of CSAM. 2) When someone is charged with child abuse, possession of child porn, etc, their physical and electronic lives will be methodically and forensically searched for CSAM material. They will likely find material they already know about, but potentially uncover new material and/or new social networks. Any material acquired would need to be analysed and classified for the purpose of effective prosecution. My understanding is (from other comments made by people on other websites) that images in the NCMEC database are tagged based on the severity of their content and that Apple is only scanning for the most extreme "A1" material. I wasn't sure what A1 meant so I googled it. According to this[0] PowerPoint presentation, page 22: A = prepubescent minor B = pubescent minor 1 = sex act 2 = "lascivious exhibition" If you want to ruin your day, the PDF provides very specific—depressingly, grossly specific—definitions for the above. [0] https://www.prosecutingattorneys.org/wp-content/uploads/Presentation-Slides2.pdf https://www.prosecutingattorneys.org/wp-content/uploads/Pres...
- tick_tock_tick 5y ago3) No one really knows and since the content is illegal to possess no one can audit it either.
- vineyardmike 5y agoThat 1/1t rate apple gave is post human review according to a recent interview
- CommieBobDole 5y agoDo you have a link to that interview? That's a really big asterisk on the "one in a trillion" claim if true.
- simondotau 5y agoAccording to Apple's technical summary: "The threshold is selected to provide an extremely low (1 in 1 trillion) probability of incorrectly flagging a given account. This is further mitigated by a manual review process wherein Apple reviews each report to confirm there is a match..." So no, "one in a trillion" doesn't include manual review.
- vineyardmike 5y ago> Apple Privacy head Erik Neuenschwander addresses concerns about its new systems to detect CSAM > We want to ensure that the reports that we make to NCMEC are high-value and actionable, and one of the notions of all systems is that there’s some uncertainty built in to whether or not that image matched. And so the threshold allows us to reach that point where we expect a false reporting rate for review of one in 1 trillion accounts per year. https://techcrunch.com/2021/08/10/interview-apples-head-of-privacy-details-child-abuse-detection-and-messages-safety-features/ https://techcrunch.com/2021/08/10/interview-apples-head-of-p... Leaves some ambiguity but it sounds like the reporting to gov is 1/1t
- hnaccy 5y agoEven if this is true (I don't trust apple's claim) I think this should still be used as a talking point to scare normal uninformed people about the surveillance tech. The other side started it first. Look at the government's dubious claims about terrorism prevention. John Walsh a founder of NCMEC testified to congress that millions of children were abducted every year and that america was "littered with mutilated, decapitated, raped, strangled children," (this was and is not true). If fear mongering about Big Brother throwing normal people in prison for pictures of their children is what it takes to blunt the expansion of the surveillance state I say fair play.
- chillfox 5y agoI have tried to play around with perceptual hashing on an image set consisting of very similar images (flowers) and there were clashes all the time.
- bonestamp2 5y ago> Apple requires a specific threshold of matches before a report is triggered True, but the wording of that condition was very vague... the threshold could be 1.
- deleted 5y ago[deleted]
- cyanydeez 5y agoright, this tech is primarily about fingerprinting existing child porn, hashing that, and trying to prevent dissemination and what they might call "ininiation" into trading child porn for the casuals. like most privacy invasions these days, its casting a widenet to put a dent in a problem that sill inevitable just route around it, and soon enough itll just be turned into a copyright cashgrab
- devwastaken 5y agoToday PhotoDNA, tomorrow NeuralNet smart AI. PhotoDNA is limited to hashes because performance 15 years ago, which makes it limited to known CSAM. However there is no doubt a push to pick up new CSAM content before it's ever distributed, which will absolutely introduce false positives.
- heavyset_go 5y ago> It's true that they're not cryptographic hashes, but the false positive rate is vanishingly small. What are you talking about? Ask anyone who has worked in this space: false positives are abound[1][2], especially when you're looking for fuzzy matches. And you have to look for fuzzy matches otherwise slight modifications to illegal images would bypass the detection system. > Photos of your own kids won't trigger it, nude photos of adults won't trigger it; This is also incorrect. In general, if two images kind of look like one another when you squint, they're going to have similar perceptual hashes. A lot of unrelated things look similar to one another when you squint, and a lot of unrelated things are going to have similar perceptual hashes. And, again, you'll be doing fuzzy matches on these hashes, so you're going to pick up those unrelated things even more so than when you just have hash collisions. [1] https://news.ycombinator.com/item?id=28091750 https://news.ycombinator.com/item?id=28091750 [2] https://news.ycombinator.com/item?id=28110159 https://news.ycombinator.com/item?id=28110159
- deleted 5y ago[deleted]
- BiteCode_dev 5y agoYoutube current censorship outrage with abusive video take downs is a proof that perceptual hash false positive are a massive problem. Even if the tech has 1 in a trillion chance, it will happen a lot with billions of people generatin thousands of ilage every years. And of course, the hash database being abused.
- admax88q 5y agoAre those YouTube takedowns happening due to perceptual hashes? I assume they are just standard ML.
- deleted 5y ago[deleted]
- intricatedetail 5y agoI worked with perceptual hashes. The false positive rate makes it unusable. Even when combined with AI it does not work well. It can get you a list of possible matches, but only a human can really tell if the image is the image. Then you have a problem of images not making to the list. That was in 2016. Maybe things changed.
- jeffe 5y agoEven assuming this is true, this is just the 'if you're not doing anything wrong, you have nothing to worry about' argument, which has been shown to have numerous flaws and counterarguments.
- xtiansimon 5y agoThank you for the paper. >"...report iCloud users who store known Child Sexual Abuse Material (CSAM) in their iCloud Photos accounts." OK. They're not trying to tag and bag your images for 'abuse content'. If you collect your child abuse porn on iCloud, we're going to report you?
- __app_dev__ 5y agoEven if that were the case this will now make it possible to "swat" or target someone by using spyware to send pics to someones phone. Obviously not many people have access to this type of spyware (think NGO Pegasus) but if law enforcement really wanted to target someone this way they will be able to do so. For everyone else getting access to the phone and manually loading the pic would work. Additionally it's possible to fool AI: https://slazebni.cs.illinois.edu/fall18/lec12_adversarial.pdf https://slazebni.cs.illinois.edu/fall18/lec12_adversarial.pd...
- Siira 5y agoSo we just believe Apple. We can call it security by infinite-trust. Beats even security by obscurity!