5 ms·
No comment on the policy etc, but I think the presentation of the numbers is a bit misleading. That 10% is the percent of flagged images which are actually OK.
by sparsely 4y ago
No comment on the policy etc, but I think the presentation of the numbers is a bit misleading.
That 10% is the percent of flagged images which are actually OK. Whether this represents a large fraction of all legal content depends on how much illegal content there is. It would be better if they quoted the false positive rate and false negative rate as a fraction of legal/illegal images respectively.
e.g. if 1/100,000,000 legal images are flagged incorrectly, and 100% of illegal images are flagged correctly, then a corpus of 100,000,000 legal images + 9 illegal images would result in the stats in the headline. That seems like a pretty good system (ignoring any principled objections to the scanning in the first place).
- Retr0id 4y ago10% of all OK images may be falsely flagged, but what about specific categories of images? What would the error rate be for an album of legal pornography? What would the error rate be for an album of a family spending the day at a beach?
- sparsely 4y ago> 10% of all OK images may be falsely flagged, but what about specific categories of images? The article doesn't say what % of OK images are falsely flagged, only what percent of flagged images are OK. Agree that subsets might have different stats, but there's 0 information in the featured article about any of that!
- Retr0id 4y agoAh, I see what you mean.
- kmeisthax 4y agoThe (false negative) error rate on a parent taking pictures of an unusual growth on their child's naughty bits and texting it to their doctor is 0% - as in, the algorithm correctly identifies this as CSAM every time. Nevertheless, nobody[0] would consider this CSAM, based purely on context that... absolutely will not be available to any of the intermediaries charged with scanning for it. What system does the scanning won't matter, because it's not a question of accuracy; it's a question of missing information. This isn't even a hypothetical. The case I mentioned above actually happened. Google narc'd on someone who did exactly the thing I mentioned, and they got a police investigation for their trouble. And while the police - armed with exhonorating context - did eventually drop charges, they are still banned from Google for life. Yes, Google was contacted by journalists about this guy and they said they stand by their initial decision to permaban him. The core danger with any automated scanning system is not the false positive or false negative rate - that is an engineering problem. It will get better. The danger is the amount of legal-but-crime-adjacent activities that will be effectively prosecuted as crimes. I mentioned child telehealth above, but there's other instances of automated prosecution expanding far beyond the letter or spirit of the law. YouTube Content ID is top-of-mind for me; though that has two problems. It both kills fair use, and correctly prosecutes copyright infringements that people expect to fly under the radar. [0] Child anti-explotation agencies still caution against this because they don't want your kid to get used to getting photographed down there. However, they wouldn't consider this sexual abuse.
- viknesh 4y agoThis headline seems factually incorrect given what the post actually claims was said. What's described isn't a 10% error rate, but rather 90% precision. It seems like the actual thing which would be called (type I) error rate isn't discussed at all. I can only imagine that the reddit post was written with that title in bad faith/to promote fearmongering. It seems like this nuance escaped the majority of the commenters on HN as well.
- croes 4y agoThe 90% precision seems to be text only and without the rate of false negatives it is pretty useless. Could be the software only finds text that are pretty obvious grooming then 90% precious is not so good.