4 ms·
I don't really get what this repository is trying to achieve and what's the point of collecting collisions. Collisions will happen, that's just how it is with h
by StrLght 5y ago
I don't really get what this repository is trying to achieve and what's the point of collecting collisions. Collisions will happen, that's just how it is with hashes.
It's already a public knowledge that Apple has 2 more systems (some server-side verification and a manual check later) to prevent false-positives. So what's the point of researching collisions in NeuralHash?
- woofie11 5y agoNo. Most proper cryptographic hash systems (e.g. used for verifying files, rather than data structures) never have collisions. Try to find a SHA256 collision. Anywhere, ever, in the history of mankind. This isn't for lack of looking. A lot of very smart people have looked for them. If you find one, I bet you'll be eligible for a tenured faculty slot at a good university, if not more. A whole world of secure systems would need to be re-engineered. Hypothetical collisions of course exist, by the pigeonhole principle, just not in the real world.
- StrLght 5y agoYes, but cryptographic hashes are irrelevant here because they'd allow to easily bypass CSAM by modifying/appending a single byte.
- woofie11 5y agoApple is claiming to have a visual equivalent to a cryptographic hash -- one which won't change with a single byte, but only if the image is substantially different. At least their security analysis relies on that. From their whitepaper: "The threshold is selected to provide an extremely low (1 in 1 trillion) probability of incorrectly flagging a given account" If your claim is that their hash algorithm isn't cryptographic, their security analysis is incorrect.
- kbelder 5y agoTheir security analysis is obviously incorrect. "equivalent to a cryptographic hash" "change... only if the image is substantially different" Both are not true, cannot be true.
- nullc 5y agoIt's trivial to cause a neural hash non-match-- either imperceptibly with a little noise, or by adding an overlay on the image. If you downsample and quantize an image before sha256ing it you get a bit of robustness to accidental false-negatives. While both schemes are trivially bypassable.
- mlajtos 5y agoIf I found one collision by accident, would that be any significant?
- saithound 5y agoYes, absolutely. It's technically possible to find SHA256 collisions accidentally, but it's so unlikely, that if you found one, it would merit serious investigation. People would not believe your statement that you found them accidentally, and "oh, I guess mlajtos really just found the colliding pair by chance" wouldn't be declared until after a very thorough investigation. In the meantime, major stakeholders (e.g. Bitcoin) would probably move away to another hash function, just in case. Dwyer calculated 1431168 NeuralHashes and found two collisions. Humanity collectively calculates over 120000000000000000000 SHA-256 hashes every second. Still, we're reasonably sure that this immense brute-force search will not lead to any collisions in any reasonable amount of time.
- jtsiskin 5y agoA good perspective on how big the SHA-256 hash space is: https://youtu.be/S9JGmA5_unY https://youtu.be/S9JGmA5_unY
- oauea 5y agoAre you fine with an apple employee looking at all your private pictures just because some hashing algorithm decided you're a pedophile? Personally, I'm not.
- StrLght 5y agoFirst, this doesn't change my opinion on CSAM, I still consider the thing way too intrusive until Apple announces E2E for iCloud. Now, I can't really call something I voluntarily uploaded to Apple's servers a "private picture". But that's just a matter of perspective, and I understand that many people would disagree with me on this.
- oauea 5y agoGood luck convincing all the couples who back up all their pictures by default to iCloud that those pictures are not private.
- dhosek 5y agoI'd argue that the hash collisions (both natural and synthetic) that I've seen give me more confidence in the system, not less. On the natural hash collisions (of which there are two), we have objects of similar shape against a solid background. It seems that a natural hash collision of a CSAM image would be unlikely (or if it does occur, it would be something that perhaps is also an infringing image). As for the synthetic hash collisions, there are visible artifacts in the picture that, if you compare with the original picture, make the overlay of the original picture on the synthetically generated hash collision obvious. Could people get tricked into downloading memes¹ with synthetically generated hash collisions? Sure, people are idiots. But I'm guessing the majority of folks will look at the picture and say, this is a sh*t picture in this meme and download something else. 1. And that, of course, assumes that meme hosters don't apply similar scanning techniques to what they serve up.
- nullc 5y ago> this is a sh*t picture in this meme and download something else. NO. Adversarial preimages can be created that look like perfectly normal images. Please stop repeating this falsehood. Here are some examples I generated (with a link to more): https://github.com/AsuharietYgvar/AppleNeuralHash2ONNX/issues/1#issuecomment-903094036 https://github.com/AsuharietYgvar/AppleNeuralHash2ONNX/issue...
- dhosek 5y agoI saw two examples. The dog/girl one has obvious artifacts in the picture of the girl. The directly linked image doesn't have artifacts, but putting them side by side, I can see what got matched up and they're still very approximately the same picture in that they're both pictures of women and the eyes are in the same part of the picture which gives support to your perspective. I do still wonder whether it would be possible to take, say, a (legal) nude and turn it into an innocent-looking image that still matches the hashes or not. I'm more inclined to believe it now than before, but theoretical possibilities don't usually map to realistic concerns.¹ 1. One example of this would be that theoretically, LaTeX's cross-reference mechanism can get caught in a cyclic state. This can only happen with page references and the most likely scenario is a reference to a roman-numeraled page number where if the page reference is output as ix the referenced location moves to page x and when the page reference is updated to x the referenced location moves to page ix (in practice, functioning examples required a shift between xcix and c, but either way, the probability of this happening in a real document is vanishingly small).