4 ms·
I think in this case the fact that there are no false negatives means that the low false positive rate is enough to infer that a positive result implies high li
by gopiandcode 6y ago
I think in this case the fact that there are no false negatives means that the low false positive rate is enough to infer that a positive result implies high likelihood of actually being malicious.
You can reason roughly as follows:
P[ pos | mal ] = 1 (no false negatives)
= P[ pos /\ mal ] = P[ mal ] (Bayes)
A low false positive rate means that:
P[ pos | ¬ mal ] ~= 0 (low false positive rate)
P[ pos /\ ¬ mal ] / P[ ¬ mal ] ~= 0
(P[ pos ] - P[ pos /\ mal])/(1 - P[mal]) ~= 0
(P[ pos ] - P[ mal ])/(1 - P[ mal ]) ~= 0
From this fraction we can conclude:
P[ pos ] - P[ mal ] ~= 0
Returning back to a the likelihood of being malicious given a positive result:
P[ mal | pos ]
= P[ mal /\ pos ] / P[pos]
= P[ mal ] / P[ pos ]
~= 1.0
- hanoz 6y agoI think you're right. I'll stick to the day job.
- sukilot 6y agoNo, you were right, and OP was wrong.
- hanoz 6y agoNo, wait, I think I was right after all. Say we have a million urls, and a thousand of them are malicious. Our filter returns a positive result for all the malicious 1000 (no false negatives) and for the safe 999000 urls only 1% will return positive (low false positive), but that's still 9900 false positives. So a positive result only has a (1000 / (1000 + 9900)), i.e. 9%, chance of actually being malicious. Even with a false positive result of only 0.1%, the probability of a positive result actually being malicious only rises to 50%, so still not in "high likelihood" territory.
- gopiandcode 6y agoNice example - you were right, my wording on that section was incorrect. sukilot quite nicely presents the error in my derivation - I guess I probably shouldn't try coming up with proofs on the spot - especially if they're not machine checked. I think what I probably should have said instead was that the low false positive rate means that only a small proportion of the honest URLs will be sent up.
- sukilot 6y ago> = P[ mal ] / P[ pos ] ~= 1.0 You didn't prove this. You can't swap subtraction for multiplication. It's not equivalent to > [ pos ] - P[ mal ] ~= 0 1e-100 - 1e-200 ~= 0, but 1e-200 / 1e-100 = 1e-100 That's the point of the base rate fallacy. https://en.m.wikipedia.org/wiki/Base_rate_fallacy https://en.m.wikipedia.org/wiki/Base_rate_fallacy False negative rate doesn't matter. A low false positive rate can be much larger than the true positive rate. Imagine an extremely rare toxin that always turns your skin blue. Only 100 people in the world have it. But 1000 people are wearing blue face paint at any given time. A diagnostic test "are you blue?" Would have no false negatives and very low false positives for the toxin, just like a Bloom filter. But a positive test would not mean highly likely to have the toxin; it would mean highly likely to be wearing face paint.
- gopiandcode 6y agoGood point, mea culpa. Yes, I guess the low false positive rate doesn't actually justify believing that the the URL has a high likelihood of actually being malicious. The correct wording should probably be that the low false positive rate means that non-malicious urls are unlikely to be sent out. I'll fix this.