4 ms·
> What is not mentioned is that Claude Code also found one thousand false positive bugs, which developers spent three months to rule out. Source? I haven't see
by mtlynch 6mo ago
> What is not mentioned is that Claude Code also found one thousand false positive bugs, which developers spent three months to rule out.
Source? I haven't seen this anywhere.
In my experience, false positive rate on vulnerabilities with Claude Opus 4.6 is well below 20%.
- r9295 6mo agoIn my experience, the issue has been likelihood of exploitation or issue severity. Claude gets it wrong almost all the time. A threat model matters and some risks are accepted. Good luck convincing an LLM of that fact
- j16sdiz 6mo agoIn TFA: I have so many bugs in the Linux kernel that I can’t report because I haven’t validated them yet… I’m not going to send [the Linux kernel maintainers] potential slop, but this means I now have several hundred crashes that they haven’t seen because I haven’t had time to check them. —Nicholas Carlini, speaking at [un]prompted 2026
- mtlynch 6mo agoThose aren't false positives; they're results he hasn't yet inspected. I wrote a longer reply here: https://news.ycombinator.com/item?id=47638062 https://news.ycombinator.com/item?id=47638062
- bethekidyouwant 6mo agosome of them certainly are…
- coldtea 6mo ago>Those aren't false positives; they're results he hasn't yet inspected. It's not a XOR
- Ukv 6mo agoThe article quote was being given as the supposed source for "Claude Code also found one thousand false positive bugs, which developers spent three months to rule out", so should substantiate that claim - which it doesn't. If the claim was instead just "a good portion of the hundreds more potential bugs it found might be false positives", then sure.
- tptacek 6mo agoYes it is. They're not not false positives until they're reported and consume maintainer time.
- lambdaone 6mo agoFalse positives can be eliminated mechanistically by testing if they actually work, in a sufficiently isolated automated test apparatus. The hard thing is reducing detected crashes to well-formulated test cases that help rather than hinder maintainers.
- sobiolite 6mo agoThe comment said "Claude Code also found one thousand false positive bugs, which developers spent three months to rule out.". Please explain how a bug can both be unvalidated, and also have undergone a three month process to determine it is a false positive?
- christophilus 6mo agoSame. Codex and Claude Code on the latest models are really good at finding bugs, and really good at fixing them in my experience. Much better than 50% in the latter case and much faster than I am.
- paulddraper 6mo agoSource: """AI is bad"""
- Supermancho 6mo agoTo the issue of AI submitted patches being more of a burden than a boon, many projects have decided to stop accepting AI-generated solutioning: https://blog.devgenius.io/open-source-projects-are-now-banning-ai-generated-pull-requests-8e1dd3e8d41c https://blog.devgenius.io/open-source-projects-are-now-banni... These are just a few examples. There are more that google can supply.
- deleted 6mo ago[deleted]
- literalAardvark 6mo agoNo, they haven't. Read the ai slop you posted carefully. It's a policy update that enables maintainers to ignore low effort "contributions" that come from untrusted people in order to reduce reviewing workload. An Eternal September problem, kind of.
- coldtea 6mo agoDidn't you just restate what the parent claimed?
- cwillu 6mo agoNo, that's not at all the same thing: ai-generated contributions from people with a track record for useful contributions are still accepted.
- dpark 6mo agoRight. AI submissions are so burdensome that they have had to refuse them from all except a small set of known contributors. The fact that there’s a small carve out for a specific set of contributors in no way disputes what Supermancho claimed.
- phanimahesh 6mo ago