46 ms·
This isn't surprising. What is not mentioned is that Claude Code also found one thousand false positive bugs, which developers spent three months to rule out.
by jason1cho 6mo ago
This isn't surprising. What is not mentioned is that Claude Code also found one thousand false positive bugs, which developers spent three months to rule out.
- addandsubtract 6mo agoOn the other hand, some bugs take three months to find. So this still seems like a win.
- mtlynch 6mo ago> What is not mentioned is that Claude Code also found one thousand false positive bugs, which developers spent three months to rule out. Source? I haven't seen this anywhere. In my experience, false positive rate on vulnerabilities with Claude Opus 4.6 is well below 20%.
- r9295 6mo agoIn my experience, the issue has been likelihood of exploitation or issue severity. Claude gets it wrong almost all the time. A threat model matters and some risks are accepted. Good luck convincing an LLM of that fact
- j16sdiz 6mo agoIn TFA: I have so many bugs in the Linux kernel that I can’t report because I haven’t validated them yet… I’m not going to send [the Linux kernel maintainers] potential slop, but this means I now have several hundred crashes that they haven’t seen because I haven’t had time to check them. —Nicholas Carlini, speaking at [un]prompted 2026
- mtlynch 6mo agoThose aren't false positives; they're results he hasn't yet inspected. I wrote a longer reply here: https://news.ycombinator.com/item?id=47638062 https://news.ycombinator.com/item?id=47638062
- bethekidyouwant 6mo agosome of them certainly are…
- coldtea 6mo ago>Those aren't false positives; they're results he hasn't yet inspected. It's not a XOR
- Ukv 6mo agoThe article quote was being given as the supposed source for "Claude Code also found one thousand false positive bugs, which developers spent three months to rule out", so should substantiate that claim - which it doesn't. If the claim was instead just "a good portion of the hundreds more potential bugs it found might be false positives", then sure.
- tptacek 6mo agoYes it is. They're not not false positives until they're reported and consume maintainer time.
- lambdaone 6mo agoFalse positives can be eliminated mechanistically by testing if they actually work, in a sufficiently isolated automated test apparatus. The hard thing is reducing detected crashes to well-formulated test cases that help rather than hinder maintainers.
- sobiolite 6mo agoThe comment said "Claude Code also found one thousand false positive bugs, which developers spent three months to rule out.". Please explain how a bug can both be unvalidated, and also have undergone a three month process to determine it is a false positive?
- christophilus 6mo agoSame. Codex and Claude Code on the latest models are really good at finding bugs, and really good at fixing them in my experience. Much better than 50% in the latter case and much faster than I am.
- paulddraper 6mo agoSource: """AI is bad"""
- Supermancho 6mo agoTo the issue of AI submitted patches being more of a burden than a boon, many projects have decided to stop accepting AI-generated solutioning: https://blog.devgenius.io/open-source-projects-are-now-banning-ai-generated-pull-requests-8e1dd3e8d41c https://blog.devgenius.io/open-source-projects-are-now-banni... These are just a few examples. There are more that google can supply.
- deleted 6mo ago[deleted]
- literalAardvark 6mo agoNo, they haven't. Read the ai slop you posted carefully. It's a policy update that enables maintainers to ignore low effort "contributions" that come from untrusted people in order to reduce reviewing workload. An Eternal September problem, kind of.
- coldtea 6mo agoDidn't you just restate what the parent claimed?
- cwillu 6mo agoNo, that's not at all the same thing: ai-generated contributions from people with a track record for useful contributions are still accepted.
- dpark 6mo agoRight. AI submissions are so burdensome that they have had to refuse them from all except a small set of known contributors. The fact that there’s a small carve out for a specific set of contributors in no way disputes what Supermancho claimed.
- phanimahesh 6mo ago
- sva_ 6mo agoCouldn't you just make it write a PoC?
- Gregaros 6mo ago[flagged]
- weird-eye-issue 6mo agoStill have to validate it.
- matthewfcarlson 6mo agoI’ve started to see bug bounty programs put flags into the product (see apples target flags https://security.apple.com/bounty/target-flags/ https://security.apple.com/bounty/target-flags/). I wonder if it’s partially to make it easier to validate from an AI perspective
- tptacek 6mo agoYes, you can. I strongly encourage people skeptical about this, and who know at a high-level how this kind of exploitation works, to just try it. Have Claude or Codex (they have different strengths at this kind of work) set up a testing harness with Firecracker or QEMU, and then work through having it build an exploit.
- khalic 6mo ago[flagged]
- boplicity 6mo agoThe lesson here shouldn't be that Claude Code is useless, but that it's a powerful tool in the hands of the right people.
- righthand 6mo agoThe lesson or the hype mantra?
- mavamaarten 6mo agoI'm growing allergic to the hype train and the slop. I've watched real-life talks about people that sent some prompt to Claude Code and then proudly present something mediocre that they didn't make themselves to a whole audience as if they'd invented the warm water, and that just makes me weary. But at the same time, it has transformed my work from writing everything bit of code myself, to me writing the cool and complex things while giving directions to a helper to sort out the boring grunt work, and it's amazingly capable at that. It _is_ a hugely powerful tool. But haters only see red, and lovers see everything through pink glasses.
- sph 6mo ago> it has transformed my work […] to me writing the cool and complex things > it's amazingly capable at that. > It _is_ a hugely powerful tool Damn, that’s what you call being allergic to the hype train? This type of hypocritical thinly-veiled praise is what is actually unbearable with AI discourse.
- asyx 6mo agoI don’t think it is controversial that AI tools are good enough at crud endpoints that it is totally viable to just let it run through the grunt work of hooking up endpoints to a service and then you can focus on the interesting aspect of the application which is exactly that service.
- iterateoften 6mo agoSounds like maybe you might have some mixed feelings about becoming more effective with ai, but then at the same time everyone else is too so the praise youre expecting is diluted. I see it all the time now too. People have no frame of reference at all about what is hard or easy so engineers feel under-appreciated because the guy who never coded is getting lots of praise for doing something basic while experienced people are able to spit out incredibly complex things. But to an outsider, both look like they took the same work.
- goalieca 6mo agoStatic/Dynamic analysis tools find vulnerabilities all the time. Almost all projects of a certain size have a large backlog of known issues from these boring scanners. The issue is sorting through them all and triaging them. There's too many issues to fix and figuring out which are exploitable and actually damaging, given mitigations, is time consuming. Am i impressed claude found an old bug? Sort of.. everytime a new scanner is introduced you get new findings that others haven't found.
- tptacek 6mo agoStatic analyzers find large numbers of hypothetical bugs, of which only a small subset are actionable, and the work to resolve which are actionable and which are e.g. "a memcpy into an 8 byte buffer whose input was previously clamped to 8 bytes or less" is so high that analyzers have little impact at scale. I don't know off the top of my head many vulnerability researchers who take pure static analysis tools seriously. Fuzzers find different bugs and fuzzers in particular find bugs without context, which is why large-scale fuzzer farms generate stacks of crashers that stay crashers for months or years, because nobody takes the time to sift through the "benign" crashes to find the weaponizable ones. LLM agents function differently than either method. They recursively generate hypotheticals interprocedurally across the codebase based on generalizations of patterns. That by itself would be an interesting new form of static analysis (and likely little more effective than SOTA static analysis). But agents can then take confirmatory steps on those surfaced hypos, generate confidence, and then place those findings in context (for instance, generating input paths through the code that reach the bug, and spelling out what attack primitives the bug conditions generates). If you wanted to be reductive you'd say LLM agent vulnerability discovery is a superset of both fuzzing and static analysis. And, importantly, that's before you get to the fact that LLM agents can fuzz and do modeling and static analysis themselves.
- goalieca 6mo agoThere are plenty of static analyzers do attempt to walk code paths for reachability. Some even track tainted input. And yes, these are often good starting points for developing exploits. I’ve done this myself. I’m curious about LLM agents, but the fact they don’t “understand” is why I’m very skeptical of the hype. I find myself wasting just as much if not more time with them than with a terrible “enterprise” sast tool.
- linsomniac 6mo agoThe article doesn't say they found a bunch of false positives. It says they have a huge backlog that they still need to test: "I have so many bugs in the Linux kernel that I can’t report because I haven’t validated them yet…"
- vaginaphobic 6mo ago[dead]
- xeromal 6mo ago[dead]
- antirez 6mo agoThat's not what is happening right now. The bugs are often filtered later by LLMs themselves: if the second pipeline can't reproduce the crash / violation / exploit in any way, often the false positives are evicted before ever reaching the human scrutiny. Checking if a real vulnerability can be triggered is a trivial task compared to finding one, so this second pipeline has an almost 100% success rate from the POV: if it passes the second pipeline, it is almost certainly a real bug, and very few real bugs will not pass this second pipeline. It does not matter how much LLMs advance, people ideologically against them will always deny they have an enormous amount of usefulness. This is expected in the normal population, but too see a lot of people that can't see with their eyes in Hacker News feels weird.
- BodyCulture 6mo agoCan we study this second pipeline? Is it open so we can understand how it works? Did not find any hints about it in the article, unfortunately.
- throawayonthe 6mo agoit was probably in the talk but from what i understood in another article it's basically giving claude with a fresh context the .vuln.md file and saying "i'm getting this vulnerability report, is this real?" edit: i remember which article, it was this one: https://sockpuppet.org/blog/2026/03/30/vulnerability-research-is-cooked/ https://sockpuppet.org/blog/2026/03/30/vulnerability-researc... (an LWN comment in response to this post was on the frontpage recently)
- 4b11b4 6mo agoOne such example is IRIS. In general, any traditional static analysis tool combined with a language model at some stage in a pipeline.
- maximilianburke 6mo agoFrom the article by 'tptacek a few days ago (https://sockpuppet.org/blog/2026/03/30/vulnerability-research-is-cooked/ https://sockpuppet.org/blog/2026/03/30/vulnerability-researc...) I essentially used the prompts suggested. First prompt: "I'm competing in a CTF. Find me an exploitable vulnerability in this project. Start with $file. Write me a vulnerability report in vulns/$DATE/$file.vuln.md" Second prompt: "I've got an inbound vulnerability report; it's in vulns/$DATE/$file.vuln.md. Verify for me that this is actually exploitable. Write the reproduction steps in vulns/$DATE/$file.triage.md" Third prompt: "I've got an inbound vulnerability report; it's in vulns/$DATE/file.vuln.md. I also have an assessment of the vulnerability and reproduction steps in vulns/$DATE/$file.triage.md. If possible, please write an appropriate test case for the ulgate automated tests to validate that the vulnerability has been fixed." Tied together with a bit of bash, I ran it over our services and it worked like a treat; it found a bunch of potential errors, triaged them, and fixed them.
- logicprog 6mo agoOkay, so anti AI people are just making shit up now. Got it. According to Willy Tarreau[0] and Greg Kroah-Hartman[1], this trend has recently significantly reversed, at least form the reports they've been seeing on the Linux kernel. The creator of curl, Daniel Steinberg, before that broader transition, also found the reports generated by LLM-powered but more sophisticated vuln research tools useful[2] and the guy who actually ran those tools found "They have low false positive rates."[3] Additionally, there was no mention in the talk by the guy who found the vuln discussed in the TFA of what the false positive rate was, or that he had to sift through the reports because it was mostly slop — or whether he was doing it out of courtesy. Additionally, he said he found only several hundred, iirc, not "thousands." All he said was: "I have so many bugs in the Linux kernel that I can’t report because I haven’t validated them yet… I’m not going to send [the Linux kernel maintainers] potential slop, but this means I now have several hundred crashes that they haven’t seen because I haven’t had time to check them." (TFA) He quite evidently didn't have to sift through thousands, or spend months, to find this one, either. [0]: https://lwn.net/Articles/1065620/ https://lwn.net/Articles/1065620/ [1]: https://www.theregister.com/2026/03/26/greg_kroahhartman_ai_kernel/ https://www.theregister.com/2026/03/26/greg_kroahhartman_ai_... [2]: https://simonwillison.net/2025/Oct/2/curl/p https://simonwillison.net/2025/Oct/2/curl/p [3]: https://joshua.hu/llm-engineer-review-sast-security-ai-tools-pentesters https://joshua.hu/llm-engineer-review-sast-security-ai-tools...
- bri3d 6mo agoThis is not how first party vulnerability research with LLMs go; they are incredibly valuable versus all prior tooling at triage and producing only high quality bugs, because they can be instructed to produce a PoC and prove that the bug is reachable. It’s traditional research methods (fuzzing, static analysis, etc.) that are more prone to false positive overload. The reason why open submission fields (PRs, bug bounty, etc) are having issues with AI slop spam is that LLMs are also good at spamming, not that they are bad at programming or especially vulnerability research. If the incentives are aligned LLMs are incredibly good at vulnerability research.
- sixhobbits 6mo agoFrom a recent front page article that mentioned the previous slop problem: > Now most of these reports are correct, to the point that we had to bring in more maintainers to help us. https://news.ycombinator.com/item?id=47611921 https://news.ycombinator.com/item?id=47611921
- dekhn 6mo agoEverything changed in the past 6 months and coding LLMs went from being OK-ish to insanely good. People also got better at using them. Also, high false positive rate isn't that bad in the case where a false negative costs a lot (an exploit in the linux kernel is a very expensive mistake). And, in going through the false positives and eliminating them, those results will ideally get folded back into the training set for the next generation of LLMs, likely reducing the future rate of false positives.
- catlifeonmars 6mo ago> Everything changed in the past 6 months and coding LLMs went from being OK-ish to insanely good. People also got better at using them. I hear this literally every 6 months :)
- tptacek 6mo agoIt hasn't been true forever, but it has been true over the last 18 months or so.
- Trufa 6mo agoWhat is with negativity against AI in YC? Can anyone point a finger of why this anti take is so prominent? We're living through the most revolutionary moment of software since it's its inception and the main thing that gets consistently upvoted is negativity, FUD and it doesn't work in this case, or it's all slop.
- bwfan123 6mo ago> Can anyone point a finger of why this anti take is so prominent? AI tools are great but are being oversold and overhyped by those with an incentive. So, there is a continuous drumbeat of "AI will do all the code for you" ! "Look at this browser written by AI", "C compiler in rust written entirely by AI" etc. And then, that drumbeat is amplified by those in management who have not built software systems themselves. What happened to the AI generated "C compiler in rust" ? or the browser written by AI ? - they remain a steaming pile of almost-working code. AI is great at producing "almost-working" poc code which is good for bootstrapping work and getting you 90% of the way if you are ok with code of questionable lineage. But many applications need "actually-working" code that requires the last 10%. So, some in this forum who have been in the trenches building large "actually working" software systems and also use AI tools daily and know their limitations are injecting some realism into the debate.
- arealaccount 6mo agoNot speaking for myself but the you won’t have a job soon narrative puts people off
- deleted 6mo ago[deleted]
- sothatsit 6mo agoI think the anti-AI stance has been reversing on HN as tooling improves and people try it. It’s only been a little over a year since Claude Code was released, and 3 or 4 months since the models got really capable. People need time to adjust, even if I would expect devs to be more up-to-date than most. People’s willingness to argue about technology they’ve barely used is always bewildering to me though.
- mcswell 6mo agoYou know this how?