5 ms·
From what I've seen in the bounty-related subreddits, AI is flooding bug bounty inboxes with low-value or meaningless reports, or straight-up hallucinations whe
by mapmeld 2mo ago
From what I've seen in the bounty-related subreddits, AI is flooding bug bounty inboxes with low-value or meaningless reports, or straight-up hallucinations when people use smaller models (to turn a profit, you make lots of low-value bug reports and see who pays out).
This has a negative effect on humans doing their work with or without LLMs: curl shut down their bounty program, and GitHub just announced they're "restructuring" theirs.
The author of this post also makes a case that HackerOne hasn't been honest about LLM training and use, either to hackers or to their own staff.
- charcircuit 2mo agoDoesn't that problem benefit from having automatic bug triage that can avoid fast tracking these bad reports?
- iepathos 2mo agoAn LLM finds a dubious bug, an LLM turns it into a convincing report, and now the proposed solution is to have an LLM triage it? There are a lot of turtles holding up this approach and the circular logic seems hard to miss. Automated triage can filter obvious spam, which was already fast and easy for humans to do. The hard part is independently reproducing a plausible finding and assessing its actual impact. If LLMs could already do that reliably, then the slop report problem wouldn't exist in the first place.
- charcircuit 2mo ago>If you could build the thing they wanted to build it would fix the slop problem It sounds like a reason to try and build it than a reason to not build it.
- H4lcyon 2mo agoI tried to build it (kind of). My team is going to use it to alert us to the most critical issues so we can hop on them before waiting for triage. It's decent at figuring out criticality but it's TERRIBLE at actually doing triage and assessing whether the report is plausibly or implausibly true. Security can be really nuanced, and from my experience so far with the model I'm using it's really bad at being skeptical enough to actually figure out if something is a legit issue with impact or not. I agree though, it would be awesome if we could get AI triage that worked.
- uqers 2mo agoDidn't Daniel later report that curl recently started getting mostly high-quality LLM reports on their bounty program? I can imagine that there would definitely be a few "bounty spammers" trying to get hits, but it seems like most of them are doing good work. I'd say instead that the problem is that a lot of people don't care anymore about the quality of the work being done, and LLMs are accelerating it. Bounty programs have shifted from ways for people to report security bugs to ways for people to try to make money.