3 ms·
An LLM finds a dubious bug, an LLM turns it into a convincing report, and now the proposed solution is to have an LLM triage it? There are a lot of turtles hold
by iepathos 2mo ago
An LLM finds a dubious bug, an LLM turns it into a convincing report, and now the proposed solution is to have an LLM triage it? There are a lot of turtles holding up this approach and the circular logic seems hard to miss.
Automated triage can filter obvious spam, which was already fast and easy for humans to do. The hard part is independently reproducing a plausible finding and assessing its actual impact. If LLMs could already do that reliably, then the slop report problem wouldn't exist in the first place.
- charcircuit 2mo ago>If you could build the thing they wanted to build it would fix the slop problem It sounds like a reason to try and build it than a reason to not build it.
- H4lcyon 2mo agoI tried to build it (kind of). My team is going to use it to alert us to the most critical issues so we can hop on them before waiting for triage. It's decent at figuring out criticality but it's TERRIBLE at actually doing triage and assessing whether the report is plausibly or implausibly true. Security can be really nuanced, and from my experience so far with the model I'm using it's really bad at being skeptical enough to actually figure out if something is a legit issue with impact or not. I agree though, it would be awesome if we could get AI triage that worked.