3 ms·
When did we enter the twilight zone where bug trackers are consistently empty? The limiting factor of bug reduction is remediation, not discovery. Even develope
by Veserv 6mo ago
When did we enter the twilight zone where bug trackers are consistently empty? The limiting factor of bug reduction is remediation, not discovery. Even developer smoke testing usually surfaces bugs at a rate far faster than they can be fixed let alone actual QA.
To be fair, the limiting factor in remediation is usually finding a reproducible test case which a vulnerability is by necessity. But, I would still bet most systems have plenty of bugs in their bug trackers which are accompanied by a reproducible test case which are still bottlenecked on remediation resources.
This is of course orthogonal to the fact that patching systems that are insecure by design into security has so far been a colossal failure.
- reactordev 6mo agoThat might have been true pre LLMs but you can literally point an agent at the queue until it’s empty now.
- batshit_beaver 6mo agoYou literally cannot, since ANY changes to code tend to introduce unintended (or at least not explicitly requested) new behaviors.
- reactordev 6mo agoI’ve had mine on a Ralph loop no problem. Just review the PR..
- k_roy 6mo agoWhich still means a single person with Claude can clear a queue in a day versus a month with a traditional team.
- worthless-trash 6mo agoYour example must have incredible users or really trivial software.
- lll-o-lll 6mo agoEventual convergence? Assuming each defect fix has a 30% chance of introducing a new defect, we keep cycling until done?
- Kinrany 6mo agoWhy would it converge?
- saintfire 6mo agoAssuming you can catch every new bug it introduces. Both assumptions being unlikely. You also end up with a code base you let an AI agent trample until it is satisfied; ballooned in complexity and redudant brittle code.
- charcircuit 6mo agoYou can have an AI agent refactor and improve code quality.
- abakker 6mo agoBut, have you any code that has been vetted and verified to see if this approach works? This whole Agentic code quality claim is an assertion, but where is the literal proof?
- WithinReason 6mo agoIf it can be trained with reinforcement learning then it will happen
- wredcoll 6mo agoDid we have code quality before llms?
- intended 6mo agoIt’s agents all the way down - until you have liability. At some point, it’s going to be someone’s neck on the line, and saying “the agents know” isn’t going to satisfy customers (or in a worst case, courts).
- bsder 6mo agoThe fact that KiCad still has a ton of highly upvoted missing features and the fact that FreeCAD still hasn't solved the topological renumbering problem are existence proofs to the contrary.
- rybosworld 6mo agoShouldn't be down voted for saying this. There are active repo's this is happening in. "BuT ThE LlM iS pRoBaBlY iNtRoDuCiNg MoRe BuGs ThAn It FiXeS" This is an absurd take.
- array_key_first 6mo agoIt probably is introducing more bugs because I think some people dont understand how bugs work. Very, very rarely is a bug a mistake. As in, something unintentional that you just fix and boom, done. No no. Most bugs are intentional, and the bug part is some unintended side effects that is a necessary, but unforseen, consequence of the main effect. So, you can't just "fix" the bug without changing behavior, changing your API, changing garauntees, whatever. And that's how you get the 1 month 1-liner. Writing the one line is easy. But you have to spend a month debating if you should do it, and what will happen if you do.
- missingdays 6mo agoSo, you have already fixed all the bugs and now just cruising through life?
- jen729w 6mo agoI wonder whether people like you have actually used Claude for any length of time. I use it all day. I consider it a near-miracle. Yet I correct it multiple times daily.
- rybosworld 6mo ago> I wonder whether people like you have actually used Claude for any length of time. I stated the LLMs are actively being used in repo's today, to chew through backlog items, and your response is to wonder if I've ever used Claude. To me it's surprising that someone like you, who appears to have a reading comprehension deficiency, is able to use Claude.
- bawolff 6mo agoBugs are not the same as (real) high severity bugs. If you find a bug in a web browser, that's no big deal. I've encountered bugs in web browsers all the time. You figure out how to make a web page that when viewed deletes all the files on the user's hard drive? That's a little different and not something that people discover very often. Sure, you'll still probably have a long queue of ReDoS bugs, but the only people who think those are security issues are people who enjoy the ego boost if having a cve in their name.
- kackerlacker 6mo agoEh, with browsers you can tell the user to go to hell if they don't like a secure but broken experience. The problem in most software is that you commit to bad ideas and then have to upset people who have higher status than the software dev that would tell them to go to hell.