3 ms·
How many of those automated fixes were reverted? How many introduced a new bug? What's the false positive rate on the finding agents? The post has counts for ev
by dabedee 2mo ago
How many of those automated fixes were reverted? How many introduced a new bug? What's the false positive rate on the finding agents? The post has counts for everything that went right and nothing for what could go wrong.
- staszewski 2mo ago[flagged]
- cognitiveinline 2mo agoOr people don't ask questions for which the answer is known to be "None, really." I get that many don't like what LLMs are doing to the industry, but this is just incorrect reaction to a very specific benefit that's proven beyond doubt (Security hardening). Accept it imo - LLMs are solving very large problems that have plagued software security.
- fg137 2mo agoSorry it is your reaction that is really weird. These are legitimate questions, and I have definitely seen Claude finding the wrong cause and then implementing completely incorrect/irrelevant fixes, only to find that it didn't work and need to start over. Not saying humans don't do the same thing, but LLMs are far from perfect, and it would be delusional to only talk about successes.
- cognitiveinline 2mo agoI've done security work before, and I've done it now with frontier. I've seen the difference first hand, and as the OP and so many other articles show, so have the leading experts in the world. So either they are lying, or you may not yet be seeing and experiencing what they are. If that's wierd, ok.
- fg137 2mo agoYou did not address my points or those in OP but kept repeating yours (that are not even relevant) in a handwavy way. And that's fine. You can continue to live in your bubble. Others simply have different experience, and your asserting "AI is perfect" on a forum is not going to matter in terms of everyone's own, first hand experience.
- iLoveOncall 2mo ago> The post has counts for everything that went right and nothing for what could go wrong. That's AI for you. At Amazon we have many forums to share our AI wins, but none to share AI failures or disappoinments. No wonder execs make bad decisions regarding AI, they only hear completely one-sided stories.
- Cthulhu_ 2mo agoAt my company there's a lot of discussion about AI and complaints that people run out of tokens within a day, but zero results are shown. No measurable (or measured) gains. Or nothing that people are willing to talk about, in any case.
- KptMarchewa 2mo agoAt my company the performance improvements channels has been exploding, with people claiming giant improvements in latency, throughput, and decreased cost of the services. The cost decreases itself is order of magnitude more (at annualized run rate) more than we pay for tokens. YMMV.
- jpc0 2mo agoI hope to see more of this across the industry. I definitely make use of AI but in my experience I almost always could have done it better myself, the places where I threw AI at the problem I didn't care about the results being good, only good enough. When we see memory and compute requirements for version x+1 of software decrease instead of increase I will happily say AI is the oracle people proclaim it to be.
- dgellow 2mo agoI’ve been asking since months on HN for proofs that companies using AI see positive ROI from it. So far I didn’t get a single concrete example
- inigyou 2mo ago"It can't be that stupid—you must be prompting it wrong." - Ed Zitron
- albinahlback 2mo agoAnd how many were introduced by AI?
- cognitiveinline 2mo agoIt's the opposite: humans. The most critical one (sandbox escape) has been sitting there for 13 years.
- tredre3 2mo agoThe question is how many new bugs were introduced through those AI fixes. Is it less than a human would have? Is it more? Is it the same type of bugs? I could be convinced either way and it's an interesting thing to ponder in my opinion.
- tstrimple 2mo agoThis is the thing the anti-ai zealots will never admit. Humans fucking suck at writing code. They talk about software development as if it's only ever performed by the top 1% of the top 1%. They never acknowledge that humans make mistakes. No. Humans create perfect code every fucking time while LLM's only produce slop. It's such a fucking mind-numbingly stupid position that I have to think they have never actually worked in an organization which produced code as a value. These fucking morons want to pretend a human never introduces a memory leak when it's plain as fucking day that vulnerabilities that these "expert programmers" introduce to software are rampant. But no. The anti-ai zealot likes to pretend that only LLMs ever produce bad code. The only bugs in existence are due to LLM slop and not the literal decades of fucking slop produced by humans without any LLM assistance. It's so fucking tired at this point. They will never be able to admit that current day LLMs are far beyond the median human programmer. Their opinions are fucking useless at this point. These idiots think The Daily WTF started with LLM code. They are basically anti-vaxxers wrt to their grasp on reality. No one should take them seriously to any degree.
- mf2hd 2mo agoFucking humans and their shitty creations eh?
- Seb-C 2mo agoExactly, also it lacks a lot of context as to why they did that. Possible (probable?) scenario: - Marketing: "we found and fixed lots of bugs thanks to AI" - Reality: the KPI is now to fix as many bugs as possible with the help of AI, so they used AI to search old and easy bugs in the backlog, and then fixed it manually
- simianwords 2mo agoHuh? Maybe a third scenario is that AI helped fix the bugs? Like described in the actual post you are replying to? This level of conspiracy theory is getting a bit ridiculous.
- IanCal 2mo agoI think you’re right, my base assumption is that the models can code and can fix bugs, and can code more in parallel and faster than humans at a lower cost. If Google are tackling lower value bugs with AI the number is in a way inflated compared to some utility measure (fixing a smaller number of worse bugs could be preferable) but it’s still things fixed.
- tstrimple 2mo agoThere's no reasoning with the folks who are anti-ai and claim LLM's cannot produce anything worthwhile. They are basically flat-earthers or anti-vaxxers at this point. There is literally no amount of evidence that will convince them. They will always say that LLMs cannot produce anything but garbage just like the anti-vaxxers who state vaccines just cause autism. They are both fucking morons who no one with any grip on reality should indulge.
- gblargg 2mo agoI assumed it was really just AI that found them, which allowed Google to know to fix them.
- dwroberts 2mo agoThey also don’t really provide an explanation for the big uptick in bugs found M146+ in the first place. Is that better testing? Or more bugs were being introduced in the first place?
- TacticalCoder 2mo ago> How many introduced a new bug? I'd say that one is not really an issue. In 2012 the Pinkie Pie exploit chain already required chaining 6 bugs to lead to an exploit [1]. Since then we've seen chains requiring more than 10 bugs (!). If you fix any one of those bugs, the exploit is non-functional anymore. Sorry out of luck. So if, say, for every ten bugs you fix, you introduce two new ones then it's still a very net win. Unless of course it introduces a bug so bad it becomes a simple exploit not requiring a long chain of exploits. But in the case of browsers we've only ever been moving to longer and longer chains of exploits required to pwn a browser. A great many window of opportunities are closing for dark-side hackers / north korean intelligence etc.: there were probably exploit chains still open for exploitation in April that just got closed by Google. If anything, besides the supply chains attacks in amateur-land, the world didn't stop working: projects (not just browsers but OSes too) are being hardened left and right. Using AI to find potential bugs is an amazing use case and there really aren't many downsides. > The post has counts for everything that went right and nothing for what could go wrong. I'm not saying there aren't a few downsides but the benefits are just too good to ignore. [1] https://blog.chromium.org/2012/05/tale-of-two-pwnies-part-1.html https://blog.chromium.org/2012/05/tale-of-two-pwnies-part-1....