36 ms·
Google fixed more Chrome bugs in June than over the past two years, thanks to AI
- heiogalma 2mo ago[dead]
- cyberjunkie 2mo agoStill can't stand Chrome.
- geek_at 2mo agowon't forgive them either for making it harder to filter out ads
- onesandofgrain 2mo ago[flagged]
- edhelas 2mo agoWe don't.
- echelon 2mo agoYou see, grandma is supposed to see the ads and buy the Mighty Putty. I'm supposed to get the adblock version. That's just how the world works.
- flohofwoe 2mo agoThat's something those products need to figure out, not the users, one potential solution is to erect paywalls and find out how much those products are actually worth to their users ;)
- 83642736392 2mo ago[dead]
- onesandofgrain 2mo agoThat's my point.
- mw888 2mo agoAI critique often funnels itself into a narrow bucket: creating code blindly with AI is bad. That's easy to grant. Adversarial testing, checking developer assumptions, refactor suggestions, small dev tools and even some guided coding all sit on the other side of the spectrum of what you can do with coding and AI. For larger and larger codebases even simple things like tracing dependencies or behavior might be greatly aided. And the critiques reserved for that narrow bucket on the other end, blindly generating code, are too easily conflated with the rest.
- ikekkdcjkfke 2mo agoYes. AI is a tool, it’s supposed to be used a certain way, anything else is a misunderstanding og what AI is. You have to aim it in the direction you want it to go, not expect it to solve all your problems magically
- cubefox 2mo agoLLMs are increasingly not tools anymore but agents. The difference is still one of degree, but it's clear in which direction we are currently moving. It was even clear to some people 10 years ago: https://gwern.net/tool-ai https://gwern.net/tool-ai
- danielbln 2mo agoThat article is surprisingly prescient.
- layer8 2mo agoAI isn’t “supposed to be” anything (other than “intelligent”). It just turned out that LLMs have to be used in particular ways to be useful. Nobody planned or designed them to be that way. We all as users are put to do the job of discovering the ways.
- ramanbanka 2mo agoAI has skills no individual can have. That part is true. But AI is not smarter than you, AI is as smart as the user who uses it. It often makes wrong decisions unless user corrects it.
- truncate 2mo agoNot that I don't believe its possible to fix a lot of bugs, I also wonder what the actual dynamic was. Were the people in team working much more than usual as well? Given its Google, I wouldn't be surprised if there was an "internal push" to fix more bugs over next X sprints so that they can publish this blog and some manager can show impact and AI adaption to his superior.
- brador 2mo agoMore likely just getting ahead of the AI attacks before they hit. The threat risk increase caused by AI has gone off the chart.
- deeringc 2mo agoExactly this. And there are few bigger targets than Chrome when it comes to finding exploits (OSes and network equipment are probably on par). I'm sure they have devoted large compute resources and human staffing at making sure that they find and fix these issues before anyone else does.
- nevi-me 2mo ago1. Our backlog of bugs gets processed quicker because instead of staring at the code for 10 minutes fiuring out what's happening, there's a tool that can reason about it quicker. 2. Code reviews and security reviews happen quicker and produce more findings. I would think that (m)any team(s) using AI might also be seeing a higher rate of finding and fixing issues. Even the Linux Kernel (I'd say Windows and Apple too) are seeing the same phenomenon.
- Supermancho 2mo agoLinus: "it keeps finding embarrassing bugs" Linux Kernel: https://lore.kernel.org/all/CAHk-=wi4zC+Ze8e+p3tMv8TtG_80KzsZ1syL9anBtmEh5Z40vg@mail.gmail.com/ https://lore.kernel.org/all/CAHk-=wi4zC+Ze8e+p3tMv8TtG_80Kzs... The idea that software has gotten so complex that a machine can evaluate code paths better than a human, seems to bristle the fur of many. Some people didn't think we would see the day where that comparative human limitation was laid bare in simpler tasks than they expected. I believe older developers are less likely to be offended, having to deal with this as a matter of course (as the mind declines).
- ed_mercer 2mo agoAI meaning Claude or Codex, because Gemini is a laughingstock for coding. Also Gemini was mentioned only once in the article.
- greenknight 2mo agoThe Gemini (and Claude / Codex) that we have, is much different to their internal models / harnesses.....
- Oras 2mo agoTrust me bro?
- Cthulhu_ 2mo agoKind of, but you can also assert it yourself; the big tech companies have billions of lines of code not publicly available, they can train their in-house code assistants on those instead of or in addition to what's available out in the open.
- brainwad 2mo agoThe latest and greatest internally at Google is more or less just Antigravity 2.0 with the latest Gemini models.
- MadsRC 2mo agoI’ve heard from people inside Google that they’re prohibited from using anything but their own models. AlbertaTech on YT, ex-Google working on YT said in a video that right before she quit someone had threatened to quit if they couldn’t get access to Claude Code
- lilerjee 2mo agoMore important thing: Do you import new more bugs?
- kotaKat 2mo agoSeems like they have enough AI they could find a way to put ManifestV2 back into the browser and fix things we actually wanted instead of putting more bullshit "AI improvements" we didn't?
- edhelas 2mo agoIf only there was alternative browsers...
- deleted 2mo ago[deleted]
- TuxPowered 2mo agoSo will the Chromium bug 5569 finally get fixed?
- Cthulhu_ 2mo agoI'm not sure why you're asking here instead of opening the ticket and having a look at the current status.
- SweetSoftPillow 2mo agoIt's not a bug but a suggestion rejected by the team (cuttent status: Won't fix, obsolete)
- dabedee 2mo agoHow many of those automated fixes were reverted? How many introduced a new bug? What's the false positive rate on the finding agents? The post has counts for everything that went right and nothing for what could go wrong.
- staszewski 2mo ago[flagged]
- cognitiveinline 2mo agoOr people don't ask questions for which the answer is known to be "None, really." I get that many don't like what LLMs are doing to the industry, but this is just incorrect reaction to a very specific benefit that's proven beyond doubt (Security hardening). Accept it imo - LLMs are solving very large problems that have plagued software security.
- fg137 2mo agoSorry it is your reaction that is really weird. These are legitimate questions, and I have definitely seen Claude finding the wrong cause and then implementing completely incorrect/irrelevant fixes, only to find that it didn't work and need to start over. Not saying humans don't do the same thing, but LLMs are far from perfect, and it would be delusional to only talk about successes.
- cognitiveinline 2mo agoI've done security work before, and I've done it now with frontier. I've seen the difference first hand, and as the OP and so many other articles show, so have the leading experts in the world. So either they are lying, or you may not yet be seeing and experiencing what they are. If that's wierd, ok.
- fg137 2mo ago
- VBprogrammer 2mo agoI've recently been using AI a lot for performance optimisation during a particularly busy period at work. I would say it was almost completely useless at the high-level direction - it would point out suspicious parts of SQL queries for example but on back to back testing these almost never resulted in any performance change. In fact, if it wasn't for the fact that it made making the actual changes I identified much easier (move these joins into a CTE etc) it would have been a detriment. Not only did I get sidetracked by a bunch of useless suggestions but I also had to put up with others dumping their raw AI output at me as if it was somehow a meaningful contribution.
- colechristensen 2mo agoThen you're not grounding it to reality properly. LLMs get bad pretty quickly when you just tell them to pull the answer out of thin air. Pin them to reality with actual performance tests to run to test theories and the results will be much better. Performance optimization involves simulating real world loads and measurable results. Give them the opportunity to actually have a closed loop if you want more than trivial improvements.
- maccard 2mo agoHave you any example materials that show a harness that an agent can be pointed at an app with some sort of telemetry tool, the results it gained and the cost of doing so? Because my experience is the same as the parents - the LLM goes on massive tangents, and the more tangents it goes on the worse the results get.
- colechristensen 2mo agoUltimately the harness is me and the experience is like managing an unruly toddler. You have to pay attention to what it does and issue corrections. Skills and prompts and AGENTS.md do some, nested sets of agents do some, but over it all is me keeping track of what it's doing and needs to do. You have to size the unit of work to its useful attention span, you have to have code architecture that is conducive to units of work, and you have to have tools like ticket managers, etc. that keep the big picture and smaller pictures in mind. Reverting from agile practices to more pre-planning whole project documentation and architecture decisions helps. But in the end it's you. LLMs have their limits and need humans to direct them.
- asdaqopqkq 2mo agoDid anyone check the code changes manually? What if it is just slop code?
- brainwad 2mo agoI don't know about Chromium specifically, but in general Google requires 2 humans to have looked at every change (normally one is the author, but not necessarily for bot-authored changes).
- Orphis 2mo agoThey did roll this policy for Chromium changes a few years ago too. And it wasn't just 2 humans but 2 Googlers. One Googler author and one Googler reviewer works. But an external author needed 2 Googlers to review changes, and all files needed to be reviewed by one author at least. So if you touched multiple files owned by different people for a more complex feature, you would naturally get there.
- Cthulhu_ 2mo agoDid you go and look at their open source repositories and open code review tools instead of fearmongering on HN? The answer can be found.
- surajrmal 2mo agoGoogle has not really embraced vibe coded slop. Humans are involved before any code merges and the quality bar remains high.
- feverzsj 2mo agoMaybe fix all the decade-long old bugs first.
- glimshe 2mo agoA lot of people here seem to be living in a different universe than me or simply don't know how to work with AI. I think detractors believe you should just let AI do the job blindly instead of leveraging it as a tool to accelerate you. They get mad at Excel for the poor investment returns. At this point, this is such a strawman, it isn't worth counter arguing. I think I'll abandon this discussion and keep using AI quietly while exchanging tips with like-minded people who are interested in using it properly and efficiently.
- maccard 2mo ago> A lot of people here seem to be living in a different universe than me I feel this way on this topic too. > I think detractors believe you should just let AI do the job blindly instead of leveraging it as a tool to accelerate you. The problem is; how _should_ I use AI? On a previous thread, I had two replies to the same comment, one saying "provide the LLM all the context it needs and let it go ham", and the other saying "Carefully guide and craft it and review everything". Both subthreads had people agreeing. My experience with LLMs is leaving them unattended gives very poor results within a handful of iterations, and that carefully guiding them can be a value add, but the effort ot do that is very often as much as just writing the damn code myself, particularly if there's a few iterations that go on.
- icepush 2mo agoThe reality of the current situation is that we have dozens of families of different model; each of which is differentiated further into various revision numbers and power levels, and depending on which one you use and what the nature of the tasks is, you very much need a different usage style that can only be discerned after multiple (Sometimes very many) sessions of using that individual model for that category of task. It is not one size fits all, and trying to pretend that it is will lead to failure.
- maccard 2mo ago> It is not one size fits all, and trying to pretend that it is will lead to failure. I'm not asking for a one size fits all solution. I'm asking for "vibe coding with Codex on GPT 5.6 with high is the way, but if you want to be more ivnolved, Sonnet 4.6 on Claude is the path. Here's the proof." Saying that it depends and it's impossible to quantify allows the camp who are claiming it's more productive to say "you're holding it wrong".
- paganel 2mo agoA company valued at $4+ trillion based on the recent AI hype alone is touting the benefits of using said AI, and we're supposed to take this shit for granted. I'm wondering what the Alphabet employees still commenting on this forum have to say about it? I guess for the right comps they can keep their mouths shut no matter the high level of idiocy involved.
- surajrmal 2mo agoThere is hype, but there are also real tangible wins as well. Analyzing code for bugs is something that works very well. We use it for helping us revert changes that lead to a flaking test and it's dramatically faster than previous techniques. It's great at suggesting duplicate bug reports. It's amazing at scaling an API migration throughout a codebase. It's great at prototyping an idea before you decide to do it properly. There are many things it's not good at (eg design and architecture), but that doesn't invalidate the things it is good at. Everyone is learning how to best use it and there are growing pains. Ignore the hype and try to find the common patterns in what people say.
- shevy-java 2mo agoThe biggest bug still has not been fixed here, which is: Google. We really need an alternative to this greedy and evil corporations de-facto controlling a huge portion of the modern www stack. All decision-making processes are here subjugated to what fits Google's adEmpire. This is a perpetual system of control amplifying abusive systems in place. We already see this with the recent age sniffing; Google just yesterday announced that all android users must hand over their age to everyone else. Before that google kicked out or locked out open source alternatives to android (or, at the least, made their life significantly harder than before). These are not isolated ways of abuse - this is systematic abuse. We really need alternatives to Google here. It can not be that one tyrant company dictates so much of people's open lives (and no, Firefox is no alternative; Mozilla is basically a nerfed and bribed entity that has given up on firefox many years ago already; ladybird may become an alternative at one point in time, but right now is not and there are questions over how Kling handles the overall project, but these complaints are significantly below what Google is doing globally here). Fixing more bugs in the adChromium code base does not change the underlying problem at hand.
- Cthulhu_ 2mo ago> We really need an alternative to this greedy and evil corporations de-facto controlling a huge portion of the modern www stack. We do / have? Multiple even. Of course, the vast majority of consumers doesn't care so they don't get as much traction.
- Razengan 2mo agoMe too! It's actually tiring! In a good way I've mostly only used Codex for reviews, for Godot/GDScript code, and it helped me catch bugs that would've taken me ages to even notice on my own; games be tricky like that. Claude was mostly useless or annoying up until the last time I tried it (about 3 months ago) You have to be vigilant though: Codex often picks out obscure edge cases and suggests adding multiple new functions and flags to avoid issues that would be better off left as fast-failure crashes. Maybe that's because of the 5.6 Sol Max I leave it on.
- iepathos 2mo agoNice marketing Google, now tell us how many of the bugs that were fixed were actually introduced by AI usage to begin with?
- goldenarm 2mo agoElephant in the room : how many of these bugs were written by LLMs in the first place ? Because creating 100x more bugs and fixing 100x more isn't something to be proud of.
- Cthulhu_ 2mo agoCan't have been much, as the Chrome project is 20+ years old and LLMs have only recently become a thing; they did not slacken their code review and testing practices with the advent of LLM code generators. And I doubt they do much development on the affected areas at the moment, they mention a 13 year old issue for example. But it's an open source project, you can go and figure out whether your assertion is correct.
- goldenarm 2mo agoThe single 13yo issue is anecdata. They obviously have the full git blame statistics but chose not to include them in the blog post, which is a bit concerning.
- rpdillon 2mo agoIt's weird to assume that this is all smoke and mirrors where AI is fixing AI's bugs. When we dove into the rsync issue, what was happening is that AI was revealing bugs that had been previously unknown, resulting in the need to fix them. I would be stunned if that were not the case here.
- lbrito 2mo ago>Can't have been much, as the Chrome project is 20+ years old and LLMs have only recently become a thing That's not a great argument. GPT3.5 was launched over 4 years ago, so that's already significant compared to Chrome's age. LLMs notoriously also produce more verbose code than a human, so its not out of question that it might produce 5x more bugs on average given the same time frame.
- llm_nerd 2mo agoThat is a stat you invented out of thin air, based upon nothing but pejorative hope, against all logic or rational consideration of the timeline or the scale of the project, and you declare it the "elephant in the room"? HN is so weird on the topic of AI, and so many desperately are trying to contrive a reality. Odd stuff.
- mlacks 2mo agoI am actually using AI to learn about Windows OS performance. I have a couple of prompts scheduled to research, test, and implement performance gains in the OS, then post them here [0] and here [1]. I always admired Brendan Greg and Bryan Cantrill at Sun/Illumos/ for their work on OS performance but never had the time to sit down and catch up to their level due to my time in the military. I got Codex to find and apply a bunch of settings I didn't know about and in the future I hope to use AI to make real changes in performance that venders probably don't have time or resources for. that being said, what I've learned so far is essentially WhyIsItAlwaysHN's entire comment "The place to find them would be performance profiles, query plans, telemetry." I just find it so frustrating that for a user interface that was basically solved in the 90s, Windows GUI still struggles to survive on hardware exceptionally more powerful than what we had 30 years ago 0 - https://www.lacksan.com/updates/ https://www.lacksan.com/updates/ 1 - https://github.com/Lacksan-Dev/HP-ZBook-Performance https://github.com/Lacksan-Dev/HP-ZBook-Performance
- galkk 2mo agoGood. Now please fix youtube. Amount of bugs that I'm getting in youtube app on iphone should be embarassing for company like Google. I cannot believe how bad it become in last couple years.
- nalekberov 2mo agoReal question is: how much those "fixed" bugs actually bothered the users to begin with? But since Google cares mostly about its investors, the numbers and a mention of Gemini in such a blog post are more important.
- why_only_15 2mo agoThe whole point is that these are security bugs that didn't necessarily bother people so far but could be exploited
- Orphis 2mo agoOld bugs at the bottom of the prioritization list deserve to be fixed. Some were probably just theoretical "if we have over 2B of this item, it'll crash because X/Y/Z and that can be exploited", some are probably in rarely used features. Once the LLMs are done going through the code, the baseline will be all better and they will probably fix a lot less than they have now. Or they will have LLMs automatically check every single crash bug that has been submitted and take care of those too. I am assuming that they are probably already using this as an input source for their AI.
- cyral 2mo agoDidn't expect to see the goalposts moved to "AI can fix bugs but do they even matter"
- seanmcdirmid 2mo agoThis is only a flex if AI also didn’t cause an increase in bugs that needed to be fixed.
- buzzin__ 2mo agoSo, your conjecture is that the LLM's skill to detect bugs somehow magically disappears the moment they start writing new code?
- ryukoposting 2mo agoActually yeah, that would align with my experience with it. You're absolutely right! Let me fix that.
- williamdclt 2mo agoI wouldn't say "magically", it's what happens with me and every other engineer I know: we're all detecting bugs and yet we're all writing them too. Not saying that LLMs work the same as humans, but saying it's not a logical fallacy.
- seanmcdirmid 2mo agoNo, the extra code velocity allowed by AI without proper QA support means there could have been a lot more bugs introduced into products over the last year or so, giving AI more bugs to find. It’s tongue in cheek but not completely uncalled for. It’s been long known that one dishonest way to inflate your bug fix count is to put more bugs into the product to fix. Goodhart’s law is merciless.
- greenimpala 2mo agoa lot of these bugs wouldn't have ever been found without AI
- jongjong 2mo agoBut how many new ones did it introduce?
- ymolodtsov 2mo agoI wish they'd spent more time on improving the UI now. It's largely been the same since Chrome appeared. Their recent vertical tabs are so much inferior to Dia, which is built by a much smaller team.
- luciana1u 2mo ago[flagged]
- fuadnafiz98 2mo agoSad they didn't say `Thanks to Gemini`
- cubefox 2mo agoConcerning even
- sevenzero 2mo agoBig AI player promotes AI through blog to make AI look better woo
- diwesh871 2mo ago[dead]
- ahartmetz 2mo agoMaybe they could ask the AI to fix their atrocious build times. Only in Chrome do you have 300 line source files that blow up to 20 megabytes after preprocessing as a matter of course. 3 GB per compile job is getting more common - and RAM is expensive right now!
- benob 2mo agoThe next step is to build software without bugs
- glauber 2mo agoConvenient that the LLM vendor is the one saying their own product works, isn’t it?
- ChrisRR 2mo agoCompany advertises their products. In other news, water is wet
- contagiousflow 2mo agoThat's why you should be critical of their press releases.
- brtkwr 2mo agoI assume this was thanks to Project Glasswing + Mythos Preview?
- ppljudge 2mo agoI’d rephrase this to something like: “People fixed more Chrome bugs (…), thanks to people who leveraged AI tools”
- lapcat 2mo agoI can't comment on Chrome bugs, but I can comment on Safari bugs. According to a search of the WebKit Bugzilla, I'm the most prolific filer of web inspector bugs outside of the WebKit team itself. Recently I noticed an unusually high number of web inspector bugs fixed by one longtime WebKit engineer, and I suspected AI assistance, though the commit messages include no disclosure of this. Nonetheless, the increase in volume was quite dramatic. The other day, an update to Safari Technology Preview was released, and I attempted to verify whether a couple of my bugs were indeed fixed. However, it was impossible for me to verify, because the latest version of Safari Technology Preview introduced a new web inspector bug that totally broke the features I was testing. Thus, I remain unimpressed. Without a doubt, LLMs have demonstrable skills and can produce code much faster than humans. I never thought that producing code fast was wise, though, even before LLMs arrived. For many years I've criticized software development based on management-driven release schedules, where developers are forced to pump out code regardless of quality, regardless of whether it's ready. Your boss may not like it, but code is done when it's done, not when your boss says it has to be done. If software complexity were predictable and reproducible, then indeed management could replace engineers with automation, but that's not how it works in reality. Back to Chrome, the browser is notorious for continually adding invasive new "features" to the web that nobody wants except advertisers. I wonder how many Chrome bugs were the result of Google push push pushing all of this new crap on users over the years?
- Tistron 2mo agoI have a software degree and some work experience but have been doing other things for a decade or so. I'm curious to try out some "vibe engineering" but it seems a bit daunting to get started, there are so many new tools available on top of the AI itself. Are there good resources for somebody like me to get started, that'll guide me though how to think about prompting, and using CI and how go have the agent successfully write specs and tests and what-not. I don't mind paying some for a good course, but I find it hard to figure out which one would be worthwhile, or whether there are some youtubers that would be better to follow. Any tips?
- pigpop 2mo agoThe best option at this point is to just sign up for a paid plan with either ChatGPT or Claude and then ask the model the same thing. My preference would be for ChatGPT and if you've been out of the game for a long time then using the desktop app might be the best choice https://chatgpt.com/codex/ https://chatgpt.com/codex/ Then try starting with voice mode (if you're comfortable chatting out loud) and just talk your way through it.
- ssl-3 2mo agoJust jump in. Build a sandbox, download Codex CLI or Claude Code or whatever and spend some time doing some creative stuff with it just for the sake of learning how it all fits together. Pay attention to the inputs you provide and the outputs that they result in. Keep your bullshit detector engaged: Bots often lie. If you get stuck, or it gets stuck, or you want prompting advice or whatever: Ask any frontier-level bot for help. Sometimes, it's very instructive to get help from Claude for an issue with Codex, or vice-versa. Want better tools? Ask the bot to suggest some that exist. (None of the existing tools fit? Have the bot write new ones.) All of this stuff is always in a state of flux, so even with the lies they'll do better at teaching than any fixed reference will. They're LLMs and processing language is what they tend to be best at... so go ahead and use that. And remember: They're designed to behave kinda-sorta like humans, but they are not humans. They're just computer programs. If you don't like their output style, or they don't like your input style: Ask them how to implement rules that make them knock that shit off. :)
- andai 2mo agoThis situation seemed familiar... https://i.ibb.co/RGvQgfmX/i-fixed-more-bugs.png https://i.ibb.co/RGvQgfmX/i-fixed-more-bugs.png
- whywhywhywhy 2mo agoThey keep breaking inertial scroll on Mac though, which is something that irritates me thousands of times a day. First they made it so when it came to rest it would then start scrolling again and overshoot by 2 lines before coming to a halt. They fixed that after a few weeks then it worked until last week and now it’s left in a state where as it’s coming to a halt it jerks every time like dropping frames of the scroll almost. I know it’s not my machine, config or trackpad because I’m seeing it on multiple computers.
- rarestoma 2mo ago[dead]
- cahoot_bird 2mo agospent a several months working on their bug bounty program earlier this year while working full time and lost a lot of motivation after a duplicate combined with lower bounty awards and claude safety guards all at the same time
- brazukadev 2mo agoThis makes zero sense with the amount of engineers they have, most of those bugs even would be fixed for free by highly competent people.
- josefresco 2mo agoDid they fix spell check? Because it still doesn't work and all the solutions are basically "wipe and start over".
- Phemist 2mo agoWorrying. Extrapolating a (speculative) future, this means (Google will feel) that soon the chromium base will no longer need the crowd-sourced bug hunting that is open-source. I expect Google to eventually stop working on chromium (in the open) and all current chromium-flavours will become de-facto forks of the last published version of chromium. These forks won't be equally easy to maintain given that the groups running them do not have access to the same level of subsidy as Chrome does with Gemini.
- dybber 2mo agoGoogle already controls Chrome/Chromium direction so much that you should use Firefox if you value an open web.
- Phemist 2mo agoWhich I do (actually a fork of firefox - waterfox), but it is still worrying as the viability of chromium-base browsers is what's keeping us having "merely" a chromium-monoculture (with viable ad blockers e.g. and website operators that check compatibility with more than one browser), rather than a chrome-monoculture (where ad-blockers surely would have been killed by now).
- Alpha3031 2mo agoMicrosoft surely has enough resources to keep their fork going if they choose to, so I see no reason why Google would try that given they would likely end up with less control over the web.
- shepherdjerred 2mo agoConversely, it has never been more feasible for alternative web browsers to be created and maintained outside of big tech.
- fragmede 2mo agoYet, no one's taken Cursor's fastrender to production.
- mnmnmn 2mo ago[dead]
- Arshad-Talpur 2mo agoThe article is great reflection how LLMs are actually doing Quality assurance on both security and codebase sides, but i still doubt the creative part that means UI side, So far in my experience LLMs at best are unable to determine what the good UI is , may be its not inherently a bug detection rather a creative process
- gverrilla 2mo agoEarlier today I opened Chrome on my iPhone, went to the Reading List. It started flicking fast and then the phone restarted. That's a first for me on iPhone.
- eviks 2mo agoNo worries, writing with AIs can improve the speed bugs enter the code base as well. Would be interesting to know the breakdown. > the “latent security issue.” Code that is safe and robust in isolation can be transformed into a critical vulnerability by an entirely unrelated, minor logic change elsewhere in the tree. Similarly, you can accidentally fix a bug without discovering/triaging it, so the pointless repeated bug lifecycle is incomplete
- xacky 2mo agoJust admit that Chromium is too complex at this point, we've all seen show HNs doing what could be considered magic in Chromium. We probably need to issue a mortarium in new platform features and focus on a deep clean in bug fixes. I've said that Mozilla is too busy doing redesigns than bug fixes as well.
- pshirshov 2mo agoDid they fix that bug with ublock not working properly on Chrome?
- running101 2mo agoI recently switched from Chrome to Edge after using Chrome for decades. It is so bloated and slow. I can have same number of tabs open in both and Edge is always 30% to 50% more efficient.
- lr1970 2mo ago> Google fixed more Chrome bugs in June than over the past two years, thanks to AI And introduced how many new bugs? thanks to AI ...
- Conol_ai 2mo ago[flagged]
- adam12 2mo agoInteresting to see how many new bugs were created from all these changes.
- throwatdem12311 2mo agoAnd how many vulnerabilities did they introduce ;)
- maelito 2mo agoIt's so hard to believe these kind of affirmations coming from a firm that has so many hundreds (thousands ?) of billions invested directly and indirectly in generative AI...
- qarl2 2mo agoLies, obviously. AI is worthless.
- hndbwksam7 2mo ago[dead]
- reloty 2mo agoThe privatization and obfuscation of the security industry is worrying. Google claims to help open source developers, but donates millions to the Alpha Omega Foundation that did not give access to Mythos to Curl and only found one issue. The Alpha Omega people appear to be selected on physical appearance and take the money away from open source without doing much. The entire Chrome article is obviously directed at suits and unsurprisingly hypes AI for better promotion chances.
- deleted 2mo ago[deleted]
- Kiro 2mo agoI can believe it. So many small CSS quirks that have bothered me for years have suddenly been fixed, meaning I can remove my work-arounds.
- rcstank 2mo agoLike what?
- 4lx87 2mo agoTitle is inaccurate. Fixed more “security bugs” not “bugs”. The article is specifically about security vulnerabilities.
- xnx 2mo agoTitle is accurate if imprecise. Security bugs are a subset of bugs.
- iamgopal 2mo agowith excellent test cases, AI can do wonders, human's task in AI world is to just write as many tests cases as we can.
- cshores 2mo agolol given the state of Gemini at the moment, they probably used Claude.
- iknowstuff 2mo agoThis is a tacit admission it was Mythos, not Gemini, despite alluding to experiments with Gemini earlier on: > Added support for model interoperability to leverage the unique strengths of both open-weights and proprietary models
- azornathogron 2mo agoNeither Mythos nor Gemini are open-weight models (right?) so I'm not sure what model(s) this is actually referring to.
- dbmnt 2mo agoI'm also curious if they used Mythos. Google definitely had access to it via Project Glasswing. However, your logic doesn't hold. That sentence only tells us they were using multiple models, but offers no real clues as to which proprietary model(s) they were using.
- naveen99 2mo agoWeird, why would they use open weight models.
- unprovable 2mo agoThe real datapoint was Firefox not paying any money in Berlin's Pwn2Own competition round this May just gone. Unheard of to have nothing confirmed... they've paid out every event since 2007 (I checked). Does this mean we must move past the low-hanging fruit now? Probably... Certainly indicates some usefulness of these models.
- ayewo 2mo agoPerhaps this was due to their red-teaming partnership [1][2] with Anthropic which they wrote about a few months earlier in March? 1: https://www.anthropic.com/news/mozilla-firefox-security https://www.anthropic.com/news/mozilla-firefox-security 2: https://blog.mozilla.org/en/firefox/hardening-firefox-anthropic-red-team/ https://blog.mozilla.org/en/firefox/hardening-firefox-anthro... Previous discussion: https://news.ycombinator.com/item?id=47273854 https://news.ycombinator.com/item?id=47273854
- MostlyStable 2mo agoI just did a search and apparently this fact (the specific one about no payouts for the first time in almost 20 years) has not gotten a discussion on HN. Given the degree of skepticism around the utility of AI bug finding and fixing (this very thread is full of it), I would have thought that concrete evidence that it can help actually make real software more secure against attacks would have gotten a write-up somewhere.
- warkdarrior 2mo agoWe KNOW AI is bad, so why would we have a use for "concrete evidence that it can help"??
- MostlyStable 2mo agoI'm honestly unsure if this is a Poe's law thing or not. I'm going to go ahead assume that you are doing the honorable thing of purposefully not including a /s for the integrity of the joke.
- MSkill1 2mo agoDid they fix the bug where Chrome tries to track you and your behavior wherever you go and whatever you do in the world?
- righthand 2mo agoSo Google could have been using AI to fix the XSLT bugs instead of acting like a cult and forcing everyone to remove it? What does this mean for things Google deems too bug riddled to be a viable feature? Will they remove the HTML surface now or fix it? I would love for the geniuses on the WHATWG to answer.
- paulddraper 2mo agoThe current title says "Google fixed more Chrome bugs in June than over the past two years, thanks to AI." This is wrong on almost every count. 1. The numbers are by milestone not month. 2. The latest milestone (or month) did not exceed the previous two years of milestones. 3. The only thing discussed are security bugs, not bugs in general. What is true: the article is about Chrome.
- khanhnguyen8386 2mo ago[flagged]
- seemaze 2mo agoMy fear is that on average all software will get worse and be less optimized. There will still be high quality products at a high cost with high touch humans shepherding development, but I think there will also be a bunch of 'good enough' slop for everyday functional tasks where there exists neither the talent nor the budget to create truly wonderful solution. Fast foot didn't kill fine dining, but it is killing us..
- simultsop 2mo agoWhen I see this, and I recall that same guys mobile OS been acting crappy software wise (10fold). I find this as a joke.
- lbrito 2mo agoOkay, now how many bugs would have been fixed if Google had spent $130B (Google's AI capex) in human resources instead of AI? Just to have a sense of proportion, considering an average software engineer salary of $200k, that same money would buy you 650,000 engineer-year salaries. Of course there are other expenses, but that gives you an idea of the order of magnitude.
- branko_d 2mo agoOn a related note: Visual Studio still has dialogs that cannot be resized, even after years of people begging Microsoft to make them resizable. Shouldn't this be a trivial fix for AI? I'll believe in AI when I get to resize my Configuration Manager! https://developercommunity.microsoft.com/t/Resize-configuration-manager-window/482335 https://developercommunity.microsoft.com/t/Resize-configurat...
- AussieWog93 2mo agoHonestly, there's a 50-50 chance that you'd be able to do this in Claude today, even without access to the source code. Just open up a configuration window manually, tell if that the Window is open and it can pull apart the executable from there. Biggest impediment would be Fable safeguards if it tries to decompile.
- ryaniscool 2mo agoWhat I like is that this is a measurable productivity increase that can be directly linked to AI and they go over their methodology. AI-positive posts skew heavily anecdotal and hyperbolic.
- voska 2mo agoSo it uses less RAM now, right? ... right??
- ghm2199 2mo agoCode related Vulnerability discovery is one of those bright spots where AI can shine because its shaped as a learning-test: every example it runs gives it feedback that is deterministic. If you have ever pointed claude/codex to your a binary executable and asked it to figure out something it can run against full throttle, I fond it always gets back to you with some good insights.
- hn_submit 2mo agoTo me this merely signals how broken C++ development really is. Most if not all of the bugs being uncovered are memory related and therefore intimately tied to the mental memory model of C and C++, namely manual memory management. It's fine for a C or C++ program encompassing a couple hundred lines but beyond that it's a liability. C and C++ are simply not fit for purpose when large scale software projects are concerned. All of these need to be ported to Rust or another memory-safe language ASAP to prevent mayhem. The hundreds if not thousands of developers working on Chrome weren't idiots who didn't know what they're doing. The complexity of programming in C/C++ is simply beyond most intelligent individuals' ability to get perfect all the time.
- rhdunn 2mo agoC++ has had smart pointers for memory (and other resource) management for a long time now (see e.g. the Windows ATL classes for working with COM objects and resources). There are a number of challenges that make browsers more challenging (even in memory safe languages like Swift and Rust). 1. Back references/pointers like `parentElement` to other objects in the graph that create dependency cycles (where traditional/simple reference counting will prevent the objects being deallocated). 2. Interacting with (and creating resources in) a garbage collector when running/evaluating JavaScript code. 3. Just-in-Time (JIT) compilation of JavaScript and other complicated interpretation-compilation pipelines that can allocate and transfer objects between the different stages.
- mbac32768 2mo agoActually, if you implement a JavaScript runtime at all your choice is either too slow for modern web users or `unsafe` everywhere. Runtime values are often simultaneously either words that encode immediates or pointers to heap blocks and doing this the proper way blows up time and space. Doing it the fast way means you have `unsafe` everywhere and Rust's lifetime model can't help you at all. Everyone chooses the fast way. Maybe Rust let's you encapsulate the unsafe a little better than C++ but the wins are going to be surprisingly small. If you had to build it today you'd maybe use Rust but the case for rewriting an existing C based runtime in Rust is not that strong. Also for things like interpreter loops the absence of computed goto in Rust stable is a real performance killer. Explicit tail calls (the `become` keyword in nightly) work but maybe you don't want your big rewrite to rely on that.
- artrockalter 2mo agoIn my experience LLMs do find a lot of embarrassing bugs as Linus says but it can constantly turn into a game of whack-a-mole where most of the bugs it finds were written in previous LLM sessions. It's a huge struggle to get it to actually fix the root cause of the bug instead of patching the symptom.
- coliveira 2mo agoYes, it is relatively easy to find any solution to a bug. It is hard to find the root cause and change only what is necessary to fix the bug. AI makes it very efficient to "fix" bugs with code that needs to be fixed again later. The big problem in fixing bugs is understanding what is causing it, not to suggest a temporary "bug fix".
- manbash 2mo agoInteresting, but IMHO the title of this post misses the point of the linked page.
- eschneider 2mo agoCool. How many did they add?
- viktorcode 2mo agoThere's no way in hell those are all exploitable bugs. I'll bet my hat that those are "insecure practices".
- cameronh90 2mo agoHave LLMs shown any decent capabilities in resolving non-security bugs yet? Security bugs are important, but my day to day frustration with tech tends to be caused by UX/functionality bugs. And Google seem to struggle more than most other companies with those. Google Home became so unbelievably buggy, and I felt so betrayed by that, that I stopped using every single Google product - from Gmail to Android to Google Cloud Platform. The few physical Google devices that I'm yet to replace account for 10% of my tech, but still cause 90% of my problems.
- epinephrinios 2mo agoThe Google Home is hot garbage. I do like the hardware though.
- charlieyu1 2mo agoI wrote a pretty insignificant book. So many complaints over the years because there are a lot of mistakes. Not something to be proud of, but I’m the only person working on a pretty large project. Thanks to AI I fixed a large number of mistakes last month.
- jayd16 2mo agoThis is a bit off topic but is there something off about the bold "n" in the font used for bullet headers?
- effseven 2mo ago[dead]
- MollyRealized 2mo agoI reported and got my first bug [1] accepted and fixed - in all honesty, the bug has been one of the primary reasons I've stayed on Firefox. [1] - https://issues.chromium.org/issues/531793712 https://issues.chromium.org/issues/531793712
- nozzlegear 2mo agoTitle seems egregiously editorialized compared to the source title.
- userbinator 2mo agoWhat remains to be seen is whether Google also introduced more Chrome bugs in June than over the past two years, thanks to AI. The big problem is that AI output can be very convincing and look "right", even appear to work, until you examine it in detail and realise all the edge-cases it didn't handle.
- murukesh_s 2mo agoIf that is the case, isn't that the failure of your testing strategy/setup? If your style doesn't fit full vibe coding - you can do the path of AI generate code - human reviews and tests all edge cases flow - I think in any serious software - that should be the flow until next few iterations. But either way, don't see a going back to coding every line by hand now, those days are gone. Was fun while it lasted!
- calmingsolitude 2mo agoNo. I’ve heard this argument multiple times before - “let the human review and write tests” - and sure it might work, but this is almost never practiced. AI makes coding a lot faster, so no one is really ready to spend time on manual review and writing tests when LLMs can do that as well and take you most of the way.
- sothatsit 2mo agoThis is evidence of culture problems in whatever teams you are a part of, or extrapolating what you see on social media to all of software engineering. We still have a very strong review culture, and people work hard to review their own code before making PRs to avoid wasting other people's time.
- userbinator 2mo agoThe AI will not only write you subtly wrong code, it'll also write you as many tests as you want that are also subtly wrong and pass when run.
- cineticdaffodil 2mo ago
- stefanlindbohm 2mo agoThey show how many more bugs they fixed compared to previous periods, but not how the amount of resources allocated to fixing bugs changed over time. If they allocated proportionally more resources to bug fixing, it really doesn’t tell us anything about the tools being used. Might as well be that the learning here is that if you put resources on fixing bugs, you can in fact fix a lot of bugs.
- hegstal 2mo agoI've also seen similar recently at my place of work with LLM linting. Lots of hard to spot bugs in old very critical code have been caught. I remember seeing similar kinds of impacts back when linting or sanitizers were introduced. They're good for making it easy to encode more complex rules/areas-of-focus etc. that may have not been meaningfully possible with the more traditional tooling. Commercial LLMs in their current incarnations make for very expensive, but very effective intelligent linters. On the other hand, my experience so far has been that they are terrible code-reviewers, to the point almost everyone ignores any kind of generic LLM review entirely (correctly - most of it is noise). I feel the industry obsession with full-automation causes them to go down the wrong rabbit-holes on full-automation vs augmentation. In my view would be significantly more effective if the LLM review process in forge-tooling was designed with a human reviewer in the loop. The goal imo should be drastically improving the review efficacy & throughput of the human element, who needs to be acquiring confidence in & socializing the change in any sane org anyway. For example iteratively & interactively rubber-ducking out analysis + remediation, or interactively providing focusing guidance on the patch, instead of as agentic batch or harness processes as is cool nowadays. Sometimes a custom REPL is just better for an expert system. These things matter more as more and more code is LLM generated.
- rustfreeforme 2mo ago[flagged]
- fluffybucktsnek 2mo agoHow about you elaborate instead of vague posting?
- rustfreeforme 2mo ago[flagged]
- Aeolos 2mo agoAll the evidence is against you here, sorry. Just read the google or microsoft security blogs, linux kernel mailing list, or any of the published security research about bug density in Rust vs C++. Or you can bury your head in the sand while the rest of the industry is moving on, I guess.
- fluffybucktsnek 2mo agoThis post is basically a massive projection. When I guess, I make my guesses explicit. You, however, consistently make guesses and pretend they are true. You misinterpret facts so you can conveniently fit them into your narrative. I don't want "To Be Right", I just don't want dunces slopping nonsense into this comment chain. Not my fault you fall into that group. I don't need to assume Google's intentions and goals here. I want to see what they've done and the results they've reached with each action. I want to learn. You just want to reaffirm your delusions, that's why you have to pull the "you will learn" card rather than show objective metrics. If you really cared about quality, you would have them ready to present. Given that you haven't and even avoided doing so, it's safe to assume, as most of us have learned through experience, this to be a case of placebo at best. And placebo is not the sign of a good programmer. Next time, provide some actual data rather than vague self-gloating. Then, we can have a discussion.
- rustfreeforme 2mo ago[dead]
- fluffybucktsnek 2mo agoYou seem to be proficient at monologuing towards walls. Perhaps that's why you keep assuming people will just take your word.
- yourewrongsorry 2mo ago[flagged]
- dangisafag1 2mo ago[dead]
- lucianguyen 2mo ago[flagged]