5 ms·
2025 State of AI Code Quality
- esafak 1y agoLots of numbers. I'm interested in seeing the trends over time. I bet with their products they could track this daily.
- diggan 1y ago> are 2.5x more likely to merge code without reviewing it What the fuck? Are people taking "vibe coding" as a serious workflow? No wonder people's side projects feel more broken and buggy than before. Don't get me wrong, I "work with" LLMs, but I'd never merge/use any code that I didn't review, none of the models or tooling is mature enough for that. Really strange how some people took a term that supposed to be a "lol watch this" and started using it for work...
- dartos 1y ago> Really strange how some people took a term that supposed to be a "lol watch this" and started using it for work... don't forget about the insane amount of marketing around AI code companies and how they put "vibe coding" in front of everyone's face all the time. You tell someone something enough times and they'll belive it
- uludag 1y agoPlus, vibe coding is the surest way for these companies to insert themselves as the ultimate middleman in the entire software development process. As an aside, I really hate how cynical I feel I've been compelled to become at the arrival of such a genuinely innovative technology. Like, with this very article I can't help but to think there are ulterior motives behind it's production.
- InitialLastName 1y agoThis is key. They've found a way to sell extremely-high-brand-stickiness shovels in a gold rush.
- bluefirebrand 1y ago> As an aside, I really hate how cynical I feel I've been compelled to become at the arrival of such a genuinely innovative technology Yeah I'm feeling like this too. This should be so exciting! We're getting close to the Star Trek dream of just telling the computer to do work and it works! I've been trying to examine why it's not exciting for me, and I'm actually pretty repulsed by it. I think it's a combination of things To start, I'm pretty disgusted by the blatant and unapologetic scraping of every single scrap of public data, regardless of license or copyrights. I'm also really discouraged by how this is turning out to be another tool in the capitalist toolbox to justify layoffs, increase downward pressure on salaries, and once again extract more value per hour worked from employees I also don't feel like the technology is actually that good or reliable yet. It has transformed my workflow but for the worse. Because my company is very bullish on AI it has resulted in me losing what little control I had to choose the tools that I feel are best for my job, in favor of what they want me to use because of the hype Ultimately I'm cynical because I don't feel like this is making my life better. It feels like it is enriching other people at my expense and I am very bitter about it
- jasonthorsness 1y agoThe absolute number who merged without reviewing was only 24% so maybe there is still hope!
- simonw 1y ago1/4 programmers in a survey merging code without reviewing it is terrifying to me. That number should be as close to 0% as possible.
- msgodel 1y agoI've done it for "nice to have" features in their own modules that I don't really care about and aren't consumed by anything else (recently an SVG plot generator for a program I wrote.) The LLM one-shotted it and I left it alone for a long time. Stuff like that is great application for literal vibe coding. I can't imagine doing it for anything serious though.
- diggan 1y agoYeah, for one-off, never-to-be-touched again I guess that kind of makes sense. But this survey seems to span much more than just one-off tiny things, and gives the impression people working as professionals in companies are actually doing "vibe-coding" not as a joke, but as a workflow, for putting software into production.
- mattgreenrocks 1y agoAs awful as it is, it is entirely understandable: it follows naturally from the claims that LLMs can replace programmers entirely. As capable as the models are, what matters more is how competent they are perceived to be, and how that is socialized. The hype machine is at deafening levels currently.
- orangebread 1y agoNot for nothing, but I did create an entire game in browser using phaser as the engine. But I'm also an experienced developer and at this point, an experienced "vibe coder". I use that last term loosely because I have a structured set of rules I have AI follow. To really understand AI's capability you have to have experienced it in a meaningful way with managed expectations. It's not going to nail what you want right away. This is also why I spend a lot of time up front to design my features before implementing.
- diggan 1y ago> I use that last term loosely because I have a structured set of rules I have AI follow Right, but what defines if what you're doing is "vibe-coding" or not is if you actually view the code it produces, at any point of the workflow. You're "vibe-coding" if you're merging/pushing without reviewing the code. I'm also an experienced developer, and used LLMs a lot, but never pushed/merged anything into production that I haven't read and understood myself.
- namanyayg 1y ago> "65% of developers using AI for refactoring and ~60% for testing, writing, or reviewing say the assistant “misses relevant context." > "Among those who feel AI degrades quality, 44% blame missing context; yet even among quality champions, 53% still want context improvements." Is this even true anymore? Doesn't happen to me with claude 4 + claude code.
- jmsdnns 1y ago> 25% of developers 1 in 5 AI-generated suggestions estimate that contain factual errors or misleading code. I cannot believe what's said in the report because it doesnt even reflect what my pro-AI coding friends say is true. Every dev I know says AI generated suggestions are often full of noise, even the pro-AI folks.
- bluefirebrand 1y agoI think this really highlights the difference between "pro ai" and "anti ai" people "It's full of noise but I'm confident I can cut through it to get to the good stuff" - Pro AI "It's full of noise and it takes more effort to cut through than it would take to just build it myself" - Anti AI I'm pretty Anti myself. I think "I can cut through the noise" is pretty misplaced overconfidence for a lot of devs
- diggan 1y agoI don't think I would place myself on either sides, I guess I'm in the "AI is OK at some stuff" camp. But if you're getting a lot of noise, I'd immediately try to adjust my system/user prompt to never get that noise in the first place. I'm currently using a variation of https://gist.github.com/victorb/1fe62fe7b80a64fc5b446f82d3137398 https://gist.github.com/victorb/1fe62fe7b80a64fc5b446f82d313... which is basically my personal coding guidelines but "codified" as simple rules for LLMs to understand. For anything besides the dumb models, I get code that more or less looks exactly like how I would have written it myself. When I find I get code back that I'm not happy with, I adjust the system/user prompt further so this time and the next it returns code like how I would have done it.
- bluefirebrand 1y agoI feel I should clarify When it comes to judging the quality of AI output, I do agree with "AI is ok at some stuff" When I say I tend to fall on the Anti AI side, I am saying "But I still don't think it's worth using much" I don't really want to lean on tools that are just ok at some stuff.
- 1y ago
- ilitirit 1y agoI currently have a big problem with AI-generated code and some of the junior devs on my team. Our execs keep pushing "vibe-coding" and agentic coding, but IMO these are just tools. And if you don't know how to use the tools effectively, you're still gonna generate bad code. One of the problems is that the devs don't realise why it's bad code. As an example, I asked one of my devs to implement a batching process to reduce the number of database operations. He presented extremely robust, high-quality code and unit tests. The problem was that it was MASSIVE overkill. AI generated a new service class, a background worker, several hundred lines of code in the main file. And entire unit test suites. I rejected the PR and implemented the same functionality by adding two new methods and one extra field. Now I often hear comments about AI can generate exactly what I want if I just use the correct prompts. OK, how do I explain that to a junior dev? How do they distinguish between "good" simple, and "bad" simple (or complex)? Furthermore, in my own experience, LLMs tend to pick up to pick up on key phrases or technologies, then builds it's own context about what it thinks you need (e.g. "Batching", "Kafka", "event-driven" etc). By the time you've refined your questions to the point where the LLM generate something that resembles what you've want, you realise that you've basically pseudo-coded the solution in your prompt - if you're lucky. More often than not the LLM responses just start degrading massively to the point where they become useless and you need to start over. This is also something that junior devs don't seem to understand. I'm still bullish on AI-assisted coding (and AI in general), but I'm not a fan at all of the vibe/agentic coding push by IT execs.
- hiq 1y ago> OK, how do I explain that to a junior dev? They could iterate with their LLM and ask it to be more concise, to give alternative solutions, and use their judgement to choose the one they end up sending to you for review. Assuming of course that the LLM can come up with a solution similar to yours. Still, in this case, it sounds like you were able to tell within 20s that their solution was too verbose. Declining the PR and mentioning this extra field, and leaving it up to them to implement the two functions (or equivalent) that you implemented yourself would have been fine maybe? Meaning that it was not really such a big waste of time? And in the process, your dev might have learned to use this tool better. These tools are still new and keep evolving such that we don't have best practices yet in how to use them, but I'm sure we'll get there.
- sathomasga 1y agoSurvey from a company that's in the business of AI coding and thus has a monetary interest in promoting the technology. No details on who conducted the survey (the company itself?) or how the 609 respondents were selected. If limited to the company's own customers, massive selection bias. The results may or may not reflect reality, but this "report" is just marketing bullshit.
- hiq 1y agoI'd be interested in seeing comparisons between languages. I expect that a terse language with an expressive type system (is that Haskell maybe?) can lead to way better results in terms of usefulness than, say, bash, because I can rely on the type system and the compiler to have gotten rid of some basic mistakes, and I can read the code faster (since it's more concise). I've mostly used LLMs with python so far and I'm looking forward to using them more with compiled languages where at least I won't have mismatching types a compiler would have detected without my help.
- hippari2 1y agoI think what really matters is how much code of that language is on StackOverflow :)
- wbharding 1y agoIt's hard to reconcile how 59% of devs in their survey are "confident" AI is improving their code quality, with prior empirical research that shows a surge in added & copy/pasted lines w/ a corresponding drop in moved (refactored) lines https://www.gitclear.com/ai_assistant_code_quality_2025_research https://www.gitclear.com/ai_assistant_code_quality_2025_rese... My experience (using a mix of Copilot & Cursor through every day) is that AI has become very capable of solving problems of low-to-intermediate complexity. But it requires extreme discipline to vet the code afterward for the FUD and unnecessary artifacts that sneak in alongside the "essential" code. These extra artifacts/FUD are to my mind the core of what will make AI-generated code more difficult to maintain than human-authored code in the long-term.
- elpocko 1y agoI wish LLMs were generally viewed as Eliza on steroids, a thing to generate plausible sounding text with, in places where we used primitive generators based on Markov models before. To implement smarter NPCs in games, and virtual chat partners to talk to, just for fun. They are, after all, really fun to play with. They should be used as smart autocomplete in your IDE, not to generate whole projects from scratch. As an idea generator when you're stuck. This requirement to be commercially useful and valuable, and to aid all kinds of businesses everywhere, gave a bad reputation to what is otherwise an amazing technological achievement. I am an outspoken AI enthusiast, because it is fun and interesting, but I hate how it is only seen as useful when it can do actual work like a human.