7 ms·
> A challenge with this kind of study is that coding agents (Claude Code, OpenAI Codex) only started working really well in 2026? 5? 4? 3? Heard this one way
by ballsac 2mo ago
> A challenge with this kind of study is that coding agents (Claude Code, OpenAI Codex) only started working really well in
2026? 5? 4? 3?
Heard this one way too many times.
- simonw 2mo agoNobody was saying coding agents started working in 2023 or 2024, because the category was defined by Claude Code which was first released in February 2025.
- dgellow 2mo agoI would say that Aider is what defined coding agents. That was at least multiple months before Claude code. I remember seeing a coworker use aider for a hackathon project Adeline November-December 2024 , and it was already decent and pretty close to the DX we consider coding agents to have
- CuriouslyC 2mo agoDevin and Cline were the original true "agents" IIRC. Aider was the original terminal "agent" but it was't a true agent, it was sort of a hybrid agentic chat, since it had limited recursion ability (3 turns by default) and you had to configure it carefully to make it consistently take multiple turns in a row.
- simonw 2mo agoAider was missing a key feature: automatically executing code. That write-execute-fix loop was the big unlock. I believe Aider avoided adding that for safety concerns. Claude Code demonstrated that throwing safety to the wing somehow kind of worked out.
- bigstrat2003 2mo agoYep, the goalposts just keep shifting. In reality: they still don't work well, unless you're content with producing low quality work.
- throwaway7783 2mo ago"unless you're content with producing low quality work." - With the right guiding hand, it is a productivity multiplier without compromising quality. As a fully autonomous developer, it is a disaster.
- Forgeties79 2mo ago> With the right guiding hand, it is a productivity multiplier without compromising quality This just reads like another variation of “it’s the user not the tool,” which is just endless runway for always blaming people and never acknowledging the limitations of LLM’s. I’d be curious to hear how the recipients of your work enabled by the “productivity multiplier” feel about the quality.
- deleted 2mo ago[deleted]
- inglor_cz 2mo agoYou can't play an entire orchestra's sheet music on a single guitar either, but your playing ability still matters a lot. I would say that as of July 2026, with the right scaffolding, you can get reasonably good output out of a LLM, or better a combination of LLMs. For example, it pays off to prepare an implementation plan with one LLM and then let another LLM check it for flaws, then again. After several iterations like this, you will have a plan better than whatever you could come up with yourself. It often is the user and not the tool. LLMs are complicated, have nontrivial failure modes, and the user needs to steer them carefully. They might be the most complicated tools on the planet right now. Anecdotally, the recipients of my work have become visibly more happy in the last months. LLMs are great at diagnosing subtle problems which tend to appear at Friday night only, and this is the sort of problem that bugs actual people the most.
- Forgeties79 2mo ago> I would say that as of July 2026, with the right scaffolding, you can get reasonably good output out of a LLM, or better a combination of LLMs. Totally agree, I don’t think I said or implied otherwise. And yes can it can be the user and often even is, but when it comes to any LLM conversation I’ve been a part of it seems people think the only answer is “you’re using it wrong.” Evangelists swear it’s a 100x multiplier and anything counter to that means you’re either a Luddite who is blinded by politics or are too dumb to use the tool.
- cowanon77 2mo agoIt seems to be true this time though; I have observed it myself and heard it from several experienced developers I personally know and respect. It feels like some threshold was crossed with Opus 4.5 and Gpt 5.3, where the models are now able to reliably solve certain classes of problems that were previously unreliable. Time will tell of course, and it’s early, but inflection points do exist with progress.
- bluefirebrand 2mo agoI wonder how much of it is real and how much of it is people just being worn down by the hype to the point they can't fight it anymore Very smart people aren't immune to being worn down over time
- simonw 2mo agoI really don't think that's how it works. Smart, experienced developers who thought coding agents were junk for most of 2025 and think they're useful now in 2026 are not saying that because they got "worn down over time".
- SpaceNoodled 2mo agoAs a smart, experienced developer who's getting worn down over time, I disagree.
- trashface 2mo agoI was getting useful coding work done with GPT 3.5. I think devs saying "the models are finally good enough" this year are just trying to save face from their own previous irrational denials.
- ben_w 2mo agoUseful, yes, sometimes, but it wasn't fully automated "Here's our JIRA board URL, fix everything that's rated 1-3 story points and in the current sprint". Now it is.
- smrtinsert 2mo agoClaude 4.5 was it (nov 2025?), without a doubt. It went from frequent hallucinations to highly usable with much less garbage output. If you were making demos of AI tools around this time your demo/pitch/product was saved and you probably looked like a genius.
- ben_w 2mo agoI get the point having read much the same from Tesla (and fans) regarding self driving cars that still haven't done half the things that Musk said was just around the corner pending regulators a decade ago and repeatedly since then. And myself I keep making comparisons between AI and the progress in 90s video games where every minor improvement got called "photo realistic" and then forgotten with the next game engine: https://archive.org/details/nextgen-issue-26 https://archive.org/details/nextgen-issue-26 So I'm not gonna say "this is it" when the software quality really matters, and I absolutely won't speak to progress (or lack of it) outside of software. But I will say "you can look around and easily see small businesses using AI to generate posters, quite a lot of small business software and websites are in the same category: the mistakes are real but increasingly don't matter".
- qsera 2mo ago>the mistakes are real but increasingly don't matter... I think it would start to matter once again. People will get fed up of AI posters and art. I think they already are...and once some threshold is crossed, the business won't dare to use AI generated assets/designs. Turns out humans are much better at recognizing patterns in stuff that is generated ONLY using patterns from human generated content.
- ben_w 2mo ago> People will get fed up of AI posters and art. I think they already are Agreed, but will this look like a meme/fashion cycle? If so, re-prompt each year with a different look. Yes, there are still issues here, a friend found an image he was amazed was AI generated, but to me it was obviously so, so I showed him a screenshot of ChatGPT making something just it and included my prompt: create image: hand drawing of cute springer spaniel puppy looking sideways, various geometric shapes drawn in layer behind and in front of the puppy, all done in style of 7 year old using crayons with mediocre colouring-in skills As I said to them: yeah, the line thickness feels AI, to me, the bad colouring-in scribbles feel like just the art style it was propmpted with it's like: it gets the big picture of the composition, and it knows how to colour in badly, but it doesn't know how to draw a dog as badly as the colouring in > Turns out humans are much better at recognizing patterns in stuff that is generated ONLY using patterns from human generated content. We're better at recognising patterns full stop. All biological brains are, and needed to be better than the current state of the art in machine learning because if a living organism was as poor at learning patterns as the SotA in machine learning, the organism would starve to death before being able to pick up anything and eat it. AI also has a second disadvantage, because there are so few models: the laziest of ChatGPT "thinkpiece" blog posts being everywhere is hard to miss, and 5000 fake bloggers all prompting the same model with "find biggest news story of today and write a blog post about it in a way that maximises my ad revenue" will get 5000 almost identical posts. This will remain true while each instance of the most commonly used AI fail to talk to each other in a way that at least mimics them collectively getting bored with writing the same thing 5000 times, it does not depend on e.g. quality.
- GolfPopper 2mo agoPerhaps the LLM companies need to start hiring true Scotsmen?
- ChrisMarshallNY 2mo agoUniversity of Edinburgh is a good school.
- cmenge 2mo agoI guess it's important who one hears this from. I just spoke to a fried who is a headhunter and who's been trying to automate his processes for a while (he likes to fiddle and certainly has skills, but he's not an engineer). He kept trying, but it just wasn't good enough. Now he said with GPT Work and Sol, it worked, but the key point is: all of it suddenly worked. The problem was one of reliability, of handling edge cases. All previous attempts / model-harness-combinations were too brittle and needed too much observation and fiddling - cheaper to do it yourself. Now he says "I don't know why I would ever hire a recruiter [the folks doing the cold outreach] again. I can focus on the candidate screening and acquiring projects, everything else is fully automated". This doesn't come from an engineer or an AI lab, but a technically inclined power user, and I think this is where things get interesting.
- mstaoru 2mo agoThen it just becomes a new baseline (everyone have access to the same LLMs), and recruiting moves up the philosophical ladder where human can add more value. What will it be? I don't know, I'm not a recruiter.
- noosphr 2mo agoAgain I've heard this since 2022 when gpt3.5 came out. This is like microprocessors in the 80s. Sure they double in capability every 18 months but the start is so pathetic it will be 30 years before they are good enough for everyday tasks.
- StilesCrisis 2mo agoCPUs in the 90s were amazing! They were over-specced for "everyday tasks." Our problem is that we overbuilt CPUs too much, so software is now written with ten unnecessary layers of abstraction because there's no real reason to simplify.
- noosphr 2mo ago3gb of ram aren't enough everyday tasks at any sane resolution. That we still don't have 1200 ppi desktop monitors is as stupid as using black and white screens in 2000.
- IshKebab 2mo agoI haven't. Around the start of 2026 is pretty widely mentioned as when they went from "this is broken slop" to "huh this is actually 90% what I would have written", which matches my experience.
- yuye 2mo ago>"huh this is actually 90% what I would have written" The last 10% is always the hardest part, though
- the_gastropod 2mo agoI’ve been feeling gaslit about this too. Getting major “we’re still early!” crypto bro vibes from this constant goalpost moving.
- dgellow 2mo agoThe „you will be left behind if you don’t fully embrace the whole thing right now“ is a 1:1 match with cryptocurrency hype
- CuriouslyC 2mo agoThe capex is still "early" for sure (i.e. data centers are still being planned and built out, we're hardware/energy constrained). If model scaling holds out, we're "early-ish" in terms of the reliability and performance of these systems, just based on utilization of the compute from the planned capex. If we hit hard diminishing returns and we don't find architectural/data workarounds, that would put a wrinkle in things, but I suspect that the AI we have now is capable of helping us find those workarounds and keep things moving.
- 3ff3 2mo ago[dead]
- CuriouslyC 2mo agoWhere did you hear me mention AGI? I'll settle for fully automated software engineering and math, which are both most certainly coming.
- siva7 2mo agoThe timescale is well established: Late '25 was the start of agentic ai when capabilities of model + api + scaffold reached autonomous state. Any study comapring events before that timeframe is comparing apples with oranges.
- erispoe 2mo agoNovember 2025: https://simonwillison.net/tags/november-2025-inflection/ https://simonwillison.net/tags/november-2025-inflection/
- jcranmer 2mo agoSince so many people are doubting you here, here is a post from ~a year ago that's pulling the same "LLMs 6 months ago were crap, now they're awesome" shtick: https://fly.io/blog/youre-all-nuts/ https://fly.io/blog/youre-all-nuts/. There's more posts along this vein being put out from 2024 on or so.
- simonw 2mo agoI find this attitude baffling. Things are allowed to get better more than once! The idea that "yeah, you said technology had improved in the past, and now you're saying it has improved again" is a gotcha just seems incoherent to me.
- hfthkgghnd 2mo ago[dead]
- codinhood 2mo agoI keep running into this as well. Someone shared a link to a study from last year telling me that AI doesn't make programmers as productive as they think. The study was obviously done months prior to it's release. So evaluating the state AI even further back. But this didn't seem to concern them. The study said X, therefore it applies to today. I'm just not sure it's worth it to argue with others about it at this point. Not that I'm 100% all behind AI coding, but I'm just shocked people are still this resistant.
- dolebirchwood 2mo ago> I'm just not sure it's worth it to argue with others about it at this point. This is the correct response. Let closed-minded people do their thing. Makes them less competitive against you. There's nothing for you to gain by trying to help them understand what they are missing.
- ballsac 2mo agowait, who's being closed-minded?
- satvikpendem 2mo agoStep changes in functions exist.
- deleted 2mo ago[deleted]