5 ms·
It seems to be true this time though; I have observed it myself and heard it from several experienced developers I personally know and respect. It feels like s
by cowanon77 2mo ago
It seems to be true this time though; I have observed it myself and heard it from several experienced developers I personally know and respect. It feels like some threshold was crossed with Opus 4.5 and Gpt 5.3, where the models are now able to reliably solve certain classes of problems that were previously unreliable.
Time will tell of course, and it’s early, but inflection points do exist with progress.
- bluefirebrand 2mo agoI wonder how much of it is real and how much of it is people just being worn down by the hype to the point they can't fight it anymore Very smart people aren't immune to being worn down over time
- simonw 2mo agoI really don't think that's how it works. Smart, experienced developers who thought coding agents were junk for most of 2025 and think they're useful now in 2026 are not saying that because they got "worn down over time".
- SpaceNoodled 2mo agoAs a smart, experienced developer who's getting worn down over time, I disagree.
- trashface 2mo agoI was getting useful coding work done with GPT 3.5. I think devs saying "the models are finally good enough" this year are just trying to save face from their own previous irrational denials.
- ben_w 2mo agoUseful, yes, sometimes, but it wasn't fully automated "Here's our JIRA board URL, fix everything that's rated 1-3 story points and in the current sprint". Now it is.
- gymbeaux 2mo agoI didn’t start using Claude Code until late 2025. Prior to that I would use ChatGPT to give me snippets of code but I was still doing most of the actual code writing. Coworkers told me in late 2025 about how they hadn’t written a line of code in “months” and just use Claude Code/agentic “whatever” so I tried out Claude Code and was pleasantly surprised. It is passable to have entire apps written by LLMs (I’ve made several that I otherwise never would have had the time to create by hand), but I wouldn’t say maintainable or easily extendable. It’s hard to be specific, but there’s something about LLM code that doesn’t look “natural”, and I’m not talking about the excessive use of comments in code. The code itself is unnatural. Functional, but unnatural. I wouldn’t want to suddenly lose LLMs and have to read through and understand and continue enhancing a codebase created by an LLM.
- fcatalan 2mo agoFor me it feels a lot like generated images or video. I've made lots of things now, but those that are 100% LLM written "work" but are uncanny, weird and the details are wrong everywhere you care to look in detail.
- r_lee 2mo agoI think how I'd put it is like, it's a deterministic common denominator thing iterated on, but it's not really intuitive, it's not really "what makes sense" but rather "what would x look like", similar to how you point out it is when it comes to image generation it's useful for scaffolding but after that I'm not sure how you could rely on it without being in the loop and directing how the code should be like
- chrsw 2mo agoThis has been my experience has well. For code generation, you have to really constrain these LLMs on your coding style, design goals, test cases, and overall expectations. I mainly use these tools to help my understanding of the code and to generate code for very specific problems. Even after all that setup and careful review I'd say it's still a net big speed up for certain software engineering tasks. Jason Turner gave an excellent talk at last year's CppCon explain how he thinks tools can be used to make generative AI coding assistance safer and more productive. https://www.youtube.com/watch?v=xCuRUjxT5L8 https://www.youtube.com/watch?v=xCuRUjxT5L8
- TeriyakiBomb 2mo agoIt tends to be when the training data wanders into their area of expertise temporarily and they go “OMG, they hype is real. I was so wrong” and then a few releases later they’re on the train and furious that the skills in their domain space have not just stopped improving, but regressed. Cue someone else in a different part of the world starting the same cycle. Meanwhile the guy who leaned in a year ago and gave up reading the output is beginning to see work grind to a halt and throwing more agents at it is increasingly not working. You can see these tropes all over social media near constantly.
- Kiro 2mo ago> You can see these tropes all over social media near constantly. You should stop using social media as your yardstick.
- wolvesechoes 2mo ago> You should stop using social media as your yardstick. Yes, don't believe people posting on HN.
- simianwords 2mo agoWhat about terry tao?
- wolvesechoes 2mo agoWho? But in all seriousness - you can even bring up Pope himself. Don't care. Show me data, show me the leaps our software made with all this 100x productivity boost. Show me a myriad of better LLVM projects, new usable kernels, and so on. Show me sharp decline in bugs and defects in existing projects. Keep blog posts and HN comments.
- simianwords 2mo agoDo you want a graph of exploits found vs time? Erdos problems solved vs time?
- iLoveOncall 2mo ago[flagged]
- simonw 2mo agoPaul Ford wrote a great piece about this in the NY Times: https://www.nytimes.com/2026/02/18/opinion/ai-software.html?unlocked_article_code=1.NFA.UkLv.r-XczfzYRdXJ&smid=url-share https://www.nytimes.com/2026/02/18/opinion/ai-software.html?... > November was, for me and many others in tech, a great surprise. Before, A.I. coding tools were often useful, but halting and clumsy. Now, the bot can run for a full hour and make whole, designed websites and apps that may be flawed, but credible. I spent an entire session of therapy talking about it. Max Woolf is a good one: https://minimaxir.com/2026/02/ai-agent-coding/ https://minimaxir.com/2026/02/ai-agent-coding/ > The real annoying thing about Opus 4.6/Codex 5.3 is that it’s impossible to publicly say “Opus 4.5 (and the models that came after it) are an order of magnitude better than coding LLMs released just months before it” without sounding like an AI hype booster clickbaiting, but it’s the counterintuitive truth to my personal frustration. [...] A year ago, I was one of those skeptics who was very suspicious of the agentic hype. DHH - https://newsletter.pragmaticengineer.com/p/dhhs-new-way-of-writing-code https://newsletter.pragmaticengineer.com/p/dhhs-new-way-of-w... > Six months ago, in an episode of the Lex Fridman podcast, David shared how he doesn’t use AI tools to write code: he types out all his code. But things have changed a lot since then. > In this episode, we discuss his approach to building software, how it’s changed in the last six months, and why he now takes an agent-first approach, and how he barely writes any code by hand. Linus Torvalds, two weeks ago (though I don't have evidence that he was a skeptic before) https://lore.kernel.org/linux-media/CAHk-=wi4zC+Ze8e+p3tMv8TtG_80KzsZ1syL9anBtmEh5Z40vg@mail.gmail.com/ https://lore.kernel.org/linux-media/CAHk-=wi4zC+Ze8e+p3tMv8T... > There are other questions around AI (like what the economy of it will actually look like in the end), but "is it useful" is no longer one of those questions. Anybody who doubts that clearly hasn't actually used it. Donald Knuth! https://www-cs-faculty.stanford.edu/~knuth/papers/claude-cycles.pdf https://www-cs-faculty.stanford.edu/~knuth/papers/claude-cyc... > Shock! Shock! I learned yesterday that an open problem I'd been working on for several weeks had just been solved by Claude Opus 4.6 - Anthropic's hybrid reasoning model that had been released three weeks earlier! It seems that I'll have to revise my opinions about "generative AI" one of these days. What a joy it is to learn not only that my conjecture has a nice solution but also to celebrate this dramatic advance in automatic deduction and creative problem solving.
- bluefirebrand 2mo agoYou're being one of the voices trying to do the wearing down right now in this very thread imo Of course it doesn't seem that way to you. Preachers view themselves as spreading the good word, they don't see how annoying it is being preached at
- TeriyakiBomb 2mo agoThing is. You can find an extremely similar paragraph written about Claude 4.x or some equivalent gpt. And simultaneously, many people expressing their frustration and the shortcomings of <insert any model> “But it’s different this time” - several people, several times over the last couple of years. This is not at all a dig at you, I’m very sorry if it reads that way. My point is these things only get truly better in anecdotes. The ways in which they fail is yet to change. Just yesterday I had gpt 5.3 generate completely awful code for the Cinema 4D Python API. Also an anecdote. But for all of the people saying they are truly intelligent and truly reason, they still make obvious mistakes, write around problems, fail entirely at architectural decisions, fail at random, generate FAR too much code. And no amount of harnesses, methodologies, loops make much of a difference. If you listen to people on the internet they say it’s all working. You listen to people on the job and they mostly say it’s creating tech debt and a review bottleneck. Also burnout, so much burnout. I think LLMs are mediocre. I think it’s fine they’re mediocre. You can work with low expectations. But the hype cycles are so tiresome.
- user43928 2mo agoI believe it is a widely accepted opinion that agentic coding took off with Opus 4.5 late in 2025. Why would you attempt to use GPT 5.3 to generate code today and form an opinion on that basis? I do not think it is even still available in Codex, I believe it only has the smaller, distilled GPT 5.3 Codex Spark.
- CuriouslyC 2mo agoLLM capabilities are spiky. They're amazing at some things, and poor at others. Over time, the set of things they're amazing at has grown, while the set of things they're poor at has shrunk. If you think LLMs are just "mediocre" without any nuance, that's a sign you haven't spent the time to evaluate them in order to make an informed opinion.
- ACS_Solver 2mo agoI don't buy into the huge LLM hype but I certainly think late 2025 was the inflection point. Up until then, I thought LLMs universally sucked at code. GPT-3.5, GPT-4, o1, o3, Claude 3.5, Sonnet 4, the whole bunch. Each iteration got marginally better, hallucinated APIs less and so on, but my overall evaluation of them all was that they wrote crap code and were unsuitable for anything other than one-off scripts. Then the incremental improvements did, in my experience, cross some kind of threshold in late 2025 where the things became useful. It is of course anecdotal and personal judgment. But I asked LLMs to implement a small feature in my codebase (my usual test) and finally it produced code I was happy with. They've also been able to locate and diagnose a problem based on logs. In my view it's now a markedly different level of capability than we had a year ago, though I would call the previous two years equally useless.
- ModernMech 2mo agoIf it's so evident, why can't someone prove it with something more than "it seems better and everyone agrees"?
- StilesCrisis 2mo agoYou think the world is lacking in LLM benchmarks?
- ModernMech 2mo agoA benchmark doesn't prove "although we said it before, this time it's true". People pointed to the benchmarks then as well.
- StilesCrisis 2mo agoAnecdotes aren't real. Don't believe any of them if you don't want to.
- ModernMech 2mo agoOkay thanks. But what I'm saying is all of these anecdotes should add up to something measurable if there's something to the idea that an inflection point was reached in November 2025, right? There's a lot of marketing and hype and motivated reasoning going on, so that's why I don't trust broadly-reported notions.
- Kiro 2mo agoWhat measurable thing would convince you?
- ModernMech 2mo agoExactly, "productivity" and similarly "efficiency" are meaningless weasel word unless they're qualified and measured against something. So which measurable things? Let's start with just anything really because then we could ground the conversation in something real rather than just feelings of being more productive, which can be deceptive. AnimalMuppet in their reply to my OP comment makes the point that these things haven't been around long enough to measure end-to-end productivity gains and come to a conclusion either way, and I agree with that. But we can still at least be measuring something. For example, in my case AI has allowed me to write 10x more LOC than I usually would in a similar amount of time. But having to review it all, I've also deployed 1/4 the number of releases I normally would in the same period. By one measure I'm more productive, by another I'm less productive. People could claim to be more productive by skipping the review. But in that case did AI make you more productive or did you lower standards? People could say they're using AI to do the review but is the impact of that being measured and has that caused more or fewer bugs? If more bugs, has the time to fix those been factored into overall productivity? In my experience it's common for people to eagerly count immediate productivity gains and discount long-term productivity sinks. For this reason I think case studies are the best convincing thing, because they properly contextualize the usage and consider a longer-term window. They're also backwards looking instead of in-the-moment, so have the benefit of hindsight. But they're harder to come by and we probably won't see any meaningful case studies for a thing that people say happened in November. But the very least people can be doing is just defining what they mean when they say "productivity" because otherwise everyone is talking past one another.