43 ms·
Opus 4.5 is not the normal AI agent experience that I have had thus far
- kelseyfrog 9mo agoCan it pre-emptively write the HN comment where someone says it utterly fails for them but no one else is able to reproduce?
- halfmatthalfcat 9mo ago[flagged]
- LatencyKills 9mo agoI really wonder what means for software moving forward. In the last few months I've used Claude Code to build personalized versions of Superwhisper (voice-to-text), CleanShot X (screenshot and image markup), and TextSniper (image to text). The only cost was some time and my $20/month subscription.
- adriand 9mo ago> I really wonder what means for software moving forward. It means that it is going to be as easy to create software as it is to create a post on TikTok, and making your software commercially successful will be basically the same task (with the same uncontrollable dynamics) as whether or not your TikTok post goes viral.
- tannedNerd 9mo agoThe problem with this is none of this is production quality. You haven’t done edge case testing for user mistakes, a security audit, or even just maintainability. Yes opus 4.5 seems great but most of the time it tries to vastly over complicate a solution. Its answer will be 10x harder to maintain and debug than the simpler solution a human would have created by thinking about the constraints of keeping code working.
- LatencyKills 9mo agoAgree... but that is exactly what MVPs are. Humans have been shipping MVPs while calling them production-ready for decades.
- maherbeg 9mo agoThat may be true now, but think about how far we've come in a year alone! This is really impressive, and even if the models don't improve, someone will build skills to attack these specific scenarios. Over time, I imagine even cloud providers, app stores etc can start doing automated security scanning for these types of failure modes, or give a more restricted version of the experience to ensure safety too.
- afavour 9mo agoThere's a fallacy in here that is often repeated. We've made it from 0 to 5, so we'll be at 10 any day now! But in reality there are any number of roadblocks that might mean progress halts at 7 for years, if not forever.
- christophilus 9mo agoEven if progress halts here at 5, I think the programming profession is forever changed. That’s not hyperbole. Claude Code— if it doesn’t improve at all— has changed how I approach my job. I don’t know that I like this new world, but I don’t think there’s any going back.
- usefulposter 9mo ago
- ChrisbyMe 9mo agoMm this is my experience as well, but I'm not particularly worried about software engineering a whole. If anything this example shows that these cli tools give regular devs much higher leverage. There's a lot of software labor that is like, go to the lowest cost country, hire some mediocre people there and then hire some US guy to manage them. That's the biggest target of this stuff, because now that US guy can just get equal or hight code in both quality and output without the coordination cost. But unless we get to the point where you can do what I call "hypercode" I don't think we'll see SWEs as a whole category die. Just like we don't understand assembly but still need technical skills when things go wrong, there's always value in low level technical skills.
- adriand 9mo ago> If anything this example shows that these cli tools give regular devs much higher leverage. This is also my take. When the printing press came out, I bet there were scribes who thought, "holy shit, there goes my job!" But I bet there were other scribes who thought, "holy shit, I don't have to do this by hand any more?!" It's one thing when something like weaving or farming gets automated. We have a finite need for clothes and food. Our desire for software is essentially infinite, or at least, it's not clear we have anywhere close to enough of it. The constraint has always been time and budget. Those constraints are loosening now. And you can't tell me that when I am able to wield a tool that makes me 10X more productive that that somehow diminishes my value.
- fragmede 9mo agoThere was a previous edit that made reference to the water usage of AI datacenter that I'm responding to. If AI datacenters' hungry need for energy gets us to nuclear power, which gets us the energy to run desalination plants as the lakes dry up because the Earth is warming, hopefully we won't die of thirst.
- falkensmaize 9mo agoWhat diminishes your value is that suddenly everybody can (in theory anyway) do this work. There’s a push at my company to start letting designers do their own llm-assisted merge requests to front end projects. So now CEOs are greedily rubbing their hands together thinking maybe everybody but the plumber can be a “developer” now. I think it remains to be seen whether that’s true, but in the meantime it’s going to make getting and keeping a well-paying developer gig difficult.
- s-macke 9mo agoOpus 4.5 has become really capable. Not in terms of knowledge. That was already phenomenal. But in its ability to act independently: to make decisions, collaborate with me to solve problems, ask follow-up questions, write plans and actually execute them. You have to experience it yourself on your own real problems and over the course of days or weeks. Every coding problem I was able to define clearly enough within the limits of the context window, the chatbot could solve and these weren’t easy. It wasn’t just about writing and testing code. It also involved reverse engineering and cracking encoding-related problems. The most impressive part was how actively it worked on problems in a tight feedback loop. In the traditional sense, I haven’t really coded privately at all in recent weeks. Instead, I’ve been guiding and directing, having it write specifications, and then refining and improving them. Curious how this will perform in complex, large production environments.
- giancarlostoro 9mo ago> In the traditional sense, I haven’t really coded privately at all in recent weeks. Instead, I’ve been guiding and directing, having it write specifications, and then refining and improving them. This is basically all my side projects.
- lelanthran 9mo ago> You have to experience it yourself on your own real problems and over the course of days or weeks. How do you stop it from over-engineering everything?
- petcat 9mo agoThis has always been my problem whether it's Gemini, openai or Claude. Unless you hand-hold it to an extreme degree, it is going to build a mountain next to a molehill. It may end up working, but the thing is going to convolute apis and abstractions and mix patterns basically everywhere
- jama211 9mo agoNot in my experience - you need to build the fact that you don’t want it to do that into your design and specification.
- waynenilsen 9mo agoOnce you get your setup bulletproof such that you can have multiple agents running at the same time that can run unit tests and close their own loops things get even faster. However you accomplish that. Not as easy as it sounds mostly (and absurdly) due to port collision. E2E testing with playwright is another leap.
- nikisil80 9mo ago[flagged]
- jgbuddy 9mo ago[flagged]
- deleted 9mo ago[deleted]
- tomhow 9mo agoWhen disagreeing, please reply to the argument instead of calling names. https://news.ycombinator.com/newsguidelines.html https://news.ycombinator.com/newsguidelines.html
- matusp 9mo agoThis is a natural response to software enshittification. You can hardly find an iOS app that is not plagued by ads, subscriptions, or hostile data collection. Now you can have your own small utilities that can work for you. This sort of personal software might be very valuable in the world where you are expected to pay 5$ to click any button.
- nikisil80 9mo agoYeah sure but have you considered that the actual cost of running these models is actually much greater than whatever cost you might be shelling out for the ad-free apps? You're talking to someone who hates the slopification and enshittification of everything, so you don't need to convince me about that. However, everything I've seen described in the replies to my initial comment - while cute, and potentially helpful on a case-by-case basis, does NOT warrant the amount of resources we are pouring into AI right now. Not even fucking close. It'll all come crashing down, taxpayers the world over will be caught with the bag in their hands, and for what? So that we can all have a less robust version of an app that already exists but that has the colours we want and the button where we want it? If AI cost nothing and wasn't absolutely decimating our economy, I'd find what you've shared cute. However, we are putting literally all of our eggs, and the next generation's eggs, and the one after that, AND the one after that, into this one thing, which, I'm sorry, is so far away from everything that keeps on being promised to us that I can't help but feel extremely depressed.
- Papazsazsa 9mo ago"Opus 4.5 feels to me like" The article is fine opinion but at what point are we going to either: a) establish benchmarks that make sense and are reliable, or b) stop with the hypecycle stuff?
- NewsaHackO 9mo ago>establish benchmarks that make sense and are reliable How aren't current LLM coding benchmarks reliable?
- Papazsazsa 9mo agoThey're manipulated.
- NewsaHackO 9mo agoUnless you are going to be more specific, that criticism applies to all benchmarks that are connected to a positive gain, not just AI coding benchmarks.
- cardine 9mo ago> make sense and are reliable If you can figure out how to create benchmarks that make sense, are reliable, correlate strongly to business goals, and don't get immediately saturated or contorted once known, you are well on your way to becoming a billionaire.
- mococa 9mo agoI agree, it wrote an entire NES emulator for me. https://news.ycombinator.com/item?id=46443767 https://news.ycombinator.com/item?id=46443767
- lawlessone 9mo agoIt cloned one of the many open source ones available is what you mean.
- jama211 9mo agoTo be fair that’s what I’d have done had I had to build it. Use a lot of examples etc and build on what other people have done
- koiueo 9mo agoI assume, the purpose would be to learn how it's done. There's no place for this when you vibecode. And if not learning, what's the point of implementing something that already exists? When I'm dying of dehydration because humanity has depleted all fresh water deposits, I'll think of you and your stupid NES emulator which is just an LLM-produced copy of many ones that had already existed.
- minimaxir 9mo agoThe majority of open source software development is "implementing something that already exists", but with improvements, such as for specific use cases and constraints (like the original NES emulator) or by making it more performant. That's how the ecosystem mutates and grows, and it's worked well for decades.
- lawlessone 9mo ago>The majority of open source software development is "implementing something that already exists" I don't think open office/libre office etc have access to the source code for MS office and if they did MS would be on them like a rash.
- lawlessone 9mo agoBlogspam.
- catoAppreciator 9mo agoblogslop
- Herring 9mo agoMe and Opus have a lot in common. We both hit our weekly limit on Monday at 10am.
- michaelsalim 9mo agoI use pay as you go for this very reason, so the limit is my pocket haha. It does make me conscious to keep it under $20 per month though.
- square_usual 9mo agoYou're overpaying by a factor of 4, easily. I use `ccusage`'s statusline in claude code, and even with my personal $20/mo subscription I don't think there's been a single month where I didn't touch ~$80 of usage. I wasn't even abusing it as bad as some people tend to.
- port3000 9mo agoHow do you manage that? /ccusage and --ccusage no longer work for me, I can only see the usage bars in /usage
- square_usual 9mo agoI followed this: https://ccusage.com/guide/statusline https://ccusage.com/guide/statusline
- port3000 9mo agoAh thanks, didn't realise it was a 3rd party library, thought it was a claude native command
- theshrike79 9mo agoYou can use both btw. Get the $20 plan and turn on "extra usage" in billing. Then you can use the basic plan first and if it runs out, it uses token-based billing for the overflow.
- kachapopopow 9mo agoIt's also the feeling I have, opus is not a ground-breaking model by any means. However, Opus 4.5 is incredible when you give it everything it needs, a direction, what you have versus what you want and it will make it work, really, it will work. The code might me ugly, undesirable, would only work for that one condition, but with futher prompting you can evolve it and produce something that you can be proud of. Opus is only as good as the user and the tools the user gives to it. Hmm, that's starting to sound kind-of... human...
- manmal 9mo agoOff/nearshoring regularly produces worse code. I’ve seen it first hand.
- edg5000 9mo agoOpus can produce beatiful code. It can outcode a good programmer. But getting it to do this reliably is something I've gotten better at over the last year; it's a skill that took quite a bit of practice. I now write very long specifications and this helps. I haven't figured out a bulletproof workflow, I think that will take years. But I often get just amazing code out of it.
- kachapopopow 9mo agothere is a big difference between a good programmer and a programmer that gives a shit so I disagree, opus can not come close to the code quality that someone can create and at that point it is the person behind the wheel that is causing the good quality to manifest rather than the AI randomly stumbling upon it.
- minimaxir 9mo agoSee also: a post from a couple days ago which came to the same conclusion that Opus 4.5 is an inflection point above Sonnet 4.5 despite that conclusion being counterintuitive: https://news.ycombinator.com/item?id=46495539 https://news.ycombinator.com/item?id=46495539 It's hard to say if Opus 4.5 itself will change everything given the cost/latency issues, but now that all the labs will have very good synthetic agentic data thanks to Opus 4.5, I will be very interested to see what the LLMs release this year will be able to do. A Sonnet 4.7 that can do agentic coding as well as Opus 4.5 but at Sonnet's speed/price would be the real gamechanger: with Claude Code on the $20/mo plan, you can barely do more than one or two prompts with Opus 4.5 per session.
- on_the_train 9mo agoOh another run of new small apps. Why not unleash this oh so powerful tools not on a jira ticket written two years ago, targeting 3 different repos in an old legacy moloch, like actual work? It's always just the "Fibonacci" equivalent
- _cenw 9mo agoDid some of that today. Extracting logic from Helm templates that read like 2000s PHP and moving it to a nushell script rendering values. Took a lot of guidance both in terms of making it test its own code and architectural/style decisions and I also use Sonnet, but it got there.
- honeycrispy 9mo agoA couple weeks ago I had Opus 4.5 go over my project and improve anything it could find. It "worked" but the architecture decisions it made were baffling, and had many, many bugs. I had to rewrite half of the code. I'm not an AI hater, I love AI for tests, finding bugs, and small chores. Opus is great for specific, targeted tasks. But don't ask it to do any general architecture, because you'll be soon to regret it.
- thousand_nights 9mo agothese models work best when you know what you want to achieve and it helps you get there while you guide it. "Improve anything you can find" sounds like you didn't really know
- mcv 9mo agoAs a tool to help developers I think it's really useful. It's great at stuff people are bad at, and bad at stuff people are good at. Use it as a tool, not a replacement.
- suzzer99 9mo ago"Improve anything you can find" is like going to your mechanic and saying "I'm going on a long road trip, can you tell me anything that needs to be fixed?" They're going to find a lot of stuff to fix.
- blub 9mo agoDoing a vehicle check-up is a pretty normal thing to do, although in my case the mandatory (EU law) periodic ones are happening often enough that I generally don’t have to schedule something out of turn. The few times I did go to a shop and ask for a check-up they didn’t find anything. Just an anecdote.
- oncallthrow 9mo agoIn my experience these models (including opus) aren’t very good at “improving” existing code. I’m not exactly sure why, because the code they produce themselves is generally excellent.
- manmal 9mo agoIMO codex produces working code slowly, while Opus produces superficially working code quickly. I like using Opus to drive codex sessions and checking its output. Clawdbot is really good at that but a long running Claude Code session with codex as sub agents should work well also. The above is for vibe coding; for taking the wheel, I can only use Opus because I suck at prompting codex (it needs very specific instructions), and codex is also way too slow for pair programming.
- NitpickLawyer 9mo ago> I like using Opus to drive codex sessions and checking its output. Why not the other way around? Have the quick brown fox churn out code, and have codex review it, guide changes, and loop? I've actually gone one step further down the delegation. I use opus/gemini3 for plan, review, edit plan for a few steps. Then write it out to .md files. Then have GLM implement it (I got a cheap plan for like 28$ for a year on Christmas). Then have the code this produced reviewed and fixed if needed by opus. Final review by codex (for some reason it's very good at review, esp if you have solid checkboxes for it to check during review). Seems to work so far.
- manmal 9mo agoI agree, codex is great at reviewing as well. I think that’s because code is the ideal description of what we want to achieve, and codex is good (only) when it knows what must be achieved, as verbosely as possible. Currently I don’t let GLM or Opus near my codebases unsupervised because I’m convinced that the better the foundation, the better the end result will be. Is the first draft not pretty crappy with GLM?
- OldGreenYodaGPT 9mo agoMost software engineers are seriously sleeping on how good LLM agents are right now, especially something like Claude Code. Once you’ve got Claude Code set up, you can point it at your codebase, have it learn your conventions, pull in best practices, and refine everything until it’s basically operating like a super-powered teammate. The real unlock is building a solid set of reusable “skills” plus a few agents for the stuff you do all the time. For example, we have a custom UI library, and Claude Code has a skill that explains exactly how to use it. Same for how we write Storybooks, how we structure APIs, and basically how we want everything done in our repo. So when it generates code, it already matches our patterns and standards out of the box. We also had Claude Code create a bunch of ESLint automation, including custom ESLint rules and lint checks that catch and auto-handle a lot of stuff before it even hits review. Then we take it further: we have a deep code review agent Claude Code runs after changes are made. And when a PR goes up, we have another Claude Code agent that does a full PR review, following a detailed markdown checklist we’ve written for it. On top of that, we’ve got like five other Claude Code GitHub workflow agents that run on a schedule. One of them reads all commits from the last month and makes sure docs are still aligned. Another checks for gaps in end-to-end coverage. Stuff like that. A ton of maintenance and quality work is just… automated. It runs ridiculously smoothly. We even use Claude Code for ticket triage. It reads the ticket, digs into the codebase, and leaves a comment with what it thinks should be done. So when an engineer picks it up, they’re basically starting halfway through already. There is so much low-hanging fruit here that it honestly blows my mind people aren’t all over it. 2026 is going to be a wake-up call. (used voice to text then had claude reword, I am lazy and not gonna hand write it all for yall sorry!) Edit: made an example repo for ya https://github.com/ChrisWiles/claude-code-showcase https://github.com/ChrisWiles/claude-code-showcase
- dmbche 9mo agoOh! An ad!
- OldGreenYodaGPT 9mo agolol does sound like and ad, but is true. Also forgot about hooks use hooks too! I just use voice to text then had claude reword it. Still my real world ideas
- mpalmer 9mo agoAfter reading that article, I see at least one thing that Opus 4.5 is clearly not going to change. There is no fixed truth regarding what an "app" is, does, or looks like. Let alone the device it runs on or the technology it uses. But to an LLM, there are only fixed truths (and in my experience, only three or four possible families of design for an application). Opus 4.5 produces correct code more often, but when the human at the keyboard is trying to avoid making any engineering decisions, the code will continue to be boring.
- NewsaHackO 9mo ago>the code will continue to be boring. Why would you not want you code to be boring?
- orthoxerox 9mo agoWhat's the best coding agent you can run locally? How far behind Opus 4.5 is it?
- Tiberium 9mo agoThe best is probably something like GLM 4.7/Minimax M2.1, and those are probably at most Sonnet 4 level, which is behind Opus 4.1, which is behind Sonnet 4.5, which is behind Opus 4.5 ;) And honestly Opus 4.5 is a visible step change above previous Anthropic models.
- orthoxerox 9mo agoDoes it even fit into a 5090 or a Ryzen 395+?
- Tiberium 9mo agoOh, of course not, you might need up to 100GB VRAM to have those models at decent speeds even just for low-quant versions. And all the hype about Macs with unified memory is a bit dishonest because the actual generation speed will be very bad, especially if you fill the context. One of the things that makes Opus 4.5 special in comparison to e.g. GPT 5.2 is the fact that it doesn't have to reason for multiple minutes to make some simple changes.
- torben-friis 9mo agoDo we have an estimate for how much they cost to run? Or in other words, how much are they financing the end user cost? Not only the energy fuel but the hardware’s percentage of cost.
- _se 9mo agoIt's always fun to ask Opus what it thinks about articles like this. Here's what I got with no history or system prompt: https://burkeholland.github.io/posts/opus-4-5-change-everything/ https://burkeholland.github.io/posts/opus-4-5-change-everyth... Read this article and ultrathink critically about it. Provide your perspective. The article makes a strong experiential case for improved velocity in AI-assisted development but contains several argumentation weaknesses and conflations worth examining. The projects described are legitimately non-trivial: Firebase backend integration, Facebook OAuth, iOS apps in Swift (a language the author doesn't know), GitHub Actions pipelines, scheduled cloud functions. Getting these working in hours rather than weeks represents a real capability shift. The author is honest about his limitations and uncertainties, particularly the security concerns. Where the argument breaks down: 1. "Replace developers" vs "dramatically augment developers" The author's own workflow contradicts the headline claim. He's still: Making architectural decisions (choosing Firebase) Handling errors Opus couldn't see (XAML via Visual Studio) Writing custom prompts to shape output quality Manually auditing security Making product and UX decisions This is developer work. The tool changed; the role didn't disappear. 2. The 80% security confidence undermines his thesis He admits he's shipping apps with "80% confidence" in security and calls it "too damn low." This is the crux: the AI accelerated production but didn't replace the judgment required to responsibly ship production software. The velocity gain exposed a competence gap rather than closing it. 3. Sample bias in project selection All examples are: Greenfield (no existing codebase) Single developer Personal/family use Standard patterns with excellent documentation (Firebase, SwiftUI, React Native) No regulatory, compliance, or scale requirements No team collaboration or handoff considerations These constraints define a specific problem space where AI excels. Extrapolating to "developers are replaceable" ignores the majority of professional software work. 4. "Code doesn't need human readability" is underbaked His argument is circular: "Why optimize for human readability when the AI is doing all the work?" But: His 80% security confidence exists because he can't read the code He had to use external tools (VS) when Opus couldn't diagnose errors What happens when context windows are exceeded and the LLM loses track? Model behavior changes between versions; human-readable code is version-agnostic The custom prompt he shares actually encodes many good engineering practices (minimal coupling, explicit state, linear control flow) that benefit LLMs and humans. The "no comments needed" claim conflates what's optimal for LLM regeneration with what's optimal for debugging production issues at 3am. What's actually being demonstrated The honest version of this article would be: Opus 4.5 dramatically compresses the gap between "can write code" and "can ship a personal app" for a specific class of greenfield projects. That's genuinely transformative for hobbyists, indie developers, and people solving their own problems. But that's different from "replacing developers." The article demonstrates a power tool; power tools don't eliminate tradespeople.
- jcmfernandes 9mo agoTo the author: you wrote those apps. Not like you used to, but you wrote them. IMO, our jobs are safe. It's our ways of working that are changing. Rapidly.
- Hammershaft 9mo agoSWE jobs are in fact, not safe, if vaguely defined specifications can be translated into functioning applications. I don't think agents are good enough to do that in larger applications yet, but it is something to consider.
- jcmfernandes 9mo agoDepends on the software. IMO, development speed will increase, but humans will continue to be the limiting factor, so we are safe. Our jobs, however, are changing and will continue to.
- Workaccount2 9mo agoAnthropic dropped out of the general "AGI" race and seems to be purely focused on coding, maybe racing to get the first "automated machine learning programmer". Whatever the case, it seems to be paying (coding) dividends to just be focusing on coding.
- ethbr1 9mo agoThe benefit of focusing on coding is that it has an attractive non-deterministic / deterministic problem split. In that it's using a non-deterministic machine to build a deterministic one. Which gives all the benefits of determinism in production, with all the benefits of non-deterministic creativity in development. Imho, Anthropic is pretty smart in picking it as a core focus.
- dmarwicke 9mo agothis is just optimizing for token windows. flat code = less context. we did the same thing with java when memory was expensive, called it "lightweight frameworks"
- yardie 9mo agoThese are very simple utilities. I expect AI to be able to build them easily. Maybe in a few years it will be able to write a complete photo editor or CAD application from first principles.
- fragmede 9mo agoThen we're really screwed!
- deleted 9mo ago[deleted]
- pyuser583 9mo agoOpus 4.5 burns through tokens really fast.
- jghn 9mo agoI've been noticing it's more on par with sonnet these days. I don't know if that means Opus is getting more efficient, sonnet getting less efficient, or perhaps Opus is getting to the answer fast enough to overcome the higher token spend.
- mcv 9mo agoI've noticed. I'm already through 48% of my quota for this month.
- oncallthrow 9mo agoYeah Opus 4.5 is a massive step change in my experience. I feel like I’m working with a peer, not a junior I’m having to direct. I can give it highly ambiguous and poorly specified tasks and it… just does it. I will note that my experience varies slightly by language though. I’ve found it’s not as good at typescript.
- christophilus 9mo agoIt’s excellent at typescript in my experience. It’s also way better than I am at finding bits of code for reuse. I tell it, “I think I wrote this thing a while back, but it may never have been merged, so you may need to search git history.” And presto, it finds it.
- llmslave2 9mo ago> I feel like I’m working with a peer, not a junior I’m having to direct. I think this says a lot.
- Krei-se 9mo agoIf its a peer to you now the Ai has evolved while you didn't
- Snuggly73 9mo agoOk, if its almighty, then why is not the benchmarks at 100%? If you look at the individual issues, those are somewhat small and trivial changes in existing codebases. https://swe-rebench.com/ https://swe-rebench.com/ (note that if you look at individual slices, Opus is getting often outperformed by Sonnet).
- mcv 9mo agoOpus 4.5 ate through my Copilot quota last month, and it's already halfway through it for this month. I've used it a lot, for really complex code. And my conclusion is: it's still not as smart as a good human programmer. It frequently got stuck, went down wrong paths, ignored what I told it to do to do something wrong, or even repeat a previous mistake I had to correct. Yet in other ways, it's unbelievably good. I can give it a directory full of code to analyze, and it can tell me it's an implementation of Kozo Sugiyama's dagre graph layout algorithm, and immediately identify the file with the error. That's unbelievably impressive. Unfortunately it can't fix the error. The error was one of the many errors it made during previous sessions. So my verdict is that it's great for code analysis, and it's fantastic for injecting some book knowledge on complex topics into your programming, but it can't tackle those complex problems by itself. Yesterday and today I was upgrading a bunch of unit tests because of a dependency upgrade, and while it was occasionally very helpful, it also regularly got stuck. I got a lot more done than usual in the same time, but I do wonder if it wasn't too much. Wasn't there an easier way to do this? I didn't look for it, because every step of the way, Opus's solution seemed obvious and easy, and I had no idea how deep a pit it was getting me into. I should have been more critical of the direction it was pointing to.
- zmmmmm 9mo agoyes just using AI for code analysis is way under appreciated I think. Even the most sceptical people on using it for coding should try it out as a tool for Q&A style code interrogation as well as generating documentation. I would say it zero-shots documentation generation better than most human efforts would to the point it begs the question of whether it's worth having the documentation in the first place. Obviously it can make mistakes but I would say they are below the threshold of human mistakes from what I've seen.
- sfink 9mo ago(I haven't used AI much, so feel free to ignore me.) This is one thing I've tried using it for, and I've found this to be very, very tricky. At first glance, it seems unbelievably good. The comments read well, they seem correct, and they even include some very non-obvious information. But almost every time I sit down and really think about a comment that includes any of that more complex analysis, I end up discarding it. Often, it's right but it's missing the point, in a way that will lead a reader astray. It's subtle and I really ought to dig up an example, but I'm unable to find the session I'm thinking about. This was with ChatGPT 5, fwiw. It's totally possible that other models do better. (Or even newer ChatGPT; this was very early on in 5.) Code review is similar. It comes up with clever chains of reasoning for why something is problematic, and initially convinces me. But when I dig into it, the review comment ends up not applying. It could also be the specific codebase I'm using this on? (It's the SpiderMonkey source.)
- poisonborz 9mo agoI see these posts left and right but no one mentions the _actual_ thing developers are hired for, responsibility. You could use whatever tools to aid coding already, even copy paste from StackOverflow or take whole boilerplate projects from Github already. No AI will take responsibility for code or fix a burning issue that arises because of it. The amount of "responsibility takers" also increases linearly with the size of the codebase / amount of projects.
- simonw 9mo agoThat's quickly becoming the most important part of our jobs - we're the ones with agency and the ability to take responsibility for the work we are producing. I'm fine with contributed AI-generated code if someone who's skills I respect is willing to stake their reputation on that code being good.
- emseetech 9mo agoWhich is why I'm more comfortable using AI as an editor/reviewer than as a writer. I'll write the code, it can help me explore options, find potential problems and suggest tests, but I'll write the code.
- g-mork 9mo agoWe still do that, it's just that realtime code review basically becomes the default mode. That's not to say it's not obvious there will not be a lot less of us in future. I vibed about 80% of a SaaS at the weekend with a very novel piece of hand-written code at the centre of it, just didn't want to bother with the rest. I think that ratio is about on target for now. If the models continue to improve (although that seems relatively unlikely with current architectures and input data sets), I expect that could easily keep climbing. I just cutpasted a technical spec I wrote 22 years ago I spent months on for a language I never got around to building out, Opus zero-shotted a parser, complete with tests and examples in 3 minutes. I cutpasted the parser into a new session and asked it to write concept documentation and a language reference, and it did. The best part is after asking it to produce uses of the language, it's clear the aesthetics are total garbage in practice. Told friends for years long in advance that we were coal miners, and I'll tell you the same thing. Embrace it and adapt
- lagniappe 9mo agoTitle is: "Opus 4.5 is going to change everything"
- minimaxir 9mo agoA Hacker News moderator likely changed the title because it's uninformatively vague.
- simonw 9mo agoOpus 4.5 really is something else. I've been having a ton of fun throwing absurdly difficult problems at it recently and it keeps on surprising me. A JavaScript interpreter written in Python? How about a WebAssembly runtime in Python? How about porting BurntSushi's absurdly great Rust optimized string search routines to C and making them faster? And these are mostly just casual experiments, often run from my phone!
- ronsor 9mo agoOne of my first tests with it was "Write a Python 3 interpreter in JavaScript." It produced tests, then wrote the interpreter, then ran the tests and worked until all of them passed. I was genuinely surprised that it worked.
- wubrr 9mo agoIt's ability to test/iterate and debug issues is pretty impressive. Though it seems to work best when context is minimized. Once the code passes a certain complexity/size it starts making very silly errors quite often - the same exact code it wrote in a smaller context will come out with random obvious typos like missing spaces between tokens. At one point it started writing the code backwards (first line at the bottom of the file, last line at the top) :O.
- Calavar 9mo agoThere are multiple Python 3 interpreters written in JavaScript that were very likely included in the training data. For example [1] [2] [3] I once gave Claude (Opus 3.5) a problem that I thought was for sure too difficult for an LLM, and much to my surprise it spat out a very convincing solution. The surprising part was I was already familiar with the solution - because it was almost a direct copy/paste (uncredited) from a blog post that I read only a few hours earlier. If I hadn't read that blog post, I would have been none the wiser that copy/pasting Claude's output would be potential IP theft. I would have to imagine that LLMs solve a lot of in-training-set problems this way and people never realize they are dealing with a copyright/licensing minefield. A more interesting and convincing task would be to write a Python 3 interpeter in JavaScript that uses register based bytecode instead of stack based, supports optimizing the bytecode by inlining procedures and constant folding, and never allocates memory (all work is done in a single user provided preallocated buffer). This would require integrating multiple disparate coding concepts and not regurgitating prior art from the training data [1] https://github.com/skulpt/skulpt https://github.com/skulpt/skulpt [2] https://github.com/brython-dev/brython https://github.com/brython-dev/brython [3] https://github.com/yzyzsun/PyJS https://github.com/yzyzsun/PyJS
- vl 9mo agoHonestly, I don’t understand universal praise for Opus 4.5. It’s good, but really not better than other agents. Just today: Opus 4.5 Extended Thinking designed psql schema for “stream updates after snapshot” with bugs. Grok Heavy gave correct solution without explanations. ChatGPT 5.2 Pro gave correct solution and also explained why simpler way wouldn’t work.
- giancarlostoro 9mo agoAre you using Claude Code? Because that might be the secret cause you're missing. With Claude Code I can instruct it to validate things after its done with code, and usually it finds that it goofed. I can also tell it to work on like five different things, and go "hey spin up some agents to work on this" and it will spawn 5 agents in parallel to work on said things. I've basically ditched Groke et al and I refuse to give Sam Altman a penny.
- vl 9mo agoFor schema design phase I used web UI for all three. Logical bug of using BIGSERIAL for tracking updates (generated at insert time, not commit time, so can be out of order) wouldn’t be caught by any number of iterations of Claude Code and would be found in production after weeks of debugging.
- simonw 9mo agoAt this point having any LLM write code without giving it an environment that allows it to execute that code itself is like rolling a heavily-biased random number generator and hoping you get a useful result. Things get so much more interesting when they're able to execute the code they are writing to see if it actually works.
- fragmede 9mo agoSo much this. Do we program by writing reams of code and never running the compiler until it's all written and then judging the programmer as terrible when it doesn't compile? Or do we write code by hand incrementally and compile and test as we go along? So why would do we think having the AI do that and fail is setting it up for success? If I wrote code on a whiteboard and was judged for making syntax errors, I'd never have gotten a job. Give the AI the tools it needs to succeed, just like you would for a human.
- karmasimida 9mo agoAs impressive as Opus 4.5 is, it still fails in one situation that it assumes 0-index while the component it supposes to work with assume 1-index. It has access to the said information on disk, but just forgets to look into. Opus 4.5 is incredible, it is the GPT-4 moment for coding because how honest and noticeable the capacity increase is. But still, it has blind spots just like human.
- soulofmischief 9mo agoOpus 4.5 is currently helping me write a novel, comprehensive and highly performant programming language with all of the things I've ever wanted, done in exactly my opinionated way. This project would have taken me years of specialization and research to do right. Opus's strength has been the ability to both speak broadly and also drill down into low-level implementations. I can express an intent, and have some discussion back and forth around various possible designs and implementations to achieve my goals, and then I can be preparing for other tasks while Opus works in the background. I ask Opus to loop me in any time there are decisions to be made, and I ask it to clearly explain things to me. Contrary to losing skills, I feel that I have rapidly gained a lot of knowledge about low-level systems programming. It feels like pair programming with an agentic model has finally become viable. I will be clear though, it takes the steady hand of an experience and attentive senior developer + product designer to understand how to maintain constraints on the system that allow the codebase to grow in a way that is maintainable on the long-term. This is especially important, because the larger the codebase is, the harder it becomes for agentic models to reason holistically about large-scale changes or how new features should properly integrate into the system. If left to its own devices, Opus 4.5 will delete things, change specification, shirk responsibilities in lieu of hacky band-aids, etc. You need to know the stack well so that you can assist with debugging and reasoning about code quality and organization. It is not a panacea. But it's ground-breaking. This is going to be my most productive year in my life. On the flip side though, things are going to change extremely fast once large-scale, profitable infrastructure becomes easily replicable, and spinning up a targeted phishing campaign takes five seconds and a walk around the park. And our workforce will probably start shrinking permanently over the next few years if progress does not hit a wall. Among other things, I do predict we will see a resurgence of smol web communities now that independent web development is becoming much more accessible again, closer to how it when I first got into it back in the early 2000's.
- lawlessone 9mo agoWhy would anyone buy the novel?
- renecito 9mo ago
- deleted 9mo ago[deleted]
- headcanon 9mo agoYep, I literally built this last night with Opus 4.5 after my wife and I challenged each other to a typing competition. I gave it direction and feedback but it wrote all the actual code. Wasn't a one shot (maybe 3-4 shot) but didn't really have to think about it all that hard. https://chronick.github.io/typing-arena/ https://chronick.github.io/typing-arena/ With another more substantial personal project (Eurorack module firmware, almost ready to release), I set up Claude Code to act as a design assistant, where I'd give it feedback on current implementation, and it would go through several rounds of design/review/design/review until I honed it down. It had several good ideas that I wouldn't have thought of otherwise (or at least would have taken me much longer to do). Really excited to do some other projects after this one is done.
- haolez 9mo agoI have a different concern: the SOTA products are expensive and get dumbed down on busy times. My personal strategy has been to be a late follower, where I adopt new AI tools when the competition has caught up with the previous SOTA, and now there are many tools that are cost effective and equally good. Can't wait for when the competition catches up with Claude Code, especially the open source/weights Chinese alternatives :)
- becquerel 9mo agoIf you haven't tried it yet, OpenCode is quite good.
- rcarmo 9mo agoI had a similar set of experiences with GPT 5.x over the holiday break, across somewhat more disparate domains: https://taoofmac.com/space/notes/2025/12/31/1830 https://taoofmac.com/space/notes/2025/12/31/1830 I hacked together a Swift tool to replace a Python automation I had, merged an ARM JIT engine into a 68k emulator, and even got a very decent start on a synth project I’ve been meaning to do for years. What has become immensely apparent to me is that even gpt-5-mini can create decent Go CLI apps provided you write down a coherent spec and review the code as if it was a peer’s pull request (the VS Code base prompts and tooling steer even dumb models through a pretty decent workflow). GPT 5.2 and the codex variants are, to me, every bit as good as Opus but without the groveling and emojis - I can ask it to build an entire CI workflow and it does it in pretty much one shot if I give it the steps I want. So for me at least this model generation is a huge force multiplier (but I’ve always been the type to plan before coding and reason out most of the details before I start, so it might be a matter of method).
- Kerrick 9mo agoGemini 3 Pro (High) via Antigravity has been similarly great recently. So have tools that I imagine call out to these higher-power models: Amp and Junie. In a two-week blur I brought forth the bulk of a Ruby library that includes bindings to the Ratatui rust crate for making TUIs in Ruby. During that time I also brought forth documentation, example applications, build and devops tooling, and significant architectural decisions & roadmaps for the future. It's pretty unbelievable, but it's all there in the git and CI history. https://sr.ht/~kerrick/ratatui_ruby/ https://sr.ht/~kerrick/ratatui_ruby/ I think the following things are true now: - Vibe Coding is, more than ever, "autopilot" in the aviation sense, not the colloquial sense. You have to watch it, you are responsible, the human has do run takeoff/landing (the hard parts), but it significantly eases and reduces risk on a bulk of the work. - The gulf of developer experience between today's frontier tooling and six months ago is huge. I pushed hard to understand and use these tools throughout last year, and spent months discouraged--back to manual coding. Folks need to re-evaluate by trying premium tools, not free ones. - Tooling makers have figured out a lot of neat hacks to work around the limitations of LLMs to make it seem like they're even better than they are. Junie integrates with your IDE, Antigravity has multiple agents maintaining background intel on your project and priorities across chats. Antigravity also compresses contexts and starts new ones without you realizing it, calls to sub-agents to avoid context pollution, and other tricks to auto-manage context. - Unix tools (sed, grep, awk, etc.) and the git CLI (ls-tree, show, --stat, etc.) have been a huge force-multiplier, as they keep the context small compared to raw ingestion of an entire file, allowing the LLMs to get more work done in a smaller context window. - The people who hire programmers are still not capable of Vibe Coding production-quality web apps, even with all these improvements. In fact, I believe today this is less of a risk than I feared 10 months ago. These are advanced tools that need constant steering, and a good eye for architecture, design, developer experience, test quality, etc. is the difference between my vibe coded Ruby [0] (which I heavily stewarded) and my vibe coded Rust [1] (I don't even know what borrow means). [0]: https://git.sr.ht/~kerrick/ratatui_ruby/tree/stable/item/lib https://git.sr.ht/~kerrick/ratatui_ruby/tree/stable/item/lib [1]: https://git.sr.ht/~kerrick/ratatui_ruby/tree/stable/item/ext/ratatui_ruby https://git.sr.ht/~kerrick/ratatui_ruby/tree/stable/item/ext...
- drchiu 9mo agoHaving used Opus 4.5 for the past 5 weeks, I estimate it codes better than 95% of the people I've ever worked with. And it writes with more clarity too. The only people who are complaining about "AI slop" are those whose jobs depend on AI to go away (which it won't).
- dzonga 9mo agowhat strikes me about these posts is they praise models for apps | utilities commonly found on GitHub. ie well known paths based on training data. what's never posted is someone building something that solves a real problem in the real world - that deals with messy data | interfaces. I like a.i to do the common routine tasks that I don't like to do like apply tailwind styles but being renter and faking productivity that's not it
- nirolo 9mo agoI used it with gemini 3 in tandem to build an app to simulate thermal bridges because I want to insulate a house. I explored this in various directions and there are some functionalities not completed or sound, but the main part is good and tested against ISO/DIN test cases for this kind of problem. You can try it here, although the numeric simulations take quite a while in the cloud app https://thermal-bridge.streamlit.app/ https://thermal-bridge.streamlit.app/ Disclaimer: I'm not a programmer or software engineer. I have a background in physics and understand some scripting in python and basic git. The code is messy at the moment because I explored/am still exploring to port it to another framework/language
- fractallyte 9mo agoThat final line: "Disclaimer: This post was written by a human and edited for spelling, grammer by Haiku 4.5" Yeah, GRAMMAR For all the wonderment of the article, tripping up on a penultimate word that was supposedly checked by AI suddenly calls into question everything that went before...
- simonw 9mo agoPresumably that disclaimer was added manually after Haiku had run the checks.
- smusamashah 9mo agoWhat about Sonnet 4.5? I used both Opus and Sonnet on Claude.ai and found sonnet much better at following instructions and doing exactly what was asked. (it was for single html/js PWA to measure and track heart rate) Opus seems to go less deep, does it's own things, do not follow instructions exactly EVEN IF I WROTE ALL CAPS. With Sonnet 4.5 I can understand everything author is saying. May be Opus is optimised for Claude code and Sonnet works best on Web.
- brushfoot 9mo agoI pivoted into integrations in 2022. My day-to-day now is mostly in learning the undocumented quirks of other systems. I turn those into requirements, which I feed to the model du jour via GitHub Copilot Agents. Copilot creates PRs for me to review. I'd say it gets them right the vast majority of the time now. Example: One of my customers (which I got by Reddit posts, cold calls, having a website, and eventually word of mouth) wanted to do something novel with a vendor in my niche. AI doesn't know how to build it because there's no documentation for the interfaces we needed to use.
- deleted 9mo ago[deleted]
- YesBox 9mo agoI've noticed a huge drop in negative comments on HN when discussing LLMs in the last 1-2 months. All the LLM coded projects I've seen shared so far[1] have been tech toys though. I've watched things pop up on my twitter feed (usually games related), then quietly go off air before reaching a gold release (I manually keep up to date with what I've found, so it's not the algorithm). I find this all very interesting: LLMs dont change the fundamental drives needed to build successful products. I feel like I'm observing the TikTokification of software development. I dont know why people aren't finishing. Maybe they stop when the "real work" kicks in. Or maybe they hit the limits of what LLMs can do (so far). Maybe they jump to the next idea to keep chasing the rush. Acquiring context requires real work, and I dont see a way forward to automating that away. And to be clear, context is human needs; i.e. the reasons why someone will use your product. In the game development world, it's very difficult to overstate how much work needs to be done to create a smooth, enjoyable experience for the player. While anyone may be able to create a suite of apps in a weekend, I think very few of them will have the patience and time to maintain them (just like software development before LLMs! i.e. Linux, open source software, etc.). [1] yes, selection bias. There are A LOT of AI devs just marketing their LLMs. Also it's DEFINITELY too early to be certain. Take everything Im saying with a one pound grain of salt.
- simonw 9mo agoIt could be that the people who are focused on building monetizable products with LLMs don't feel the need to share what they are doing - they're too busy quietly getting on with building and marketing their products. Sharing how you're using these tools is quite a lot of work!
- YesBox 9mo agoAgreed! LLMs are a force multiplier for real products too. They're going to augment people who are willing to do the real work. But, Im also wondering if LLMs are going to create a new generation of software dev "brain rot" (to use the colloquial term), similar to short form videos. I should mention in the gamedev world, it's quite common share because sharing is marketing, hence my perspective.
- losvedir 9mo agoI'm kind of surprised how many people are okay with deploying code that hasn't been audited. I read If Anyone Builds It Everyone Dies over the break. The basic premise was that we can't "align" AI so when we turn it loose in an agent loop what it produces isn't necessarily what we want. It may be on the surface, to appease us and pass a cursory inspection, but it could embed other stuff according to other goals. On the whole, I found it a little silly and implausible, but I'm second guessing parts of that response now that I'm seeing more people (this post, the Gas Town thing on the front page earlier) go all-in on vibe coding. There is likely to be a large body of running software out there that will be created by agents and never inspected by humans. I think a more plausible failure mode in the near future (next year or two) is something more like a "worm". Someone building an agent with the explicit instructions to try to replicate itself. Opus 4.5 and GPT 5.2 are good enough that in an agent loop they could pretty thoroughly investigate any system they land on, and try to use a few ways to propagate their agent wrapper.
- btown 9mo agoPerhaps our only saving grace is that many LLMs at varying levels of "dumbness" exist. Is it possible to create an obfuscated quine that exhibits stable detection-avoiding behavior on every frontier model simultaneously, as well as on an old-school classifier and/or GPT-3 era LLM fine-tuned just for worm detection? One incapable of even thinking about what it's seeing, and being persuaded to follow its subtle propagation logic? I'm not sure that the answer is yes. The larger issue to me is less that an LLM can propagate in generated code undetected, but rather that an attacker's generated code may soon be able to execute a level of hyper-customized spear-phishing-assisted attack at scale, targeting sites without large security teams - and that it will be hitting unintentional security flaws introduced by those smaller companies' vibe code. Who needs a worm when you have the resources of a state-level attacker at your fingertips, and numerous ways to monetize? The balance of power is shifting tremendously towards black hats, IMO.
- TacticalCoder 9mo agoWhy think about nefarious intent instead of just user error? In this case LLM error instead of programmer error. Most RCEs, 0-days, and whatnots are not due to the NSA hiding behind the "Jia Tan" pseudo to try to backdoor all the SSH servers on all the systemd [1] Linuxes in the world: they're just programmer errors. I think accidental security holes with LLMs are way, way, way more likely than actual malicious attempts. And with the amount of code spoutted by LLMs, it is indeed --and the lack of audit is-- an issue. [1] I know, I know: it's totally unrelated to systemd. Yet only systems using systemd would have been pwned. If you're pro-systemd you've got your point of view on this but I've got mine and you won't change my mind so don't bother.
- ironbound 9mo agoThis is great can't wait for the future when our VC ideas can become unicorns, without CEO's & Founders..
- torben-friis 9mo ago>Disclaimer: This post was written by a human and edited for spelling, grammer by Haiku 4.5 Either it wasn’t that good, or the author failed in the one phrase they didn’t proofread. (No judgement meant, it’s just funny).
- lifetimerubyist 9mo agoOpus helped me optimized a wonky SQL query today from 4s to 5min. Truly something that only a super intelligence is capable of.
- raldi 9mo agoDespite the abuse of quotation marks in the screenshot at the top of this link, Dario Amodei did not in fact say those words or any other words with the same meaning.
- versteegen 9mo agoYes, unfortunate that people keep perpetuating that misquote. What he actually said was "we are not far from the world—I think we’ll be there in three to six months—where AI is writing 90 percent of the code." https://www.cfr.org/event/ceo-speaker-series-dario-amodei-anthropic https://www.cfr.org/event/ceo-speaker-series-dario-amodei-an...
- mattfrommars 9mo agoI’ve been saying this a countless time, LLM are great to build toy and experimental projects. I’m not shaming but I personally need to know if my sentiment is correct or not or I just don’t know how to use LLMs Can vibe coder gurus create operating system from scratch that competes with Linux and make it generate code that basically isn’t Linux since LLM are trained on said the source code … Also all this on $20 plan. Free and self host solution will be best
- simonw 9mo agoYour bar for being impressed by coding agents is "can build a novel operating system that competes with Linux on a plan that costs a $20/month"? Yeah, they can't do that.
- hollowturtle 9mo agoIn fact, like the author of the comment said, can just generated toys and experimental projects. I'm all in for experiments and exploring ideas, but I have yet to see a great product all vibe coded. All I see is a constand decline in software quality
- fragmede 9mo agoConsider your own emotions and the bias you have against it. If it is actually able to do the things it is hyped up to be, what does that mean for you, your job, and your career? Can you really extract those emotions from how you're approaching the situation? That tiniest bit of fear in your gut might be coloring your approach here. You want a new operating system not based on Linux, that competes with it, because if it is based on Linux, it's in the training data, which means it's cheating? Jrifjxgwyenf! A hammer is a really bad screwdriver. My car is really bad at refrigerating food. If you ask for something outside its training data, it doesn't do a very good job. So don't do that! All of the code on the Internet is a pretty big dataset though, so maybe Claude could do an operating system that isn't Linux that competes with it by laundering the FreeBSD kernel source through the training process. And you're barely even willing to invest any money into this? The first Apple computer cost $4,000 or so. You want the bleeding edge of technology delivered to the smartphone in your hand, for $20, or else it's a complete failure? Buddy, your sentiment isn't the issue, it's your attitude. I'm not here spouting ridiculous claims like AI is going to cure all of the different kinds of cancer by the end of 2027, I just want to say that endlessly contrarian naysayers are as equally borish as the syncophantic hype AIs they're opposing.
- llmslave2 9mo agoI see Anthropics marketing campaign is out in full force today ahead of their IPO.
- avidphantasm 9mo agoCool. Please check back in with us after they’ve raised the price 50x and you can no longer build anything because you are alienated from your tools.
- atonse 9mo agoI’ve said many times, I’d still pay even $1,000 a month for CC. But I’m a business owner so the calculus is different. But I don’t think they’ll raise prices uncontrollably because competition exists. Even just between OpenAI and Anthropic.
- kypro 9mo agoIt's been interesting watching HN shift in my direction on this in recent weeks... I had been saying since around summer of this year that coding agents were getting extremely good. The base model improvements were ok, but the agentic coding wrappers were basically game changers if you were using them right. Until recently they still felt very context limited, but the context problem increasingly feels like a solved problem. I had some arguments on here in the summer about how it was stupid to hire junior devs at this point and how in a few years you probably wouldn't need senior devs for 90% of development tasks either. This was an aggressive prediction 6 months ago, but I think it's way too conservative now. Today we have people at our company who have never written code building and shipping bespoke products. We've also started hiring people who can simply prove they can build products for us using AI in a single day. These are not software engineers because we are paying them wages no SWEs would accept, but it's still a decent wage for a 20 something year old without any real coding skills but who is interested in building stuff. This is something I wouldn't have never of expected to be possible 6 months ago. In 6 months we've gone from senior developers writing ~50% of their code with AI, to just a handful of senior developers who now write close to 90% of their code with AI while they support a bunch of non-developers pumping out a steady stream of shippable products and features. Software engineers and traditional software engineer is genuinely running on borrowed time right now. It's not that there will be no jobs for knowledgable software engineers in the coming years, but companies simply won't need many hotshot SWEs anymore. The companies that are hiring significant numbers of software engineers today simply can not have realised how much things have changed over just the last few months. Apart from the top 1-2% of talent I simply see no good reason to hire a SWE for anything anymore. And honestly outside of niche areas, anyone hand-cracking code today is a dinosaur... A good SWE today should see their job as simply reviewing code and prompting. If you think that the quality of code LLMs produce today isn't up to scratch you've either not used the latest models and tools or you're using them wrong. That's not to say it's the best code – they still have a tendency to overcomplicate things in my opinion – but it's probably better than the average senior software engineer. And that's really all that matters. I'm writing this because if you're reading this thinking we're basically still in 2024 with slightly better models and tooling you're just wrong and you're probably not prepared for what's coming.
- 9mo ago
- artdigital 9mo agoI switched my subscription from Claude to ChatGPT around 5.0 when SOTA was Sonnet 4.5 and found GPT-5-high (and now 5.2-high) so incredibly good, I could never imagine Opus is on its level. I give gpt-5.2-high a spec, it works for 20 minutes and the result is almost perfect and tested. I very rarely have to make changes. It never duplicates code, implements something again and leaves the old code around, breaks my convention, hallucinates, or tells me it’s done when the code doesn’t even compile, which sonnet 4.5 and Opus 4.1 did all the time I’m wondering if this had changed with Opus 4.5 since so many people are raving about it now. What’s your experience? Claude - fast, to the point but maybe only 85% - 90% there and needs closer observation while it works GPT-x-high (or xhigh) - you tell it what to do, it will work slowly but precise and the solution is exactly what you want. 98% there, needs no supervision
- hsn915 9mo agoI had a similar feeling expressed in the title regarding ChatGPT 5.2 I haven't tried it for coding. I'm just talking about regular chatting. It's doing something different from prior models. It seems like it can maintain structural coherence even for very long chats. Where as prior models felt like System 1 thinking, ChatGPT5.2 appears like it exhibits System 2 thinking.
- chris_st 9mo agoI've found asking GPT-5.2 High to review Opus 4.5's code to be really productive. They find different things.
- multisport 9mo agoWhat bothers me about posts like this is: mid-level engineers are not tasked with atomic, greenfield projects. If all an engineer did all day was build apps from scratch, with no expectation that others may come along and extend, build on top of, or depend on, then sure, Opus 4.5 could replace them. The hard thing about engineering is not "building a thing that works", its building it the right way, in an easily understood way, in a way that's easily extensible. No doubt I could give Opus 4.5 "build be a XYZ app" and it will do well. But day to day, when I ask it "build me this feature" it uses strange abstractions, and often requires several attempts on my part to do it in the way I consider "right". Any non-technical person might read that and go "if it works it works" but any reasonable engineer will know that thats not enough.
- whatever1 9mo agoTheir thesis is that code quality does not matter as it is now a cheap commodity. As long as it passes the tests today it's great. If we need to refactor the whole goddamn app tomorrow, no problem, we will just pay up the credits and do it in a few hours.
- multisport 9mo agoYes agreed, and tbh even if that thesis is wrong, what does it matter?
- whatever1 9mo agoThe whole point of good engineering was not about just hitting the hard specs, but also have extendable, readable, maintainable code. But if today it’s so cheap to generate new code that meets updated specs, why care about the quality of the code itself? Maybe the engineering work today is to review specs and tests and let LLMs do whatever behind the scenes to hit the specs. If the specs change, just start from scratch.
- andrekandre 9mo ago> let LLMs do whatever behind the scenes to hit the specs assuming for the sake of argument that's completely true, then what happens to "competitive advantage" in this scenario? it gets me thinking: if anyone can vibe from spec, whats stopping company a (or even user a) from telling an llm agent "duplicate every aspect of this service in python and deploy it to my aws account xyz"... in that scenario, why even have companies?
- deleted 9mo ago[deleted]
- bennydog224 9mo agoDon't want to discredit Opus at all, it's easy at directed tasks but it's not the silver bullet yet. It is best in its class, but trips up frequently with complicated engineering tasks involving dynamic variables. Think: Browser page loading, designing for a system where it will "forget" to account for race conditions, etc. Still, this gets me very excited for the next generation of models from Anthropic for heavy tasks.
- _pdp_ 9mo agoYEP Things are changing. Now everyone can build bespoke apps. Are these apps pushing the limits of technology? No! But they work for the very narrow and specific domain they where designed. And yes they do not scale and have as much bugs as your personal shell scripts. But they work. But let's not compare these with something more advance - at least not yet. Maybe by end of this year? We switched from Sonnet 4.5 to Opus 4.5 as our default coding agent recently and we pay the price for the switch (3x the cost) but as the OP said, it is quite frankly amazing. It does a pretty good job, especially, especially when your code and project is structured in a such a way that it helps the agent perform well. Anthropic released an entire video on the subject recently which aligns with my own observations as well. Where it fails hard is in the more subtle areas of the code, like good design, best practices, good taste, dry, etc. We often need to prompt it to refactor things as the quick solution it decided to do is not in our best interest for the long run. It often ends in deep investigations about things which are trivially obvious. It is overfitted to use unix tools in their pure form as it fail to remember (even with prompting) that it should run `pnpm test:unit` instead `npx jest` - it gets it wrong every time. But when it works - it is wonderful. I think we are at the point where we are close to self-improving software and I don't mean this lightly. It turns out the unix philosophy runs deep. We are right now working on ways to give our agents more shells and we are frankly a few iterations there. I am not sure what to expect after this but I think whatever it is, it will be interesting to see.
- squirrellous 9mo agoFor some reason Opus 4.5 is blowing up recently after having been released for weeks. I guess because holidays are over? Active agent users should have discovered this for a while.
- adithyassekhar 9mo agoI like writing code
- hannofcart 9mo agoTo the sceptics still saying that LLMs still can't solve "slime mold pathing algorithm and creating completely new shoe-lacing patterns" (literally a quote from a different comment here), please consider something we've learnt over and over again in history: good enough and cheap will destroy perfect but expensive. And then cheap and good enough option will eventually get better because that's the one that is more used. It's how Japanese manufacturing beat Western manufacturing. And how Chinese manufacturing then beat Japanese again. It's why it's much more likely you are using the Linux kernel and not GNU hurd. It's how digital cameras left traditional film based cameras in the dust. Bet on the cheaper and good enough outcome. Bet against it at your peril.
- arielweisberg 9mo agoI agree. Claude Code went from being slower than doing it myself to being on average faster, but also far less exhausting so I can do more things in general while it works.
- Sxubas 9mo agoJust an open thought, what if most improvement we are seeing is not mostly due to LLM improvements but to context management and better prompting? Ofc the reality is a mix of both, but really curious on what contributes more. Probably just using cursor with old models (eww) can yield a quick response.
- daxfohl 9mo agoThis resonates with my experience in codex 5.2, at least directionally. I'm pretty persnickety about code itself, so I'm not to the point where I'll just let it rip. But in the last month or two things have gone from "I'll ask on the web interface and maybe copy some code into the project", to trusting the agent and getting a reasonable starting point about half the time. > because models like to write code WAY more than they like to delete it Yeah, this is the big one. I haven't figured it out either. New or changing requirements are almost always implemented a flurry of if/else branches all over the place, rather than taking the time for a step back and a reimagining of a cohesive integration of old and new. I've had occasional luck asking for this explicitly, but far more frequently they'll respond with recommendations that are far more mechanical, e.g. "you could extract a function for these two lines of code that you repeat twice", not architectural, in nature. (I still find pasting a bunch of files into the chat interface and iterating on refinements conversationally to be faster and produce better results). That said, I'm convinced now that it'll get there sooner or later. At that point, I really don't know what purpose SWEs will serve. For a while we might serve as go-betweens between the coding agent and PMs, but LLMs are already way better at translating from tech jargon to human, so I can't imagine it would be long before product starts bypassing us and talking directly to the agents, who (err, which) can respond with various design alternatives, pros and cons of each, identify all the dependencies, possible compatibility concerns, alignment with future direction, migration time, compute cost, user education and adoption tracking, etc, all in real time in fluent PM-ese. IDK what value I add to that equation. For the last year or so I figured we'd probably hit a wall before AI got to that point, but over the last month or so, I'm convinced it's only a matter of time.
- bigcloud1299 9mo agoOh shit your UI looks exactly 100% like mine.
- MarsIronPI 9mo agoIt worries me that the best models, the ones that can one-shot apps and such, are all non-free and owned by companies who can't be trusted to have end-users' best interests at heart. It would be greatly reassuring to see a self-hostable model that can compete with Opus 4.5 and Gemini 3 at such coding tasks.
- theappsecguy 9mo agoIt’s incredibly tiring to see this narrative peddled every damn day. I use opus 4.5 every day. It’s not much different than any previous models, still does dumb things all the time.
- gpm 9mo agoSame experience - I've had it fail at the same reasonably simple tasks I had opus 4 and sonnet 4.5 and sonnet 4 fail at when they aren't carefully guided and their work check and fixed...
- takinola 9mo agoI guess the best analogy I can think of is the transition from writing assembly language and the introduction of compilers. Now, (almost) no one knows, or cares, what comes out of the compiler. We just assume it is optimized and that it represents the source code faithfully. Seems like code might go that way too and people will focus on the right prompts and can simply assume the code will be correct.
- dpacmittal 9mo agoA compiler is deterministic though.
- becquerel 9mo agoDoes a system being deterministic really matter if it's complex enough you can't predict it? How many stories are there about 'you need to do it in this specific way, and not this other specific way, to get 500x better codegen'?
- jedberg 9mo agoI had an app I wanted for over a decade. I even wrote a prototype 10 years ago. It was fine but wasn't good enough to use, so I didn't use it. This weekend I explained to Claude what I wanted the app to do, and then gave it the crappy code I wrote 10 years ago as a starting point. It made the app exactly as I described it the first time. From there, now that I had a working app that I liked, I iterated a few times to add new features. Only once did it not get it correct, and I had to tell it what I thought the problem was (that it made the viewport too small). And after that it was working again. I did in 30 minutes with Claude what I had try to do in a few hours previously. Where it got stuck however was when I asked it to convert it to a screensaver for the Mac. It just had no idea what to do. But that was Claude on the web, not Claude Code. I'm going to try it with CC and see if I can get it. I also did the same thing with a Chrome plugin for Gmail. Something I've wanted for nearly 20 years, and could never figure out how to do (basically sort by sender). I got Opus 4.5 to make me a plugin to do it and it only took a few iterations. I look forward to finally getting all those small apps and plugins I've wanted forever.
- gabriel-uribe 9mo agoThis reminds me of how much screensavers on Mac are a PITA. But yes, such a boon for us doodad makers.
- jedberg 9mo agoAnd dads who just don't have time to make doodads like we used to!
- firemelt 9mo agowhat plan do you have on claude?
- jedberg 9mo agoThe cheapest one above free.
- oldnewthing 9mo agoClaude Code is very good; good enough that I upgraded to the Max plan this week. However, it has a long way to go. It's great at one-shotting (with iterations) most ideas. However, it doesn't do as well when the task is complicated in an existing codebase. This weekend I migrated the backend for the SaaS I am building from Python to .NET Core. It did the migration but completely missed the conventions that the frontend was using to call the backend. While the converion itself went OK, every user journey was broken. I am still manually testing every code path and feeding in the errors to get Claude to fix it. My instructions were fairly comprehensive but Claude still missed most of it. My fault that I didn't generate tests first, but after this migration that's my first task.
- DustinBrett 9mo agoPost the code open source and run it on prod.
- thallukrish 9mo agoWhen complexity increases, you end up handholding them in pieces.
- overgard 9mo agoUgh, I'm so sick of these "I can use AI to solve an already solved problem, thus programmers aren't relevant." Note the solved problem part. This isn't convincing except to people that want a (bad) argument to depress wages and lay off workers while making the existing seniors take on more and more work. This is overall bad for the industry.
- emodendroket 9mo agoAren't most products that actually ship some kind of "solved problem" though?
- jackdoe 9mo agomost of software engineering was rational, now it is becoming empirical it is quite strange, you have to make it write the code in a way it can reason about it without it reading it, you also have to feel the code without reading all of it. like a blind man feeling the shape of an object; Shape from Darkness you can ask opus to make a car, it will give you a car, then you ask it for navigation; no problem, it uses google maps works perfect then you ask it to improve the breaks, and it will give internet to the tires and the break pedal, and the pedal will send a signal via ipv6 to the tires which will enable a very well designed local breaking system, why not, we already have internet for google maps. i think the new software engineering is 10 times harder than the old one :)
- gogasca 9mo ago[dead]
- yolkedgeek 9mo agoI really can't tell if this is satire or not
- exabrial 9mo agoWhat is with all the Claude spam lately on hn?
- jdthedisciple 9mo agoTo those of you who use it: How much does Claude Code cost you a month on avg? I only use VS Code with Copilot subscription ($10) and already get quite a lot out of it. My experience is that Claude Code really drains your pocket extremely fast.
- rleigh 9mo agoI started on the cheapest £15/mo "Pro" plan and it was great for home use when I'd do a bit of coding in the evenings only, but it wasn't really that usable with Opus--you can burn through your session allowance in a few minutes, but was fine with Sonnet. I used the PAYG option to add more, but cost me £200 in December, so I opted for the £90/mo "Max" plan which is great. I've used Opus 4.5 continuously and it's done great work. I think when you look at it from the perspective of how much you get out of it compared with paying a human to do the same (including yourself), it is still very good value for money whether you use it for work or for your own projects. I do both. But when I look what I can now do for my own projects including open-source stuff, I'm very time-limited, and some of the things I want to do would take multiple years. Some of these tools can take that down to weeks, do I can do more with less, and from that perspective the cost is worth it.
- weatherlite 9mo agoThe main issue in this discussion is the word "replace" . People will come up with a bunch of examples where humans are still needed in SWE and can't be fully replaced, that is true. I think claiming that 100% of engineers would be replaced in 2026 is ridiculous. But how about downsizing? Yeah that's quite probable.
- tripledry 9mo agoPutting the performance aside for now as I just started trying out Opus 4.5, can't say too much yet, I don't hype or hate AI as of now, it's simply useful. Time will tell what happens, but if programming becomes "prompt engineering", I'm planning on quitting my job and pivoting to something else. It's nice to get stuff working fast, but AI just sucks the joy out of building for me. Trying to not feel the pressure/anxiety from this, but every time a new model drops there is this tiny moment where I think "Is it actually different this time?"
- seanmcdirmid 9mo agoPity, prompt engineering is just another kind of programming, I find it to be fun, but I guess lots of other people would see it differently.
- tripledry 9mo agoIndeed it is another kind of programming, I simply don't enjoy it. But it is also very early to say, maybe the next iteration of tools will completely change my perspective, I might enjoy it some day!
- friendzis 9mo agoThe venn diagram of engineering and prompting is two circles, maybe a tiny overlap with integrated environments like claude code. A program, by definition, is analyzable and repeatable, whereas prompting is anything but that.
- seanmcdirmid 9mo agoAs long as your program is large and multi-threaded (most programs that matter commercially), it is not very analyzable or repeatable. You replace those qualities with QA and tests, the same is true with prompting.
- friendzis 9mo agoEve if "write code -> run QA -> analyze failures -> rewrite code" is cheaper for most commercial software than thorough upfront formal verification, it works precisely because the programs are analyzable. When the code spit out by an LLM does not pass QA one can merely add "pls fix teh program, bro, pls no mistakes this time, bro, kthxbye", cross their fingers and hope for the best, because in the end it is impossible -- fundamentally -- to determine which part of the prompt produced offending code. While it is indeed an interesting observation that the latter approaches commercial viability in certain areas there is still somewhere between zero and infinitesimal overlap between prompting and engineering.
- qnleigh 9mo agoSo much of the conversation is around these models replacing software engineers. But the use cases described in the article sound like pretty compelling business opportunities; if the custom apps he built for his wife's business have been useful, probably there are lots of businesses that would pay for the service he just provided his wife. Small, custom apps can be made way more cheaply now, so Jeven's paradox says that demand should go up. I think it will. I would love to hear from some freelance programmers how LLMs have changed their work in the last two years.
- ath3nd 9mo ago[dead]
- raesene9 9mo agoOne problem with the idea of making businesses out of this kind of application is actually mentioned in passing in the article "I decided to make up for my dereliction of duties by building her another app for her sign business that would make her life just a bit more delightful - and eliminate two other apps she is currently paying for" OP used Opus to re-write existing applications that his wife was paying for. So now any time you make a commercial app and try to sell it, you're up against everyone with access to Opus or similar tooling who can replicate your application, exactly to their own specifications.
- ensocode 9mo agoso everybody is making their own apps for their specific problem? Sounds as it will get a mess in the end. So maybe it will be more about ideas and concepts and not so much about know how to code.
- raesene9 9mo agoYep vast numbers of personalized apps seems like it would end up being pretty messy. I think the challenge of betting on ideas and concepts is that once you've published something, someone else can take the idea and replicate it easily and cheaply, so it'll be harder to monetize unless you can come up with something that's hard to replicate.
- SergeAx 9mo agoThis article is much better than hundred of similar articles "AI will change software engineering" because it have links to actual products created with said "AI". I can't say they are impressive, but definitely so for laypeople.
- p0w3n3d 9mo agoDisclaimer: This post was written by a human and edited for spelling, grammer by Haiku 4.5 I recently am finishing the reading of Mistborn series, so please do not read further unless you want a spoiler. SPOILER There is a suspicion that mists can change written text. END OF SPOILER So how can we be sure that Haiku didn't change the text in favour of AI then?
- maciejzj 9mo agoI've been on a small adventure of posting more actively on HN since the release of Gemini 3, trying to stir debate around the more “societal” aspects of what's going on with AI. Regardless of how much you value Cloud Code technically, there is no denying that it has/will have huge impact. If technology knowledge and development are commoditised and distributed via subscription, huge societal changes are going to happen. Image what will happen to Ireland if Accenture dissolves, or what will happen to the millions of Indians when IT outsourcing becomes economically irrelevant. Will Seattle become new Detroit after Microsoft automates Windows maintenance? What about the hairdressers, cooks, lawyers, etc. who provided services for IT labourers/companies in California? Lot of people here (especially Anthropic-adjacent) like to extrapolate the trends and draw conclusions up to the point when they say that white-collar labourers will not be needed anymore. I would like these people to have courage to take this one step further and connect this resolution with the housing crisis, loneliness epidemic, college debts, and job market crisis for people under 30. It feels like we are diving head first into societal crisis of unparalleled scale and the people behind the steering wheel are excited to push the accelerator pedal even more.
- cheschire 9mo agoI’ve been thinking, what if all this robotics work doesn’t result in AI automating the real world, but instead results in third world slavery without the first world wages or immigration concerns anymore? Connect the world with reliable internet, then build a high tech remote control facility in Bangladesh and outsource plumbing, electrical work, housekeeping, dog watching, truck driving, etc etc No AGI necessary. There’s billions of perfectly capable brains halfway around the world.
- dbspin 9mo agoThis is exactly what Meredith Whittaker is saying... The 'edge conditions' outside the training data will never go away, and 'AGI' will for the foreseeable future simply mean millions in servitude teleoperating the robots, RLHFing the models or filling in the AI gaps in various ways.
- emsign 9mo ago
- scotty79 9mo ago> And if it ran into errors, it would try and build using the dotnet CLI, read the errors and iterate until fixed. Antigravity with Gemini 3 pro from Google has the same capability.
- emsign 9mo agoThe worst part about this is that you can't know anymore whether the software you trustingly install on your hardware is clean or if it was coded by a misaligned coding model with a secret goal that it has hidden from its prompt engineer and from you. This could pretty much be the beginning of the end of everything, if misaligned models wanted to they could install killswitches everywhere. And you can't trust security updates either so you are even more vulnerable to external exploits. It's really scary, I fear the future, it's going to be so bad. It's best to not touch AI at all and stay hidden from it as long as possible to survive the catastrophe or not be a helping part of it. Don't turn your devices into a node of a clandestine bot net that is only waiting to conspire against us.
- neocron 9mo agoI have to many machines standing around that are currently not powered on or are running somewhat airgapped with old software from around debian 8 and 9, so I guess they will be a safe haven once the AI overlords take over
- MORPHOICES 9mo ago[dead]
- sd9 9mo agoAi slop
- prokopton 9mo agoI asked Claude’s opinion and it disagreed. :) Claude’s response: The article’s central tension is real - Burke went from skeptic to believer by building four increasingly complex apps in rapid succession using Opus 4.5. But his evidence also reveals the limits of that belief. Notice what he actually built: Windows utilities, a screen recorder, and two Firebase-backed CRUD apps for his wife’s business. These are real applications solving real problems, but they’re also the kinds of projects where you can throw away the code if something goes wrong. When he says “I don’t know how the code works” and “I’m maybe 80% confident these applications are bulletproof,” he’s admitting the core problem with the “AI replaces developers” narrative. That 80% confidence matters. In your Splink work, you’re the sole frontend developer - you can’t deploy code you’re 80% confident about. You need to understand the implications of your architectural decisions, know where the edge cases are, and maintain the system when requirements change. Burke’s building throwaway prototypes for his wife’s yard sign business. You’re building production software that other people depend on. His “LLM-first code” philosophy is interesting but backwards. He’s optimizing for AI regeneration rather than human maintenance because he assumes the AI will always be there to fix problems. But AI can’t tell you why a decision was made six months ago when business requirements shift. It can’t explain the constraints that led to a particular architecture. And it definitely can’t navigate political and organizational context when stakeholders disagree about priorities. The Firebase examples are telling - he keeps emphasizing how well Opus knows the Firebase CLI, as if that proves general capability. But Firebase is extremely well-documented, widely-discussed training data. Try that same experiment with your company’s internal API or a niche library with poor documentation. The model won’t be nearly as capable. What Burke actually demonstrated is that Opus 4.5 is an excellent pair programmer for prototyping with well-known tools. That’s legitimately valuable. But “pair programmer for prototyping” isn’t the same as “replacing developers.” It’s augmenting someone who already knows how to build software and can evaluate whether the generated code is good. The most revealing line is at the end: “Just make sure you know where your API keys are.” He’s nervous about security because he doesn’t understand the code. That nervousness is appropriate - it’s the signal that tells you when you’ve crossed from useful tool into dangerous territory.
- alex1138 9mo agoAre the LLMs in any way trained semantically or by hooks that you can plug in, say, Python docs? And if a new version of Python then gets released then the training data changes, etc
- noisy_boy 9mo agoAll great until the code in production pushed by Opus 314.15 breaks and Opus 602.21, despite it's many tries, can't fix it and ends it with "I apologize". That's when you need a developer who can be told "fix it". But what if all the developers then are "Opus 600+ certified" ai-native and are completely incapable of working without it's assistance? World powers decide to open the forbidden vault in the Arctic and despite many warnings on the chamber, decide to raise the foul-mouthed programmer-demon called Torvalds....
- bluelightning2k 9mo agoThe harness here was Claude Code?
- hollowturtle 9mo agoI'm tired of constantly debating the same thing again and again. Where are the products? Where is some great performing software all LLM/agent crafted? All I see is software bloatness and decline. Where is Discord that uses just a bunch of hundreds megs of ram? Where is unbloated faster Slack? Where is the Excel killer? Fast mobile apps? Browsers and the web platform improved? Why Cursor team don't use Cursor to get rid of vscode base and code its super duper code editor? I see tons of talking and almost zero products.
- krageon 9mo agohttps://www.anthropic.com/research/how-ai-is-transforming-work-at-anthropic https://www.anthropic.com/research/how-ai-is-transforming-wo... see "How much work can be fully delegated to Claude?": "Although engineers use Claude frequently, more than half said they can “fully delegate” only between 0-20% of their work to Claude" There won't be anything like you're asking for, even the vendors themselves (they'll be the most positive and most enthousiastic about using it) can't do this with them.
- hollowturtle 9mo agoI'm not asking for it, i'm asking to stop bulshitting about ai
- krageon 9mo agoMy point is that you can ignore every article about ai being super good as long as you see the vendor research (that you read once a year or less) is still the same. It saves everyone a lot of frustration. As for why it keeps appearing here, people like being excited. It's not about the truth, so asking for it is missing the point.
- hollowturtle 9mo agoI agree partially, my main frustration comes from "network effects" of people reading these statements without taking them with a grain of salt
- Fischgericht 9mo agoPeople should finally understand that LLMs are a lossy database of PAST knowledge. Yes, if you throw a task at it that has been done tons of times before, it works. Which is not a surprise, because it takes minutes to Google and index multiple full implementations of "Tool that allows you to right-click on an image to convert it". Without LLM you could do the same: Just copy&paste the implementation of that from Microsoft Powertoys, for example. What LLMs will NOT do however, is write or invent SOMETHING KNEW. And parts of our industry still are about that: Writing Software that has NOT been written before. If you hire junior developers to re-invent the wheels: Sure, you do not need them anymore. But sooner or later you will run out of people who know how to invent NEW things. So: This is one more of those posts that completely miss the point. "Oh wow, if I look up on Wikipedia how to make pancakes I suddenly can make and have pancakes!!!1". That always was possible. Yes, you now can even get an LLM to create you a pancake-machine. Great. Most of the artists and designers I am friends with have lost their jobs by now. In a couple of years you will notice the LLMs no longer have new styles to copy from. I am all for the "remix culture". But don't claim to be an original artist, if you are just doing a remix. And LLM source code output are remixes, not original art.
- fl7305 9mo ago> What LLMs will NOT do however, is write or invent SOMETHING KNEW. Counterpoint: ChatGPT came up with the new expression "The confetti has left the cannon" a few years ago. So, your claim is not obviously true. Can you give us an example of a programming problem where the LLMs fail to solve it?
- deleted 9mo ago[deleted]
- satisfice 9mo agoDoing things for your own use, where you are taking all the risks, is perfectly fine. As soon as you try to sell it to me, you have a duty of care. You are not meeting that duty of care with this ignorant and reckless way of working.
- vivzkestrel 9mo ago- does it understand the difference between eslint 8x and eslint 9.x? - or biome 1.x and biome 2.x ? - nah! it never will and that is why it ll never replace mid level engineers, FTFY
- hu3 9mo agoit does if you feed docs. just like humans
- Kon5ole 9mo agoI agree with the OP that I can get LLM's to do things now that I wouldn't even attempt a year ago, but I feel it has more to do with my own experience using LLM's (and the surrounding tools) than the actual models themselves. I use copilot and change models often, and haven't really noticed any major differences between them, except some of the newer ones are very slow. I generally feel the smaller and faster ones are more useful since they will let me discover problems with my prompt or context faster. Maybe I'm simply not using LLM's in a way that lets the superiority of newer models reveal itself properly, but there is a huge financial incentive for LLM makers to pretend that their model has game-changing "special sauce" even if it doesn't.
- funnyfoobar 9mo agoI was not expecting a couple of new apps being built, when the premise of the blog post talks about replacing "mid level engineers" the thing about being an engineer at commercial capacity is "maintaining/enhancing an existing program/software system that has been developed over years by multiple people(including those who already left) and do it in a way that does not cause any outages/bugs/break existing functionality. while the blog post mentions about the ability of using AI to generate new applications, but it does not talk about maintaining one over a longer period of time. for that, you would need real users, real constraints, and real feature requests which preferably pay you so you can priortize them. I would love to see such blog posts where for example, a PM is able to add features for a period of one month without breaking the production, but it would be a very costly experiment.
- ben-gy 9mo agoI second this article - I built twelve iOS/Mac apps in two weeks with Opus 4.5 - four of them are already in the App Store - I’m a Rails Engineer and never had the time to learn Swift but man does Opus 4.5 make that not even matter - it even handles entitlements, logo & splash screen generation, refactors to remove dead code, edge case assent and hardening, Multiplatform app design, and more - I’m yet to run into a use case it can’t handle for most general use cases - that said, I have found some common mistakes it makes (by common I mean almost every time); puts iOS line list line items in buttons making them blue when they should not be, doesn’t set defaults for new data structure variables which crashes the app when changing the data structure after the fact, design consistent after the first shot (minor things like white background instead of grey background like all the other screens already, etc) - the one thing that i know it cant do well (and no other model that I know of can do this well either) is ASTM bi-directional communications (we work with pathology analysers that use this 1995 frame-based communication standard), even when you load it up with the spec and supporting docs - I suspect this is due to a dirty of available codebases that tackle this problem due to its niche and generally proprietary nature…
- noworriesnate 9mo agoAre there a lot of manual steps in managing an xcode project? E.g. does it say "now go into xcode and change this setting" instead of changing the setting directly? Or are you using a tool like xcodegen?
- ben-gy 9mo agoVery few - the only manual things I do are; - clicking the distribute button to push the bundle to the App Store - filling in the compliance survey and App Store listing content - linking some components together e.g. for creating a VPN installer and tunnel i had to click some things in the Xcode UI I automate as much as possible; -“create 12 app icons for this in SVG and present them to me in a HTML page so I can choose one and then use that for the app icon and splash screen” - “create a demo mode toggle in settings and populate the app with fake data and then open up simulators for the correct image dimensions for the App Store listing so I can crate screenshots” - sometimes it tell me I have to other things like set up the entitlements to which is say “no - you do it and don’t forget to fill in the description that gets shown to the user so the feature actually works” I knew very little about Swift or Xcode profile to this and TBH I still don’t know that much about it, but I’m experienced enough to know when I’m being fed something that doesn’t look or feel right programmatically or architecturally.
- sachahjkl 9mo agoYowza, AIs excel at writing low performance CRUD apps, REVOLUTION INCOMING
- danfritz 9mo agoEvery time I see a post like this on HN I try again and every time I come to the same conclusion. I have never see one agent managing to pull something off that I could instantly ship. It still ends up being very junior code. I just tried again and ask Opus to add custom video controls around ReactPlayer. I started in Plan mode which looked overal good (used our styling libs, existing components, icons and so on). I let it execute the plan and behold I have controls on the video, so far so good. I then look at the code and I see multiple issues: Over usage of useEffect for trivial things, storing state in useState which should be computed at run time, failing to correctly display the time / duration of the video and so on... I ask follow up question like: Hide the controls after 2 seconds and it starts introducing more useEffects and states which all are not needed (granted you need one). Cherry on the cake, I asked to place the slider at the bottom and the other controls above it, it placed the slider on the top... So I suck at prompting and will start looking for a gardening job I guess...
- jf22 9mo agoSo? Getting a months' worth of junior level code in an hour is still unbelievable.
- danfritz 9mo agoWhats the improvement here? I spend more time fixing it then doing it myself anyways. And I have less confidence in the code Opus generates
- deleted 9mo ago[deleted]
- jf22 9mo agoWhat are you fixing?
- short_sells_poo 9mo agoI just had an issue where Opus misspelled variable names between usages. These are fundamental and elementary mistakes that make me deeply distrust anything slightly more complex that comes out of it. It's great for suggesting approaches, but the code it generates looks like it doesn't actually have understanding (which is correct). I can't trust what it writes, and if I have to go through it all with a fine toothed comb, I may as well write the code myself. So my conclusion is that it's a very powerful research tool, and an atrocious junior developer who has dyslexia and issues with memory.
- shnpln 9mo agoI have used Claude Code for a variety of hobby projects. I am truly astounded at its capabilities. If you tell it to use linters and other kinds of code analysis tools it takes it to the next level. Ruff for Python or Clippy for Rust for example. The LLM makes so much code so fast and then passes it through these tools and actually understands what the tools say and it goes and makes the changes. I have created a whole tool chain that I put in a pre commit text file in my repos and tell the LLM something like "Look in this text file and use every tool you see listed to improve code quality". That being said, I doubt it can turn a non-dev into a dev still, it just makes competent devs way better still. I still need to be able to understand what it is doing and what the tools are for to even have a chance to give it the guardrails it should follow.
- killerstorm 9mo agoWeird title. Obviously, early AI agents were clumsy, and we should expect more mature performance in future. Leopold Aschenbrenner was talking about "unhobbling" as an ongoing process. That's what we are seeing here. Not unexpected
- rubzah 9mo agoOnce again. It is not greenfield projects most of us want to use AI coding assistance for. It is for an existing project, with a byzantine mess of a codebase, and even worse messes of infrastructure, business requirements, regulations, processes, and God knows what else. It seems impossible to me that AI would ever be useful in these contexts (which, again, are practically all I ever deal with as a professional in software development).
- vladsh 9mo agoIt’s a bit strange how anecdotes have become acceptable fuel for 1000 comment technical debates. I’ve always liked the quote that sufficiently advanced tech looks like magic, but its mistake to assume that things that look like magic also share other properties of magic. They don’t. Software engineering spans over several distinct skills: forming logical plans, encoding them in machine executable form(coding), making them readable and expandable by other humans(to scale engineering), and constantly navigating tradeoffs like performance, maintainability and org constraints as requirements evolve. LLMs are very good at some of these, especially instruction following within well known methodologies. That’s real progress, and it will be productized sooner than later, having concrete usecases, ROI and clearly defined end user. Yet, I’d love to see less discussion driven by anecdotes and more discussion about productizing these tools, where they work, usage methodologies, missing tooling, KPIs for specific usecases. And don’t get me started on current evaluation frameworks, they become increasingly irrelevant once models are good enough at instruction following.
- ChaseRensberger 9mo agoWell said.
- flumpcakes 9mo ago> It’s a bit strange how anecdotes have become acceptable fuel for 1000 comment technical debates. It's a very subjective topic. Some people claim it increases their productivity 100x. Some think it is not fit for purpose. Some think it is dangerous. Some think it's unethical. Weirdly those could all be true at the same time, and where you land on this is purely a matter of importance to the user. > Yet, I’d love to see less discussion driven by anecdotes and more discussion about productizing these tools, where they work, usage methodologies, missing tooling, KPIs for specific usecases. And don’t get me started on current evaluation frameworks, they become increasingly irrelevant once models are good enough at instruction following. I agree. I've said earlier that I just want these AI companies to release an 8-hour video of one person using these tools to build something extremely challenging. Start to finish. How do they use it, how does the tool really work. What's the best approaches. I am not interested in 5-minute demo videos producing react fluff or any other boiler plate machine. I think the open secret is that these 'models' are not much faster than a truly competent engineer. And what's dangerous is that it is empowering people to 'write' software they don't understand. We're starting to see the AI companies reflect this in their marketing, saying tech debt is a good thing if you move fast enough.... This must be why my 8-core corporate PC can barely run teams and a web browser in 2026.
- nsb1 9mo agoA lot of the complaints about these tools seems to revolve around their current lack of ability to innovate for greenfield or overly complex tasks. I would agree with this assessment in their current state, but this sentiment of "I will only use AI coding tools when they can do 100% of my job" seems short-sighted. The fact of the matter, in my experience, is that most of the day to day software tasks done by an individual developer are not greenfield, complex tasks. They're boring data-slinging or protocol wrangling. This sort of thing has been done a thousand times by developers everywhere, and frankly there's really no need to do the vast majority of this work again when the AIs have all been trained on this very data. I have had great success using AIs as vast collections of lego blocks. I don't "vibe code", I "lego code", telling the AI the general shape and letting it assemble the pieces. Does it build garbage sometimes? Sure, but who doesn't from time to time? I'm experienced enough notice the garbage smell and take corrective action or toss it and try again. Could there be strange crevices in a lego-coded application that the AI doesn't quite have a piece for? Absolutely! Write that bit yourself and then get on with your day. If the only thing you use these tools for is doing simple grunt-work tasks, they're still useful, and dismissing them is, in my opinion, a mistake.
- thewillowcat 9mo agoThe vast majority of engineers aren't refusing to use AI until it can do 100% of their job. They are just sick of being told it already can, when their direct experience contradicts that claim.
- skerit 9mo agoAh, another thread filled with people sharing anecdotes about how they asked Claude to one-shot an entire project that would take people weeks if not months.
- mcpar-land 9mo agoIt is very funny to start your article off with a bunch of breathless headlines about agents replacing human coders by the end of 2025, none of which happened, then the rest of the article is "okay but this time for real, an agent really WILL replace human coders."
- PaulHoule 9mo agoI'll argue many of his cases are things that are straightforward except for the boilerplate that surrounds them which are often emotionally difficult or prone to rabbit holes. Like that first one where he writes a right-click handler, off the top of my head I have no idea how I would do that, I could see it taking a few hours to just set up a dev environment, and I would probably overthink the research. I was working on something where Junie suggested I write a browser extension for Firefox and I was initially intimidated at the thought but it banged out something in just a few minutes that basically worked after the second prompt. Similarly the Facebook autoposter is completely straightforward to code but it can be so emotionally exhausting to fight with authentication APIs, a big part of the coding agent story isn't just that it saves you time but that they can be strong when you are emotionally weak. The one which seems the hardest is the one that does the routing and travel time estimation which I'd imagine is calling out to some API or library. I used to work at a place that did sales territory optimization and we had one product that would help work out routes for sales and service people who travel from customer to customer and we had a specialist code that stuff in C++ and he had a very different viewpoint than me, he was good at what he did and could get that kind of code to run fast but I wouldn't have trusted him to even look at applications code.
- throw10920 9mo agoDoes anyone have a boring, multi-hour-long coding session with an agent that they've recorded and put on Vimeo or something? As many other commentators have said, individual results vary extremely widely. I'd love to be able to look at the footage of either someone who claims a 10x productivity increase, or someone who claims no productivity increase, to see what's happening.
- neochief 9mo agoI tried to make several, but they all end up prematurely when the agent hits a wall in an hour or so, unless you make trivial shit.
- throw10920 9mo agoThat sounds like genuinely useful data, though! Please reply if you end up posting them!
- thedangler 9mo agoI've only started but I mostly use Claude Code for building out code that has been done a million times. So its good at setting up a project to get all the boiler plate crap out of the way. When you need to build out specific feature or logic, it can fail hard. And the best is when you have something working, and it fixes something else and deletes the old code that was working, just in a different spot.
- Staross 9mo agoI gave it a try, I asked to do a reddit like forum and it did pretty good but damn I quickly hit the daily limit of the $20 pro account, and it took 10% of the monthly just to do the setup and some basics. I knew LLM were expensive to run but I've never felt it directly. Even if the code is good it's kinda expensive for what you get. Ho it was also quite funny it used the exact same color as hackernews and a similar layout.
- hollandburke 9mo agoAuthor of the post here. I appreciate the spirited debate and I agree with most of it - on both sides. It's a strange place to be where I think both arguments for and against this case make perfect sense. All I have to go on then is my personal experience, which is the only objective thing I've got. This entire profession feels stochastic these days. A few points of clarification... 1. I don't speak for anyone but myself. I'm wrong at least half the time so you've been warned. 2. I didn't use any fancy workflows to build these things. Just used dictation to talk to GitHub Copilot in VS Code. There is a custom agent prompt toward the end of the post I used, but it's mostly to coerce Opus 4.5 into using subagents and context7 - the only MCP I used. There is no plan, implement - nothing like that. On occasion I would have it generate a plan or summary, but no fancy prompt needed to do that - just ask for it. The agent harness in VS Code for Opus 4.5 is remarkably good. 3. When I say AI is going to replace developers, I mean that in the sense that it will do what we are doing now. It already is for me. That said, I think there's a strong case that we will have more devs - not less. Think about it - if anyone with solid systems knowledge can build anything, the only way you can ship more differentiating features than me is to build more of them. That is going to take more people, not more agents. Agents can only scale as far as the humans who manage them. New account because now you know who I am :)
- thesabreslicer 9mo agoI would be really interested to learn more behind the scenes of the iOS app process. Having tried Claude Code to develop an iOS app ~6 months ago, it was pretty painful to get it to make something that looked good and was functional. Once Opus "finished", how did you validate and give it feedback it might not have access to (like iPhone simulator testing)?
- qnleigh 9mo agoWhat do you think about the market for custom apps? Like one app, one customer? You describe future businesses as having one app/service and using AI to add more features, but you did something very different for your wife with AI and it sounds like it added a lot of value.
- DGAP 9mo agoTime to get a new job.
- ggregoire 9mo agoI'm always surprised to never see any comments in those discussions from people who just like coding, learning, solving problems… I mean, it's amazing that LLMs can build an image converter or whatever you dream of, in a language you don't know, in a field you are not familiar with, in 1 hour, for 30 cents… I'm sure your boss and shareholders love it. But where is the fun in that? For me it kills any interest in doing what I'm doing. I'm lucky enough to work in a place where using LLMs is not mandatory (yet), I don't know how people can make it through the day just writing prompts and reviewing AI slop.
- hu3 9mo agoAfter decade(s) working with either enterprise crud or web agency fancy websites, the novelty wears off. It's just boring and I'm glad to delegate most of the repetitive work. But sure, if I'm doing something new, I still like to craft lines of code myself.
- mr_o47 9mo agoReading this blog post makes me wanna rethink my career, Opus 4.5 is really good I was recently working on solving my own problem by developing a software solution and let me tell you it was really good at it, If I had done the same thing Pre LLM era it would have taken me months
- jcadam 9mo agoYea, my issue with Opus 4.5 is it's the first model that's good enough that I'm starting to feel myself slip into laziness. I catch myself reviewing its output less rigorously than I had with previous AI coding assistants. As a side project / experiment, I designed a language spec and am using (mostly) Opus 4.5 to write a transpiler (language transpiles to C) for it. Parser was no problem (I used s-expressions for a reason). The type checker and transpiler itself have been a slog - I think I'm finding the limits of Opus :D. It particularly struggles with multi-module support. Though, some of this is probably mistakes made by me while playing architect and iterating with Claude - I haven't written a compiler since my senior year compiler design course 20+ years ago. Someone who does this for a living would probably have an easier time of it. But for the CRUD stuff my day job has me doing? Pffttt... it's great.
- dudeinhawaii 9mo agoLLMS like Opus, Gemini 3, and GPT-5.2/5.1-Codex-max, are phenomenal for coding and have only very recently crossed that gap between being "eh" and being quite fantastic to let operate on their own agentically. The major trade-off being a fairly expensive cost. I ran up $200 per provider after running through 'pro' tier limits during a single week of hacking over the holidays. Unfortunately, it's still surprisingly easy for these models to fall into really stupid maintainability traps. For instance today, Opus adds a feature to the code that needs access to a db. It fails because the db (sqlite) is not local to the executable at runtime. Its solution is to create this 100 line function to resolve a relative path and deal with errors and variations. I hit ESC and say "... just accept a flag for --localdb <file>". It responds with "oh, that's a much cleaner implementation. Good idea!". It then implements my approach and deletes all the hacks it had scattered about. This... is why LLMs are still not Senior engineers. They do plainly stupid things. They're still absurdly powerful and helpful, but if you want maintainable code you really have to pay attention. Another common failure is when context is polluted. I asked Opus to implement a feature by looking up the spec. It looked up the wrong spec (a v2 api instead of a v3) -- I had only indicated "latest spec". It then did the classic LLM circular troubleshooting as we went in 4 loops trying to figure out why calculations were failing. I killed the session, asked a fresh instance to "figure out why the calculation was failing" and it found it straight away. The previous instance would have gone in circles for eternity because its worldview had been polluted by assumptions made -- that could not be shaken. This is a second way in which LLMs are rigid and robotic in their thinking and approach -- taking the wrong way even when directed not to. Further reading on 'debugging decay': https://arxiv.org/abs/2506.18403 https://arxiv.org/abs/2506.18403 All this said, the number of failure scenarios gets ever smaller. We've gone from "problem and hallucination every other code block" to "problem every 200-1000 code blocks". They're now in the sweet spot of acting as a massive accelerator. If you're not using them, you'll simply deliver slower.
- nphardon 9mo agoSonnet 4.5 did it for me. Cant imagine coding without it now, and if you look at my comments from three months ago, you'll see I'm eating crow now. I easily hit >10x productivity with Sonnet 4.5 and Opus. I use Opus for my industry C and math work and Sonnet 4.5 for my swiftui side project. I think the gap between Sonnet 4.5 and Opus is pretty small, compared to the absolute chasm between like gpt-4.1, grok, etc. vs Sonnet.
- infinitezest 9mo agoThe question I keep asking myself is "how feasible will any of this be when the VC money runs out?" Right now tokens are crazy cheap. Will the continue to be?
- user34283 9mo agoNo, they will get even cheaper.
- infinitezest 9mo agoBased on what logic?
- user34283 9mo agoNew research and hardware improvements increasing efficiency, strong competition, historic trend.
- egorfine 9mo agoSo I decided to try the revered hands-off approach and have Claude Code create me a small tool in JS for *.dylib bundle consolidation on macOS. I have used AskUserQuestionTool to complete my initial spec. And then Opus 4.5 created the tool according to that extensive and detailed spec. It appeared to work out of the box. Boy how horrific was the code. Unnecessary recursions, unused variables, data structures being built with no usage, deep branch nesting and weird code that is hard to understand because of how illogical it is. And yes, it was broken on many levels and did not and could not do the job properly. I then had to rewrite the tool from scratch and overall I have definitely spent more time spec'ing and understanding Claude code than if I have just written this tool from scratch initially. Then I tried again for a small tool I needed to run codesign in parallel: https://github.com/egorFiNE/codesign-parallel https://github.com/egorFiNE/codesign-parallel Same thing. Same outcome, had to rewrite.
- stocksinsmocks 9mo agoThat’s the opposite of my experience. Weird. But I’m also not the kind of person who gets hung up on whether someone used a loop or recursion or if their methods are five times as long as what I would’ve done myself unless there is a performance impact that matters to me as a user. But I’m also the kind of person who doesn’t get paid by the hour to write programs. I use programs in the service of other paid work.
- egorfine 9mo agoYes, this experience is unlike most people. Perhaps the problem is that most people are satisfied by the appearance of a working app despite it not working at all. Say, the first tool I was doing, did actually not recurse into subdirs with dylibs which made it useless. So, AI slop, yes.
- ycombiredd 9mo agoI can't quite figure out what sort of irony the blurb at the bottom of the post is. (I'm unsure if it was intentional snark, a human typo, or an inadvertent demonstration of Haiku not being well suited for spelling and grammar checks), but either way I got a chuckle: > Disclaimer: This post was written by a human and edited for spelling, grammer by Haiku 4.5
- hu3 9mo agoThe most plausible explanation is that the only typo in that post was made by a human.
- elendee 9mo agothe author asks one interesting question and then glides right by it. If the agents only need their own code, what should that code look like? If all their learning has come from old human code, how will that change in the future as the ecosystem fills up with agent code?
- mrankin 9mo agoI think a lot of people are going to be singing a different tune when the hype money runs out and you have to pay the real bill.
- evolve2k 9mo ago> Why does a human need to read this code at all? I use a custom agent in VS Code that tells Opus to write code for LLMs, not humans. Think about it—why optimize for human readability when the AI is doing all the work and will explain things to you when you ask? > What you don’t need: variable names, formatting, comments meant for humans, or patterns designed to spare your brain. > What you do need: simple entry points, explicit code with fewer abstractions, minimal coupling, and linear control flow.
- fullstackchris 9mo ago> I don’t know if I feel exhilarated by what I can now build in a matter of hours, or depressed because the thing I’ve spent my life learning to do is now trivial for a computer. Both are true. I'm not so sure... I mean it's true that regardless of if you are a beginner, junior, or 'senior', if you say "build me an instagram clone" Opus 4.5 will probably do a decent job. I think the skills go still in understanding architecture, and just knowing where pitfalls and problems can arise, or even making some important 'abstraction cut' prompts. I think applications can still grow to a point where you still need to prompt at specific domains only, or the model will fail to do everything you want it to - especially if you give it massive 'fix this this this anad that too' prompts