8 ms·
I was worried this time last year that by this time this year, companies would have slashed their engineering teams down to a handful and everything would be dr
by efficax 3mo ago
I was worried this time last year that by this time this year, companies would have slashed their engineering teams down to a handful and everything would be driven by mostly autonomous agents with human guidance. But it just hasn't happened. Do I write all my code with an agent now? Yes. Can you just give an agent a desired outcome and let it work, unsupervised? Absolutely not. I can produce more code than I used to, but if I want it to be good, to be stable, to do what the product manager and designers want, it's only about 2 to 3 times more code than before. And that productivity is impacted by the fact that I'm reviewing 2 to 3 times more code than before (and you have to review, even more so now than before, because if you just let opus or gpt 5 do its thing, you'll get some terrible results, and I've found a lot of engineers on my team are just letting it do it's thing without a lot of iteration).
- jknoepfler 3mo agoIck. Stop.
- alt227 3mo agoI have experienced and feel very much the same, and it is refreshing to see a realistic post about the success of agentic coding instead of the usual hype or doom.
- ramoz 3mo agoAs crazy as it may sound, my workflow today does not look too different from a year ago - where I was already heavy into claude code. Im not certain things will look too different a year from now either. We still have serious bottlenecks in terms of focus/attention you have for both delegating agent work and being able to review it. Even if we solve the "trust what ai does" problem, these cognitive deficit issues still exist - for teams coordinating work, even users adopting new shit, etc. As an industry we are leaning heavy into accepting "slop" as the status quo - we care more about efficiency of output right now. Slop will get better & we can become more adaptive to living with the paradox of amazing yet delicate systems generated by AI. But I feel big shifts coming in this regard and if/when it does we may find ourselves in the dystopia of broader unemployment with worse net outcomes. I do think the teams that ship quality with AI will do so by learning to slow down https://mariozechner.at/posts/2026-03-25-thoughts-on-slowing-the-fuck-down/ https://mariozechner.at/posts/2026-03-25-thoughts-on-slowing...
- zamalek 3mo ago> Can you just give an agent a desired outcome and let it work, unsupervised? Absolutely not. Ignoring instructions - whether in AGENTS.md or my prompt - is the worst of it, and it routinely happens. It just waives things that I explicitly told it to do as part of the design. Vibe coders (in the true sense, zero oversight) claim that you just need to prompt it carefully. That's completely untrue when faced with your careful prompt being ignored. I even have "don't overrule me without asking" in my global AGENTS.md, and it simply doesn't do that.
- codemog 3mo agoAlso noticed this. Their intelligence is very jagged. I’ve had them produce some highly optimized code yet fail to follow basic code guidelines.
- rogerrogerr 3mo agoI’m convinced the magic bullet is deterministic checks. Linters, static analyzers, etc. Whatever you can do to create deterministic gates that the LLM simply must overcome to reach a “done” state, do it. Has been making a huge difference for my team, but sister teams are so invested in writing the perfect Make No Mistakes prompt that they just can’t see it. Basically I treat it like a junior dev. We don’t get junior devs to write code correctly by cajoling them just right, we add CI gates. It still works.
- sdesol 3mo agoWhy aren't the teams using shared checks? Are the codes in different repos?
- rogerrogerr 3mo agoThey’re very, very different projects.
- zamalek 3mo agoWouldn't have helped, sibling comment: https://news.ycombinator.com/item?id=48797883 https://news.ycombinator.com/item?id=48797883 Architectural decisions are not lintable.
- supern0va 3mo ago>I was worried this time last year that by this time this year, companies would have slashed their engineering teams down to a handful and everything would be driven by mostly autonomous agents with human guidance. But it just hasn't happened. I find this somewhat puzzling. I thought things were moving quickly, but at this time last year I couldn't even get Claude (using Cursor) to spin me up a service skeleton that would compile, let alone do anything meaningful. I know it feels like a long time somehow, but it was only between November and February that things started to actually somewhat work without significant hand holding. Even now, it seems like we're still figuring out how to fully leverage the current models and tooling, even in organizations that have largely gotten on board.
- deaux 3mo ago> at this time last year I couldn't even get Claude (using Cursor) to spin me up a service skeleton that would compile, let alone do anything meaningful I've been using it to do this for 2 years now. And many people with me. The change you mention is one of is primarily one of Overton windows, of vibes.
- simonw 3mo agoWhich harness software were you using for this 2 years ago? VS Code Copilot? Cursor?
- WhatIsDukkha 3mo agoAider for me Very successful by just being careful and walking it forward. Yes its about 2 years, August 2024 from git it looks like.
- deaux 3mo agoExactly the same for me. Sonnet 3.5+Aider made it possible. I got downvoted though, apparently you and me are liars, haha. Easier to believe that than admit that their take of "I adopted it exactly when it became viable, now it's good, and before that it was a waste of time" was wrong. They're the people who were loudly claiming here last summer it was useless, asking "show me anything useful that has been coded using LLMs". Come to think of it, it has now been a few months since I saw one of those. Used to be every other thread.
- tunesmith 3mo agoI think it doesn't prove much that it hasn't happened yet. Companies might just be moving slower than you think, and are still planning on doing it. And, in many corners, "don't manually write code" is being joined by "don't manually read code" as an attractive principle.
- goatlover 3mo agoThe Yale economist Pascual Restrepo, who is well regarded researcher in automation, doesn't think it will happen for most jobs. https://fortune.com/2026/04/04/ai-jobs-future-not-important-enough-to-be-automated-yale/ https://fortune.com/2026/04/04/ai-jobs-future-not-important-...
- Imustaskforhelp 3mo agoyou might have to think the way through though and these companies are already being caught up with the huge token costs at the same time. There was an interesting comment during the cloudflare layoffs (partially driven by the fact that the company was bleeding money also because of its token costs from one estimate being 5* million$ per month (I feel so silly that I accidentally had written/meant 500 and had kentonv do the stats on that part :-( Sorry kentonv!), don't quote me on that though) The part was that there is only an enough marketshare in the first place. Cloudflare was doing some crazy experiments like operating matrix on cf workers and wordpress alternative and fediverse and so much stuff. So they basically spent 10x the amount of token (and the token costs) and I imagine as such the reading code of that part was getting sidelined as the attractive principle you are talking about. Yet the market can't bring an actual demand 10x times though. These are things which nudge a user slightly but the actual impact on user growth isn't 10x or even justifiable within some cases given the costs. Yet at the same time driving up the people who actually know their stuff and firing them because of the token costs. The people who have actually mitigated some of the largest DDOS attacks and are the backbone behind cf cash-cow (enterprise payments) is the fact that they have had the experience and entreprise knowledge about these things, yet they are literally removing that by firing workers and oh replacing them with interns. (They got 1111 interns and fired 1100 employees or something iirc) It's weird and I have talked to some people about it but there is a disconnect between what management is hearing about AI and the ground reality of things. Reviewing code is becoming the bottleneck but if you don't review code and are shipping things to production, then you can get fired as I have talked about in some of my other comments sharing a story about how a guy shipped code to prod and the response was "but claude generated it" and got fired because the company basically said, look we basically don't care if it was generated by claude but the responsibility was on you to check it (review) and because the commit was done by you, you are gonna be treated responsible and he got fired from his job. Yet this was the same company which was asking its employee to play around with claude at their free time, the manager of the employee I talked to being the most automatable person, the company employees working till 1 AM because they were saying to management that things were fine but they were being burried under the technical debt,that employee that I talked to got honest with the management and told reality and the management treated them as a person who didn't know AI or were the odd one out. Sooo I don't know actually to be honest. TLDR: reviewing code is being treated as the bottleneck but it is also the only thing stopping your company from imploding under technical debt, actual debt because of token costs etc. I remain skeptical if we should treat it as a bottleneck or as a safeguard mechanism. After all, if nobody's in the loop then whose responsible? Reviewing code isn't a bottleneck so much so its a safeguard mechanism in my opinion. Also things differ in corporate land and hobby land and I would prefer corporate to not be using the practices that I do with how I do things for fun in my hobby time. Side note: Even more so, I think I am a LiteLLM security working group maintainer and I have seen first hand on how much damage it can do in supply chain even when things were done right from LiteLLM side and the fault was within the side of ironically a security product that they used called Trivy. There are things which you can do to be better prone to supply chain attacks in general but there is no full bullet proof way of doing so and in such. Caution (should) be taken when dealing with corporate systems and as such I sweat a little when anyone suggests code review to be completely eliminated. Things (are/can be) different in hobby/prototyping world though.
- khurs 3mo ago"companies would have slashed their engineering teams down to a handful and everything would be driven by mostly autonomous agents with human guidance. But it just hasn't happened. " It never was going to happen. Always the same story: https://en.wikipedia.org/wiki/Gartner_hype_cycle#/media/File:Gartner_Hype_Cycle.svg https://en.wikipedia.org/wiki/Gartner_hype_cycle#/media/File...
- kaydub 3mo agoCompanies are putting a ton of effort into getting to that point of having agents do the work unsupervised. Whoever gets there first is going to be the winner. I personally don't think it's possible and I haven't written a line of code since Sept 2025. There's an AI psychosis going on right now, especially among the execs or management class, and we all gotta nod our heads in agreement and burn through tokens.
- eloisius 3mo agoIf you have runway, it’s a good time to start your own thing or join someone who is. Personally, I cannot force myself to wade through slop PRs from careless coworkers. If that’s the job now, I’d rather run a hotdog stand or something. Luckily, I don’t think things are that dire. I think the companies issuing AI mandates are manufacturing sawdust, and even if it works, it would just enable them to burn through customer goodwill in record time as they make user-hostile decisions free from engineer pushback. These are going to be a few tough years, but I think the opportunities to start something new are everywhere.
- dfedbeef 3mo agoPretty sure it's going to be a tough couple of decades, not just years.
- eloisius 3mo agoOf course I don't know either way. However, I feel like it's going to be different, but maybe not apocalyptic. The market I am building for is photographers, and from what I can gather (and know first hand as a customer) is that there's real discontent with the toolmakers in our ecosystem. It seems like in the quest for every tech company to become a multi-billion dollar empire, they've lost the plot on making hammers for their customers and have instead turned into some kind of strip mining operation. A totally AI agent-driven company is a MBA wet dream, and I think a fever dream. If Adobe, for example, were to achieve it, I don't think they'd use it to fix the backlog of bugs overnight. I believe they'd just become an even more incoherent zombie, trying to extract rents from creative cloud subscriptions. In the meantime, photographers still need tools. If you wanna run a software company as a regular small-to-medium-sized business, you may find some customers that are happy to buy a quality hammer. The unicorn startup days might be behind us, but I'd be okay with that. Now, if AI obviates creatives altogether I don't really know what to say. I'd morn the loss of a world that became so tasteless that AI-generated decorations are good enough, for starters.
- hysan 3mo agoI feel like the increased reviewing time is consistently understated. I’m just an IC, but it seems obvious to me that you cannot cut staff and achieve increased output. There literally aren’t enough eyeballs to go around reviewing code when everyone is 2-3xing their output. I spend so much more time reviewing code; reviews that are sorely needed because I regularly catch batshit insane “fixes” that work but would quickly turn the codebase into a mess (the most recent one being a multi-hundred line diff that I went and fixed in 2 lines in 15 min). Maybe I’m underthinking it but it seems obvious that you either maintain the same output with fewer staff or you gain increased output with the same staff. All the companies that are attempting to cut staff and gain increased output are chasing an impossibility and throwing away their opportunity to accelerate.
- sealWithIt 3mo agoI can similarly output 2-3x more code but everything stalls down to me having to review and integrate in a meaningful way the moment I am the one that has to maintain that code. It's eerie to observe collaborators output code they don't understand, spend days chatting with Claude instead of reading (like really reading) compiler's output or 3 pages of manual, and how lost and oblivious they look when the AI fixates on solving a different problem than the one they have been tasked.
- bigstrat2003 3mo agoFrankly: if you want it to be good and stable, you can't really go any faster than before. The time it takes you to review all the code is no less than it would've to just write it in the first place, because the actual typing things out was never the part which took up time.
- threethirtytwo 3mo agoIm not worried about anything within 1 or 2 years. The upheaval if it happens will likely be within 10.
- thisisit 3mo ago> I was worried this time last year that by this time this year, companies would have slashed their engineering teams down to a handful and everything would be driven by mostly autonomous agents with human guidance. But it just hasn't happened. Amara’s law: We tend to overestimate the effect of a technology in the short run and underestimate the effect in the long run. This continues to be applied to AI where people think is going to be next 12 to 18 months. Changes are coming but certainly not at the rate Zuckerberg and most people are thinking.
- ifwinterco 3mo agoI believe this in general but the issue with AI specifically is historically in every cycle we've overestimated the potential of AI in both the short and the long run - people in the 60s were genuinely quite convinced we were close to AGI for a few years, as ridiculous as that now sounds. This is why you get "AI winters" but we've never had a "steam engine winter" or a "railway winter" or a "petrochemicals winter"
- Gareth321 3mo agoIt's interesting to score our predictions of various technologies. I've been waiting for hoverboards for 40 years. Our predictions of nuclear power being in every consumer appliance were way off. We just couldn't make it safe and cheap enough. This is a common theme for "physical" technologies. On the other hand, we greatly underestimated the advent and potential of the internet. Most science fiction of the 20th century envisioned very large computers with limited interconnected capabilities. We made computers far smarter and more ubiquitous than most could ever have conceived. I see AI as a function of the kinds of technologies we consistently underestimate.
- ifwinterco 3mo agoYes, that's a fair counterpoint. I guess my counter-counterpoint would be that LLMs actually seem to have characteristics closer to the first group than the second, in that they (currently at least) need enormous quantities of physical things and energy in order to work. So the nuclear power problem of "it works fine and it's actually very good in a lot of ways, it's just too expensive" could be quite relevant
- bwhiting2356 3mo agoCould be lack of imagination on my part but I truly can't imaging shipping 1000's of lines of code that I can't understand (beyond low-stakes prototypes). That means there's a ceiling on productivity gains.
- Aperocky 3mo agoIf you end up with 2 to 3 times more code. That is HORRIBLE, because it means about 50-66% of the code is otherwise unnecessary. Those are eventually going to become unmaintainable garbage. However, if you get 2 to 3 times the code in the interim, that's probably less than what's needed. I find myself cycle through almost 10x-20x amount of code implementations to get what I want which is actually less code, simple solution and desired behavior. Given a specific behavior, there are usually just 1 simplest implementation, whether done by human or AI. However, there are 100 ways to do it with more complexity and either handwritten or AI slop, it will mean pain down the line. We used to have a lot of handwritten complexity because of certain design pattern culture, but they used to be contained because the ability to generate them is costly. Now it's much more risky and therefore more important to have simplicity as the guiding principle in ALL projects.
- chicken-stew 3mo ago“How many Kloc did you submit today? I did 12.”
- alex1138 3mo ago"How much on user account support?" "..........."
- mkozlows 3mo agoI kinda love that you made this post feel negative enough that a bunch of AI skeptics are enthusiastically agreeing with a post suggesting that the realistic, pragmatic bear case for AI is, uh, 2-3x productivity improvements.
- onion2k 3mo agoDo I write all my code with an agent now? Yes. Can you just give an agent a desired outcome and let it work, unsupervised? Absolutely not. I strongly suspect that developers moving from writing code to managing agents to write code for them is very similar to developers moving into leadership and management roles and managing ICs to write code for them. Some devs just 'get it' and thrive, leading a team really well and building a great culture. But a lot of them don't, especially if they don't get the support necessary to understand what changes when you move from IC to manager. If the team (or agent swarm) isn't performing well it often isn't a problem with them. It's a problem with the new manager still trying to stay on top of everything and micromanaging all the things. Alternatively, the new manager is completely hands off and only appears at a check-in point (one-to-one, agent completes a task, etc) where they crap on the work and get cross. I have no evidence for this, but I'd guess that putting developers through some sort of management training would make them much better at using agentic swarms.
- geraneum 3mo ago> If the team (or agent swarm) isn't performing well it often isn't a problem with them. It's a problem with the new manager still trying to stay on top of everything and micromanaging all the things. You see problems in the results? Simple, don’t check the results!
- Foobar8568 3mo agoLet codex review Claude output and the contrary. Win for everyone involved, more code, more features, more token usage, promotions, code to be debug by juniors and the cycle of life can continue.
- hatefulheart 3mo agoAbsurd. We are beset on all sides with companies declaring agentic coding a failure and here you are stating as a matter of fact some teams “thrive” with this probabilistic expensive approach to approximating working code? All the while concluding with “I have no evidence for any of this”.
- adamddev1 3mo agoDo you mean that to implement the same features, the agents write 2-3 times as much code as a human would write before?
- entrope 3mo agoMy experience is that AI agents write 20-33% more code than I would for a given feature set, mostly because they are worse at remembering what utility functions already exist and less likely to merge similar functions into more generic ones. They generate that code 2-10 times faster than I could. Defect density is harder to compare: I probably generate more "dumb" defects due to oversight or missing unit tests, but fewer defects that violate domain rules or architectural objectives.
- noosphr 3mo agoMeasuring software performance by lines of code is like measuring aircraft performance by weight. I have no idea why everyone seems to have forgotten this simple fact over the last four years.
- Tenoke 3mo agoMeasuring aircraft production per weight doesnt sound like that bad of a proxy. If I hear that Boeing produced 300 kilotons worth of aircrafts more this year, I'd be right to suspect they've ramped up production.
- noosphr 3mo agoEvery day we stray further into depravity and barbarism.
- entrope 3mo agoI think you're overlooking the lesson of Goodhart's law: you can use a metric, but if you make it a target, it stops being a good metric. Neither "tons of aircraft" nor "lines of code" should be the measuring stick -- and if they're not, then they can still be used as metrics. To be fair, the hazard with AI agents is that they generate fluent output that is often facile, so it's easy to do a lot of things while having a lot of defects. That's a sign that quality control is not prioritized enough. A change in quality will also reduce the utility of SLOC as a metric, but the mechanism is different than what Charles Good hart pointed out.
- Tenoke 3mo agoHave you honestly not noticed how dried up the tech job market is? This take is bizzare to me.
- andrewaylett 3mo agoThe thing is: I could produce 2-3 times as much code as before _without_ an LLM, if I didn't care about my colleagues' ability to review my output properly. Lines of code are a liability, not an asset. You want as few of them as you can get away with, without compromising the actual asset: the functionality. A huge part of the job of Software Engineering is producing the right amount of code at the right time.
- ben_w 3mo ago> Lines of code are a liability, not an asset. You want as few of them as you can get away with, without compromising the actual asset: the functionality. > A huge part of the job of Software Engineering is producing the right amount of code at the right time. Absolutely true, however my experience says that the correlation between "good software engineering practices" and "positive business outcomes" is, at best, small. 120 kloc mostly from one single developer copy-pasting and keeping non-compilable code for an obsolete target "for reference" for a decade, becoming both a ball of mud and a whole pantheon of god classes? No unit tests, no code review? Won awards. Properly engineered, mandatory code review, mandatory unit tests, dev meetings to knowledge-share? People with the money said too slow, closed it down. (Sometimes people bring up how bad Musk's code was at PayPal. I never bothered investigating. Successful product though, wasn't it?)
- hatefulheart 3mo agoSurvivorship bias, you don’t know of all the failed projects that couldn’t get off the ground because of incompetent development team and practices that lead a product to its demise, or a product that is possible within constraints that otherwise could have been a success, but not realised by sloppy work and incompetence. Furthermore the dependencies you choose to build your product are presumably filtered for engineering practices or world class engineers. So given the choice you yourself prefer top quality engineering, so do your customers. Much in the same way you are a customer of your projects dependencies. Difference being, as developers we get to see how the sausage is made, our customers only see second and third order effects.
- kvgr 3mo agoBack in the day there was a mantra: best amount od code is 0. Now we have agents spitting lines after lines. I am not afraid of my future. Even if one person can do work of 5, the amount of generated code will grow exponentially. And not everything can he vibecoded with 0 knowledge. There is a complexity that need understanding to change and optimize. For now :)
- throwaw12 3mo ago> Back in the day there was a mantra: best amount od code is 0. It was true for that time, because producing and maintaining the code was done by humans with limited speed of comprehension. Today, we might challenge this assumption (not saying its wrong or right), because migrations can be done in 1-2 weeks with hundreds of agents.
- yieldcrv 3mo agoMake sure to have an agent audit your codebase for redundant code and dead code When you come back to the codebase in two weeks or a few months, the agent just redoes stuff
- PunchyHamster 3mo agoso 3 times more code to maintain in long term too. With no human that actually understands it on staff I feel like even those benefits gonna melt pretty quickly. It's great as code review buddy tho
- pipes 3mo agoI'm really struggling to get an agent to write code I'm happy with. It's mostly pretty awful. I've a fairly simple c# coding style. But simple is proving a bit more difficult to convey than I thought. I get it to produce code. I then have to spend along time convincing myself it's correct. If I don't I end up embarrassing myself when a coworker reviews it, questions it and it's obvious I don't properly understand it. This is really starting to screw with me mentally. It's like everyone in the world is saying they can fly by flapping their arms (dark factories). When I try I just stay in the same spot burning a lot of energy.
- ryandrake 3mo agoI don't think everyone in the world is saying they can fly by flapping their arms. It's a small number of very vocal, very online, AI enthusiasts, many who have a financial stake in AI winning.
- efficax 3mo agowork from tests to implementation. Validate the tests, and work with the agent to ensure there are not more cases that need testing. Then you can let the agent implement the code and you can refine it until it's simple enough but covers the test cases. TDD is the only way to use agents effectively, imo
- pipes 3mo agoThanks. I have been doing this tdd abd it's a big improvement but the code is still pretty awful. Two skills I rely on are Matt pococks grill-me and his tdd. These massively helped, and technically the code generated is correct but it's still hard to follow and bloated. https://github.com/mattpocock/skills/blob/main/skills/productivity/grill-me/SKILL.md https://github.com/mattpocock/skills/blob/main/skills/produc...
- stasomatic 3mo agoAs a hobbyist coder, I wonder how much more new code, fundamentally, there needs to be? New devices need new firmware, new cars etc, but how much of that is bespoke? Sure, a new movie or a book is new entertainment, but I've already seen that movie and read that book, they just had different jackets. What do these "engineers" actually do that is novel and how much of the pizza is the novel slice is?
- tracker1 3mo agoI would add that a lot of that extra code is often in test/demo paths... I tend to think of working with an Agent as a "team" of 1 + agent... where the developer is now wearing a QA and PM hat in addition to lead/sr dev. That the work getting done is now roughly the offset of what a team would have produced and that coordination needs to step back and treat each individual with an agent as roughly a dev team. Coordination overhead and mythical man month still apply, just at a layer up.
- tangweigang 3mo ago[flagged]