6 ms·
Agreed, but: There's been a notable jump over the course of the last few months, to where I'd say it's inevitable. For a while I was holding out for them to hi
by daxfohl 9mo ago
Agreed, but:
There's been a notable jump over the course of the last few months, to where I'd say it's inevitable. For a while I was holding out for them to hit a ceiling where we'd look back and laugh at the idea they'd ever replace human coders. Now, it seems much more like a matter of time.
Ultimately I think over the next two years or so, Anthropic and OpenAI will evolve their product from "coding assistant" to "engineering team replacement", which will include standard tools and frameworks that they each specialize in (vendor lock in, perhaps), but also ways to plug in other tech as well. The idea being, they market directly to the product team, not to engineers who may have specific experience with one language, framework, database, or whatever.
I also think we'll see a revival of monolithic architectures. Right now, services are split up mainly because project/team workflows are also distributed so they can be done in parallel while minimizing conflicts. As AI makes dev cycles faster that will be far less useful, while having a single house for all your logic will be a huge benefit for AI analysis.
- smt88 9mo agoThere's no chance LLMs will be an engineering team replacement. The hallucination problem is unsolvable and catastrophic in some edge cases. Any company using such a team would be uninsurable and sued into oblivion.
- eru 9mo agoWriting software is actually one of the domains where hallucinations are easiest to fix: you can easily check whether it builds and passes tests. If you want to go further, you can even require the LLM to produce a machine checkable proof that the software is correct. That's beyond the state of the art at the moment, but it's far from 'unsolvable'. If you hallucinate such a proof, it'll just not work. Feed back the error message from the proof checker to your coding assistant, and the hallucination goes away / isn't a problem.
- ulrikrasmussen 9mo agoIn order to prove safety you need a formal model of the system and formally defined safety properties that are both meaningful and understandable by humans. These do not exist for enterprise systems
- eru 9mo agoAn exhaustive formal spec doesn't exist. But you can conservatively proof some properties. Eg program termination is far from sufficient for your program to do what you want, but it's probably necessary. (Termination in the wider sense: for example an event loop has to be able to finish each run through the loop in finite time.) You can see eg Rust's or Haskell's type system as another light-weight formal model that lets you make and proof some simple statements, without having a full formal spec of the whole desired behaviour of the system.
- tsimionescu 9mo agoThat is true and very useful for software development, but it doesn't help if the goal is to remove human programmers from the loop entirely. If I'm a PM who is trying to get a program to, say, catalogue books according to the Dewey Decimal system for a library, a proof that the program terminates is not going to help that much when the program is mis-categorizing some books.
- seanmcdirmid 9mo agoIs removing the human in the loop really the goal, or is the goal right now to make the human a lot more productive? Because...those are both very different things.
- tsimionescu 9mo agoI don't know what the goal for OpenAI or Anthropic really is. But the context of this thread is the idea that the user daxfohl launched that these companies will, in the next few years, launch an "engineering team replacement" program; and then the user eru claimed that this is indeed more doable in programming than other domains because you can have specs and tests for programs in a way that you can't for, say, an animated movie.
- solid_fuel 9mo ago> Writing software is actually one of the domains where hallucinations are easiest to fix: you can easily check whether it builds and passes tests. What tests? You can't trust the tests that the LLM writes, and if you can write detailed tests yourself you might as well write the damn software.
- eru 9mo agoUse multiple competing LLM. Generative adversarial network style.
- solid_fuel 9mo agoCool. That sure sounds nice and simple. What do you do when the multiple LLMs disagree on what the correct tests are? Do you sit down and compare 5 different diffs to see which have the tests you actually want? That sure sounds like a task you would need an actual programmer for. At some point a human has to actually use their brain to decide what the actual goals of a given task are. That person needs to be a domain expert to draw the lines correctly. There's no shortcut around that, and throwing more stochastic parrots at it doesn't help.
- eru 9mo agoJust because you can't (yet) remove the human entirely from the loop, doesn't mean that economising on the use of the humans time is impossible. For comparison have a look at compilers: nowadays approximately no one writes their software by hand, we write a 'prompt' in something like Rust or C, and ask another computer program to create the actual software. We still need the human in the loop here, but it takes much less human time than creating the ELF directly.
- exceptione 9mo agoI can see how LLMs can help with testing, but one should never compare LLMs with deterministic tools like compilers. LLMs are entirely a separate category.
- DrammBA 9mo agoYou focused on writing software, but the real problem is the spec used to produce the software, LLMs will happily hallucinate reasonable but unintended specs, and the checker won’t save you because after all the software created is correct w.r.t. spec. Also tests and proof checkers only catch what they’re asked to check, if the LLM misunderstands intent but produces a consistent implementation+proof, everything “passes” and is still wrong.
- simonw 9mo agoThis is why every one of my coding agent sessions starts with "... write a detailed spec in spec.md and wait for me to approve it". Then I review the spec, then I tell it "implement with red/green TDD".
- daxfohl 9mo agoSame, and similarly something like a "create a holistic design with all existing functionality you see in tests and docs plus new feature X, from scratch", then "compare that to the existing implementation and identify opportunities for improvement, ranked by impact, and a plan to implement them" when the code starts getting too branchy. (aka "first make the change easy, then make the easy change"). Just prompting "clean this code up" rarely gets beyond dumb mechanical changes. Given so much of the work of managing these systems has become so rote now, my only conclusion is that all that's left (before getting to 95+% engineer replacement) is an "agent engineering" problem, not an AI research problem.
- tsimionescu 9mo agoThe premise was that the AI solution would replace the engineering team, so who exactly is writing/reviewing this detailed spec?
- simonw 9mo agoThat's a bad premise.
- 9mo ago
- mohaine 9mo agoAh, most the problem in programming is writing the tests. Once you know what you need the rest is just typing. I can see an argument where you can get none programers to create the input and output of said tests but if the can do that, they are basically programmers. This is of course leaving aside that half the stated use cases I hear for AI are that it can 'write the tests for you'. If it is writing the code and the tests it is pointless.
- discreteevent 9mo agoYou need more than tests. Test induced design damage: https://dhh.dk/2014/test-induced-design-damage.html https://dhh.dk/2014/test-induced-design-damage.html
- somenameforme 9mo agoTests and proofs can only detect issues that you design them to detect. LLMs and other people are remarkably effective at finding all sorts of new bugs you never even thought to test against. Proofs are particularly fragile as they tend to rely on pre/post conditions with clean deterministic processing, but the whole concept just breaks down in practice pretty quickly when you start expanding what's going on in between those, and then there's multithreading...
- thesz 9mo ago> you can easily check whether it builds and passes tests. This link were on HN recently: https://spectrum.ieee.org/ai-coding-degrades https://spectrum.ieee.org/ai-coding-degrades "...recently released LLMs, such as GPT-5, have a much more insidious method of failure. They often generate code that fails to perform as intended, but which on the surface seems to run successfully, avoiding syntax errors or obvious crashes. It does this by removing safety checks, or by creating fake output that matches the desired format, or through a variety of other techniques to avoid crashing during execution." The trend for LLM generated code is to build and pass tests but do not deliver functionality needed. Also, please consider how SQLite is tested: https://sqlite.org/testing.html https://sqlite.org/testing.html The ratio between test code and code itself is mere 590 times (590 LOC of tests per LOC of actual code), it used to be more than 1100. Here is notes on current release: https://sqlite.org/releaselog/3_51_2.html https://sqlite.org/releaselog/3_51_2.html Notice fixes there. Despite being one of the most, if not the most, tested pieces of software in the world, it still contains errors. > If you want to go further, you can even require the LLM to produce a machine checkable proof that the software is correct. Haha. How do you reconcile a proof with actual code?
- eru 9mo ago> Haha. How do you reconcile a proof with actual code? You can either proof your Rust code correct, or you can use a proof system that allows you to extract executable code from the proofs. Both approaches have been done in practice. Or what do you mean?
- thesz 9mo agoRust code can have arbitrary I/O effects in any parts of it. This precludes using only Rust's type system to make sure code does what spec said. The most successful formally proven project I know, seL4 [1], did not extracted executable code from the proof. They created a prototype in Haskell, mapped (by hand) it to Isabelle, I believe, to have a formal proof and then recreated code in C, again, manually. [1] https://sel4.systems/ https://sel4.systems/ Not many formal proof systems can extract executable C source.
- 9mo ago
- shevy-java 9mo agoWell - the end result can be garbage still. To be fair: humans also write a lot of garbage. I think in general most software is rather poorly written; only a tiny percentage is of epic prowess.
- rezonant 9mo agoWho writes the tests?
- Marazan 9mo agoWho is writing the tests?
- fragmede 9mo agoI am but a lowly IC, with no notion of the business side of things. If I am an IC at, say, a FANG company, what insurance has been taken out on me writing code there?
- smt88 9mo ago> If I am an IC at, say, a FANG company, what insurance has been taken out on me writing code there? Every non-trivial software business has liability insurance to cover them for coding lapses that lead to data breaches or other kinds of damages to customers/users.
- kristiandupont 9mo agoI use LLM's to write the majority of my code. I haven't encountered a hallucination for the better part of a year. It might be theoretically unsolvable but it certainly doesn't seem like a real problem to me.
- deleted 9mo ago[deleted]
- smt88 9mo agoI use LLMs whenever I'm coding, and it makes mistakes ~80% of the time. If you haven't seen it make a huge mistake, you may not be experienced enough to catch them.
- kristiandupont 9mo agoHallucinations, no. Mistakes, yes, of course. That's a matter of prompting.
- tdrz 9mo ago> That's a matter of prompting. So when I introduce a bug it's the PM's fault.
- matwood 9mo agoThese types of comments are interesting to me. Pre-chatGPT there were tons of posts how so many software people were terrible at their jobs. Bugs were/are rampant. Software bugs caused high profile issues, but likely so many more we never heard about. Today we have chatGPT and only now will teams be uninsurable and sued into oblivion? LOL
- elzbardico 9mo agoLLMs were trained on exactly that kind of code.
- smt88 9mo agoIf you've ever used Claude Code in brave mode, I can't understand how you'd think a dev team could make the same categories of mistakes or with the same frequency.
- sublinear 9mo agoThis doesn't make any sense. If the business can get rid of their engineers, then why can't the user get rid of the business providing the software? Why can't the user use AI to write it themselves? I think instead the value is in getting a computer to execute domain-specific knowledge organized in a way that makes sense for the business, and in the context of those private computing resources. It's not about the ability to write code. There are already many businesses running low-code and no-code solutions, yet they still have software engineers writing integration code, debugging and making tweaks, in touch with vendor support, etc. This has been true for at least a decade! That integration work and domain-specific knowledge is already distilled out at a lot of places, but it's still not trivial. It's actually the opposite. AI doesn't help when you've finally shaved the yak smooth.
- chongli 9mo agoIf the business can get rid of their engineers, then why can't the user get rid of the business providing the software? A lot of businesses are the only users of their own software. They write and use software in-house in order to accomplish business tasks. If they could get rid of their engineers, they would, since then they'd only have to pay the other employees who use the software. They're much less likely to get rid of the user employees because those folks don't command engineer salaries.
- hypeatei 9mo agoSo instead of paying a human that "commands an engineer salary" then they'll be forced to pay whatever Anthropic or OpenAI commands to use their LLMs? I don't see how that's a better proposition: the LLM generates a huge volume of code that the product team (or whoever) cannot maintain themselves. Therefore, they're locked-in and need to hope the LLM can solve whatever issues they have, and if it can't, hope that whatever mess it generated can be fixed by an actual engineer without costing too much money. Also, code is only a small piece and you still need to handle your hosting environment, permissions, deployment pipelines, etc. which LLMs / agentic workflows will never be able to handle IMO. Security would be a nightmare with teams putting all their faith into the LLM and not being able to audit anything themselves. I don't doubt that some businesses will try this, but on paper it sounds like a money pit and you'd be better off just hiring a person.
- catlifeonmars 9mo agoI actually think it’s the opposite. We’ll see fewer monorepos because small, scoped repos are the easiest way to keep an agent focused and reduce the blast radius of their changes. Monorepos exist to help teams of humans keep track of things.
- daxfohl 9mo agoCould be. Most projects I've worked on tend to span multiple services though, so I think AI would struggle more trying to understand and coordinate across all those services versus having all the logic in a single deployable instance. The way I see feature development in the future is, PM creates a dev cluster (also much easier with a monolith), has AI implement a bunch of features to spec, AI provides some feedback and gets input on anywhere it might conflict with existing functionality, whether eventual consistency is okay, which pieces are performance criticial, etc., and provides the implementation, a bunch of tests for review, and errata about where to find observability data, design decisions considered and chosen, etc. PM does some manual testing across various personas and products (along with PMs from those teams), has AI add feature flags, launches. The feature flag rollout ends up being the long-pole, since generally the product team needs to monitor usage data for some time before increasing the rollout percentage. So I see that kind of workflow as being a lot easier in a monolithic service. Granted, that's a few years down the road though, before we have AI reliable enough to do that kind of work.
- catlifeonmars 9mo ago> Most projects I've worked on tend to span multiple services though, so I think AI would struggle more trying to understand and coordinate across all those services versus having all the logic in a single deployable instance. 1. At least CC supports multiple folders in a workspace, so that’s not really a limitation. 2. If you find you are making changes across multiple services, then that is a good indication that you might not have the correct abstraction on the service boundary. I agree that in this case a monolith seems like a better fit.
- daxfohl 9mo ago
- concats 9mo ago> Ultimately I think over the next two years or so, Anthropic and OpenAI will evolve their product from "coding assistant" to "engineering team replacement" The way I see it, there will always be a layer in the corporate organization where someone has to interact with the machine. The transitioning layer from humans to AIs. This is true no matter how high up the hierarchy you replace the humans, be it the engineers layer, the engineering managers, or even their managers. Given the above, it feels reasonable to believe that whatever title that person has—who is responsible for converting human management's ideas into prompts (or whatever the future has the text prompts replaced by)—that person will do a better job if they have a high degree of technical competence. That is to say, I believe most companies will still want and benefit if that/those employees are engineers. Converting non-technical CEO fever dreams and ambitions into strict technical specifications and prompts. What this means for us, our careers, or Anthropic's marketing department, I cannot say.
- eloisant 9mo agoThat reminds me of the time where 3GL languages arrived and bosses claimed they no longer needed developers, because anyone could write code in those English-like languages. Then when mouse-based tools like Visual Basic arrived, same story, no need for developers because anyone can write programs by clicking! Now bosses think that with AI anyone will be able to create software, but the truth is that you'll still need software engineers to use those tools. Will we need less people? Maybe. But in the past 40 years we have been increasing the developers productivity so many times, and yet we still need more and more developers because the needs have grown faster.
- deleted 9mo ago[deleted]
- OkayPhysicist 9mo agoMy suspicion is that it will be bad for salaries, mostly because it'll kill the "looks difficult" moat that software development currently has. Developers know that "understanding source code" is far from the hard part of developing software, but non-technical folks' immediate recoiling in the face of the moon runes has kept our profession pretty easy to justify high pay for for ages. If our jobs transition to largely "communing with the machines", then we'll go from a "looks hard, is hard" job, to a "looks easy, is hard" job, which historically hurts bargaining power.
- twelvedogs 9mo agohonestly i think they got the low hanging fruit already. they're bumping up against the limits of what it can do and while it's impressive it's not spectacular
- embedding-shape 9mo agoMaybe I'm easily impressed, but that LLMs even work to output basic human-like text to me is bananas, and I do understand a bit of how it works, yet it's still up there as "Amazing that huge airplanes even can fly" is for me.
- ericmcer 9mo agoIf you research how something like Cursor works I don't think you would believe it is inevitable. The jump that would have to happen for it to replace engineers entirely is insurmountable. They can keep expanding contexts and coming up with clever ways to augment generation but I don't see it ever actually having full vision on the system, product and users. Beyond that it is incredibly biased towards existing code & prompt content. If you wanted to build a voice chat app, and you said "should I use websockets or http?" It would say Websockets. It won't override you and say "Use neither, you should use webRTC", but an experienced engineer would spot that the prompt itself is flawed instantly. LLMs just will bias towards existing tokens in the prompt and won't surface data that would challenge the question itself.
- aldanor 9mo agoUnless you, well, state in AGENTS.md that prompts may offer suboptimal options in which case it's the machine's duty to question them, treat the prompter like a coworker and not a boss.
- apstls 9mo agoSit down and re-read your comment one night with your "I am an engineer and will solve this as an engineering problem" hat firmly on. If you stop thinking of LLMs as lobotimized coworkers trapped inside an API wrapper and instead as computational primitives then things become much more interesting and the future becomes clearer to see.