4 ms·
Software engineering has always worked this way, just not to ICs. “The LLMs produce non-deterministic output and generate code much faster than we can read it,
by iloveoof 4mo ago
Software engineering has always worked this way, just not to ICs.
“The LLMs produce non-deterministic output and generate code much faster than we can read it, so we can’t seriously expect to effectively review, understand, and approve every diff anymore. But that doesn’t necessarily mean we stop being rigorous, it could mean we should move rigor elsewhere.“
Direct reports, when delegated tasks by managers, product non-deterministic outputs much faster than team leads/managers can review, understand or approve every diff. Being a manager of software developers has always been a non-deterministic form of software engineering.
- devmor 4mo ago> Being a manager of software developers has always been a non-deterministic form of software engineering. Unless the manager is also a principal/architect, I don’t find this to be agreeable. It’s similar to saying that you are a non-deterministic chef when you order food from a restaurant.
- teaearlgraycold 4mo agoWell yes but if no humans at the company understand the code then no one is truly responsible for it.
- Npovview 4mo agowhat about the artifacts that were supposed to test the correctness of the code? are they passing willy nilly?
- ekidd 4mo agoNo amount of testing will save a large program with a dogshit architecture. Roughly, this is because tests increase coverage linearly with the number of tests, but weird interactions increase exponentially with code size. This might be fine if you're building a tiny app, or if you're building a medium-sized app that follows a strict existing architecture (like a web app consisting mostly of forms). In which case, have fun. But if you're building something slightly novel and interesting, then Claude is surprisingly bad at architecture and taste, and it tends to "fix" problems by spewing more slop. What you need instead is actual insight that leads to simplifying principles. This, in turn, allows breaking up the exponential complexity into disciplined patterns. This allows your code complexity to scale far more slowly, allowing an essentially linear number of tests to provide coverage. I actually download and try people's vibe-coded developer tools. And frankly, those tools are some of the worst software I've used in my life, worse than even Unix-vendor Motif implementations from the early 90s. Like, I'm super happy that people can vibe-code themselves simple, one-off personal tools. That's incredibly empowering. But that doesn't mean you can big, novel stuff the same way without a competent human actively in the loop.
- theshrike79 4mo ago> those tools are some of the worst software I've used in my life Is the code bad or don't they do what they claim they do? Both are very different issues.
- ekidd 4mo agoThey do what they claim to do maybe 20% of the time. The other 80% of the time is spent trying to figure out why they aren't working, why they corrupted their data, why they crash every 10 minutes, etc. And I want to be clear that this isn't some non-technical novice vibe coding this garbage. This is often extremely talented developers with decades of experience who have apparently decided that they don't need to look at their code anymore. You can get very good results out of AI agents. But mostly the people who get good results are the ones who still read the LLM output in detail, and who introduce the structure the LLMs are missing. But like I said, this distinction mostly becomes apparent past a certain size and novelty level.
- theshrike79 4mo agoWhere do you find these apps that fail to work 80% of the time? I must be an anomaly because all of the vibe coded apps I'm running 24/7 don't keep crashing or stop suddenly working.
- Npovview 4mo agoWould Antirez with LLMs make the same mistakes a novice would make? You are comparing your strongest contender with my weakest contender.
- alabut 4mo agoSimon Willison made a similar parallel recently: https://simonwillison.net/2026/May/6/vibe-coding-and-agentic-engineering/ https://simonwillison.net/2026/May/6/vibe-coding-and-agentic... “The thing that really helps me is thinking back to when I’ve worked at larger organizations where I’ve been an engineering manager. Other teams are building software that my team depends on. If another team hands over something and says, “hey, this is the image resize service, here’s how to use it to resize your images”... I’m not going to go and read every line of code that they wrote. I’m going to look at their documentation and I’m going to use it to resize some images. And then I’m going to start shipping my own features. And if I start running into problems where the image resizer thing appears to have bugs or the performance isn’t good, that’s when I might dig into their Git repositories and see what’s going on. But for the most part I treat that as a semi-black box that I don’t look at until I need to.”
- skydhash 4mo agoBut then, the ownership is clear. And no team would be like to be pointed that their 5th iteration is also broken and can’t be relied for production usage. That’s the difference with AI code. LLM are not aligned with your goals. Any trust in them doing the right thing is very misguided.
- esperent 4mo agoThat's why you have them write tons of tests. Way more than you generally would for human written code. And the agent writing/maintaining the tests is not the agent fixing the bugs. I've personally had a LLM write an image resizing library for me. It's a fairly basic one, I didn't need anything fancy. I could have used something off the shelf but it was at a time when I was testing what Claude could do. And to be honest, it just worked. One shot, if I recall correctly, or at least, one session with a few tweaks and never touched again. It's been embedded in a larger app for several months and I don't recall hitting a single bug with that, specifically. So I'm not sure your complaints about "the 5th iteration" being broken have much grounds here.
- 4mo ago
- simoncion 4mo ago> Being a manager of software developers has always been a non-deterministic form of software engineering. I disagree. Being a manager of programmers requires that you trust your programmers and have some way of occasionally verifying the correctness and efficacy of what they build to make sure that that trust is still properly placed. But on top of that, the user of an LLM isn't really akin to a manager of programmers. Human programmers are responsible for what they write [0], and even the ones that only cost ~50% of a senior's total comp [1] are still going to be able to fairly reliably explain to you why they made the decisions they did, and fairly reliably be able to follow instruction. LLMs just aren't there yet, and the major LLM providers may never care to get them there. A programmer who's using LLMs is a programmer who's using LLMs... not a manager of other programmers. I'm not going to say that the tech will never advance to that point, but it's simply not there yet. [0] Unless management decides otherwise, of course. [1] Nvidia's CEO recently mentioned that he'd be "deeply alarmed" if senior staff aren't spending at least half of their total compensation [2] on LLM providers, so I'm going to use that as my benchmark for "expected annual LLM spend". [2] ...meaning that each senior programmer costs their employer at least 50% more than their total compensation....
- camgunz 4mo agoNo, because those direct reports can use tools to build deterministic software. LLMs can't, because they themselves are non-deterministic. They will say they did, and they will be wrong. And the LLM you have check will also say it did, and it will be wrong. Etc etc. These things just can't be in the critical path. They are ridiculously unreliable.
- insanitybit 4mo agoWhat? Software being deterministic is not a feature of who wrote it. And how the hell is a human "deterministic"?
- camgunz 4mo agoAt the bottom of each of these arguments is a "who is accountable" question. You can tell an LLM "hey, use Lean to verify this" or "hey, do a code review" or "hey, write some tests and run them". It might do these things, it might not. You can then tell another LLM "hey, check that LLM A did these things". It might do these things, it might not. Repeat. You can tell a human (an IC) the same things. You can then tell another human (a manager) "hey, check that IC A did these things". So far, these are the same. But there's now a critical difference: you can then hold those people accountable if they don't. They can be in the critical path. LLMs can't. People can improve. LLMs can't. People can work together. LLMs can't. This doesn't always matter. You don't need things like accountability or improvement or teamwork all the time. But you do in reliable software.