5 ms·
I think the core idea here is a good one. But in many agent-skeptical pieces, I keep seeing this specific sentiment that “agent-written code is not production-
by ketzo 6mo ago
I think the core idea here is a good one.
But in many agent-skeptical pieces, I keep seeing this specific sentiment that “agent-written code is not production-ready,” and that just feels… wrong!
It’s just completely insane to me to look at the output of Claude code or Codex with frontier models and say “no, nothing that comes out of this can go straight to prod — I need to review every line.”
Yes, there are still issues, and yes, keeping mental context of your codebase’s architecture is critical, but I’m sorry, it just feels borderline archaic to pretend we’re gonna live in a world where these agents have to have a human poring over every single line they commit.
- movedx01 6mo agoNot having a code review process is archaic engineering practice at this point(at any point in history, really), be it for human written or AI written code.
- bluGill 6mo agoMaybe in the future humans won't need to pour over every line. However I quickly learn which interns I can trust and which I need to pour over their code - I don't trust AI because it has been wrong too often. I'm not saying AI is useless - I do most of my coding with an agent, but I don't trust it until I verify every line.
- bensyverson 6mo agoI did this for a while… and until Opus 4.5, I couldn't fully trust the model. But at this point, while it does make the occasional mistake, I don't need to scrutinize every line. Unit and integration tests catch the bugs we can imagine, and the bugs we can't imagine take us by surprise, which is how it has always been.
- bluGill 6mo agoEven with 4.6 I find there are a lot of mistakes it makes that I won't allow. Though it is also really good at finding complex thread issues that would take me forever...
- pixl97 6mo agoWe live in a world where every line of code written by a human should be reviewed by another human. We can't even do that! Nothing should go straight to prod ever, ever ever, ever.
- latchkey 6mo ago> Nothing should go straight to prod ever, ever ever, ever. I'm one-shotting AI code for my website without even looking at it. Straight to prod (well, github->cf worker). It is glorious.
- bikelang 6mo agoWere people reviewing your hobby projects previously? Were you on-call for your hobby website? If not - then it sounds like nothing changed?
- latchkey 6mo agoThis is my business website.
- ehsanu1 6mo agoThat a personal website? Prod means different things in different contexts. Even then, I'd be a bit worried about prompt injection unless you control your context closely (no web access etc).
- miltonlost 6mo agoYou say it's borderline archaic. I say trusting agents enough to not look at every single line is an abdication of ethics, safety, and engineering. You're just absolving yourself of any problems. I hope you aren't working in medical devices or else we're going to get another Therac-25. Please have some sort of ethics. You are going to kill people with your attitude.
- tru1ock 6mo agoAlmost nobody works on medical devices... And some of you lucky folks might be working with mega minds everyday, but the rest of us are but shadows and dust. I trust 5.4 or 4.6 more than most developers. Through applying specific pressure using tests and prompts I force it to built better code for my silly hobby game than I ever saw in real production software. Before those models I was still on the other side of the line but the writing is on the wall.
- bikelang 6mo agoWere you not reviewing every line when a human wrote it before it went to prod? I think the output of these tools is about as good as a human would write - which means it needs thorough review if I’m going to be on the hook to resolve its issues at 2AM.
- alecbz 6mo agoYeah in many places we had two humans with context on every line, and now we're advocating going to zero?
- AnimalMuppet 6mo agoMaybe that's the distinction. If I write it, you can call me at 2AM. If an AI wrote it, call the AI at 2AM. Oh, it can't take the phone call and fix the issue? Then I'm reviewing its output before it goes into prod.
- cableshaft 6mo agoThis is a weird analogy. You can ask the A.I. to fix the issue at any time of day (assuming the person asking someone with enough technical knowledge that can evaluate the fix at least). You won't always be able to get ahold of someone at 2am. You won't be able to get ahold of me at 2am, for example. It'll throw some notification on my screen and I won't see it until I wake up.
- SpicyLemonZest 6mo agoIt's a conversation I've had many times in my career and I'm sure I'll have many more. We've got code that seems plausible on a surface level, at a glance it solves the problem it's meant to solve - why can't we just send it to prod and address whatever problems we find with it later? The answer is that it's very easy for bad code to cause more problems than it solves. This: > Then one day you turn around and want to add a new feature. But the architecture, which is largely booboos at this point, doesn't allow your army of agents to make the change in a functioning way. is not a hypothetical, but a common failure mode which routinely happens today to teams who don't think carefully enough about what they're merging. I know a team of a half-dozen people who's been working for years to dig themselves out of that hole; because of bad code they shipped in the past, changes that should have taken a couple hours without agentic support take days or weeks even with agentic support.
- alecbz 6mo agoHow do you know which lines you need to review and which you don't? Does it feel archaic because LLMs are clearly producing output of a quality that doesn't require any review, or because having to review all the code LLMs produce clips the productivity gains we can squeeze out of them?
- postexitus 6mo agoYou sound like you are working on unimportant stuff. Sure, go ahead, push.
- MrScruff 6mo agoHonestly a lot of useful software is ‘unimportant’ in the sense that the consequences of introducing a bug or bad code smell aren’t that significant, and can be addressed if needed. It might well be for many projects the time saved not reviewing is worth dealing with bugs that escape testing. Also, it’s entirely possible for software to be both well engineered and useless.
- postexitus 6mo agoExactly - not so much in "important" stuff.
- layer8 6mo agoIt’s not archaic, it’s due diligence, until we can expect AI to reliably apply the same level of diligence — which we’re still pretty far off from.
- slopinthebag 6mo agoIf you keep the scope small enough it can be production ready ootb, and with some stuff (eg. a throwaway React component) who really cares. But I think it's insane to look at the output of Claude Code or Codex with frontier models and say "yep, that looks good to me". Fwiw OP isn't an agent skeptic, he wrote one of the most popular agent frameworks.
- bigstrat2003 6mo ago> It’s just completely insane to me to look at the output of Claude code or Codex with frontier models and say “no, nothing that comes out of this can go straight to prod — I need to review every line.” It's insane to me that someone can arrive at any other conclusion. LLMs very obviously put out bad code, and you have no idea where it is in their output. So you have to review it all.
- mememememememo 6mo agoDepends on your prod. For an early startup validating their idea, that prod can take it. For a platform as a service used by millions, nope.
- manmal 6mo agoThe article didn't say to read every line though. Just the interesting ones. If you don't know where the interesting ones are, you have already lost.