7 ms·
I think this is a bit of a simplistic mental approach. I've certainly seen a lot of "The engineer owns the outcome, AI is just a tool, don't release anything yo
by zug_zug 11d ago
I think this is a bit of a simplistic mental approach. I've certainly seen a lot of "The engineer owns the outcome, AI is just a tool, don't release anything you don't vouch for."
However, I just don't think that's realistic. It's asking an author to suddenly become an editor. It's asking somebody who writes code to now read and debug others code.
It can actually be harder to find the the bug in a tricky piece of code than it can be to write your own correct code from scratch. I see AI introduce all sorts of bugs all the time in my personal projects that I would never introduce, and would never think to test for, especially around anything graphical.
- christophilus 11d ago> It's asking somebody who writes code to now read and debug others code. This has been a big part of the job for anyone on a team for at least 20 years. I do agree that it’s the hardest and worst part of the job, and has now become the majority of the job for anyone who isn’t vibe coding. So, that sucks.
- phrotoma 11d agoIt's a different of degree, not kind. Anybody who has reviewed pull requests can tell you that sooner or later you approve a PR after many rounds of changes because it's finally "good enough". Fighting with a robot to just do the damned thing is less fraught because they don't get offended by critiques but it takes more round trips to get them pointed in the direction you want.
- thw_9a83c 11d agoFighting with a robot requires also a different kind of attention. When you're reviewing the human code, you can quite easily guess an overall seniority and competency level of the author and then you can adjust your level of attention to every detail. E.g. if the solution requires an understanding of some core idea, ones the human understands this core idea, you can be quite sure that it is consistently implemented everywhere. With AI, 90% of the PR could be expertly implemented but then, for no obvious reason, 10% could be low-quality surprise. I've never seen such unbalanced output from human programmers.
- zahlman 11d agoIf 90% of it was fine, maybe it would be better to just fix the 10% yourself rather than "fighting with a robot" to try to get an automated fix.
- thw_9a83c 10d agoYes, but those 10% of a problematic code is not easy to find without a very detailed study of the whole PR. And since most of the code looks (and usually is) very well-written, the human brain somehow doesn't expect to find those low-quality or sub-optimal parts in such code. That's why I wrote that reviewing the AI code requires different kind of attention.
- OptionOfT 11d agoDisagree. At least back in the day there weren't endless comments about how this widget is load-bearing, and how honestly the other widget carries the derived widget, referencing decision ADR-100 that is nowhere to be found. All these comments matter because once accepted as part of the codebase the next LLM takes these comments as canonical. The largest problem these days is the volume of code developers are expected to review. The volume went up significantly.
- Daishiman 11d ago> At least back in the day there weren't endless comments about how this widget is load-bearing By far the biggest problem 90% of developers have with AI is that they should be turning off comments, as it's clear that the training data they have is no good for developing a theory of mind for an engineer who has to read them. I've turned them off and add them myself at review time and am quite happy.
- zahlman 11d agoRequiring the coding agent to (try to) iterate on code clarity until comments are no longer necessary, probably doesn't hurt either. Save the commentary for conversation logs, agent Markdown files, and other sorts of documentation.
- theshrike79 11d ago> the next LLM takes these comments as canonical This is the best and worst thing about LLM coding agents. They trust comments way too implicitly. And then the errors just keep compounding. Or a temporary hack that becomes "load-bearing" because the agent doesn't figure out that it's supposed to be a temporary testing shim - instead it keeps building on it until it basically duplicates what it's mocking.
- geertj 11d ago> It's asking an author to suddenly become an editor. I think that’s right, and what is needed. It still gives a significant speed up for coding, while still keeping the output human maintainable. There is the idea that the agent will just produce binary code directly at some point. I don’t know if it ever comes to that but for now I’m in the ‘I’ve become an editor’ camp.
- bigstrat2003 11d agoThe time it takes you to review the code the LLM produces is the same amount of time it would take to just write the code yourself. There's no speedup to be had using these things, despite what many claim.
- sfn42 11d agoStrongly disagree. I know my codebase well, I know what I'm expecting before I ask Claude to do it, and I tell it what I'm expecting. I might tell it roughly what I want, have it make a plan, review the plan and ask for changes if necessary, then execute. This way I don't need to scrutinize every detail, I just look over the big picture. I also care a lot more about the big picture - architecture and data flow etc. Basically if you view your codebase as a tree I care much more about the trunk and the big branches than I do about the smaller branches and particularly the leaves. So the details of some little leaf function somewhere are fairly insignificant, it's trivial to change at any time. As long as it works and isn't unreasonably slow it's fine. Working this way I can get things done in minutes or hours that would previously take days or even weeks.
- abalashov 10d ago> I know my codebase well I'll bet you know it because you wrote and/or worked on it manually, likely over a period of years. The odds of you knowing a slop codebase that well, or even particularly at all, are much lower.
- arcanemachiner 11d agoI think the answer is not to debug the code, but, when possible, to debug the outputs. The code may be considered to be a black box much of the time. (This is much more true for my hobby projects than my work projects.)
- yosefk 11d ago...because your work projects are bigger, there's only so big a black box can get before you lose all comprehension of it, and splitting it to smaller black boxes the shapes of which you keep refining is programming, and the part of it LLMs currently can't do
- dist-epoch 11d agoI routinely see Astra extract related functionality into it's own file after it gets past a certain size, unprompted. And prompted it can extract the black boxes if you tell it what the boxes are or what to look for. Same for cleaning up tech debt after organic development, it's suggestions on how to simplify and modularize are good, but you need to prompt. Given that the prompts are quite generic, "look for technical debt, suggest simpler architectures, what could be extracted in a separate module", it won't be long till it will do it on it's own.
- sameerds 11d ago> It's asking somebody who writes code to now read and debug others code. That's exactly right. Open source projects are currently drowning under LLM generated PRs, where those who used to write code are simply punting that work to AI, but still expecting others to review it. It's not okay to expect such a free lunch. If you moved the labour of writing code one step away, then you are yourself the first line of defence now, so you better start reviewing code that you claim to be yours.
- sfn42 11d agoThat's what I do and expect my colleagues to do. Even before LLMs I was reviewing my own PRs before submitting them to others. I still do that. I work closely with Claude to create something good that I'm happy with, then I review it and test it to ensure it's good. And only then do I submit the PR to colleagues for final review. I expect the same from colleagues, I'm not interested in treating them as a middle man between me and Claude.
- skybrian 11d agoIf you can explain how to reproduce a bug, you can ask the AI to debug the code and it usually works, in my experience. If not, you can ask it to add logging or other tools for better observability.
- CoolestBeans 11d agoI agree. When you write your own code, you know what your intention was when writing it. Furthermore, as you gain experience and mature you know in the back of your mind that every mistake during code writing costs disproportionately more to fix later on. You only get that feeling by owning the code. AI cannot do that. It can't have skin in the game in that way.
- Daishiman 11d ago> I see AI introduce all sorts of bugs all the time in my personal projects that I would never introduce, and would never think to test for, especially around anything graphical. This is referred to in the need for E2E testing and E2E testing not being a substitute. Code review is definitely the biggest challenge of AI-driven development IMO. I still have not found good processes that work in my org, but for my personal work I independently reached the author's conclusions a while ago and am very satisfied with the results.
- zahlman 11d ago> It's asking somebody who writes code to now read and debug others code. Writing code has always involved reading and debugging your own code, at an absolute minimum, even if you did everything solo. In any remotely serious collaborative effort, it also involved code review and collaborative debugging; people use issue trackers and assign themselves and each other "tickets", which often involve fixing issues that are ultimately caused by someone else's code. > It can actually be harder to find the the bug in a tricky piece of code than it can be to write your own correct code from scratch. Part of the point is to reject tricky code exactly because it is tricky (as this is rarely actually necessary).
- DANmode 10d ago> It's asking an author to suddenly become an editor. It’s asking an author to suddenly become an editor if they decide to use the robot for a task. Certain workplaces are demanding this - but not all. Many still just want working commits without tech debt. In fact, private and public teams alike are backed up at the PR review stage, so, lots of sane places wouldn’t mind individual contributors using the robot less - especially if its use increases the complexity of reviewing the task. Speed isn’t the only variable to optimize for!