13 ms·
If you are good at code review, you will be good at using AI agents
- shakna 1y ago> bikeshedding function names ... Function names compose much of the API. The API is the structure of the codebase. This isn't some triviality you can throw aside as unimportant, it is the shape that the code has today, and limits and controls what it will have tomorrow. It's how you make things intuitive, and it is equally how you ensure people follow a correct flow and don't trap themselves into a security bug.
- glimshe 1y agoI think I'd actually have a use for an AI that could receive my empty public APIs (such as a C++ header file) as an input and produce a first rough implementation. Maybe this exists already, I don't know because I haven't done any serious vibe coding.
- IanCal 1y agoYou can do this already, the most useful things to help with this are either writing tests or having it write tests and telling it how to compile and see error messages so you can let it loop.
- stuaxo 1y agoYeah it can, though rough is definitely the word. And sometimes the LLM just won't go in the direction you want, but that's OK - you just have to go write those bits of code. It can be suprising where it works and where it doesn't. Just go with those first suggestions though and the code will end up rough.
- jeroenhd 1y agoAs long as you're reinventing the wheel (implementing some common pattern because you don't want to pull in an entire dependency), that kind of AI generation works quite well. Especially if you also have the AI generate tests for its code, so you can force it to iterate on itself while it gets things wrong the first couple of tries. It's slow and resource intensive, but it'll generate something mostly complete most of the time. I'm not sure if you're saving any time there, though. Perhaps if you give an LLM task before ending the work day so it can churn away for a while unattended, it may generate a decent implementation. There's a good chance you need to throw out the work too; you can't rely on it, but it can be a nice bonus if you're lucky. I've found that this only works on expensive models with large context windows and limited API calls, though. The amount of energy wasted on shit code that gets reverted must be tremendous. I hope the AI industry makes true on its promise that it'll solve the whole inefficiency problem because the way things are going now, the industry isn't sustainable.
- simonw 1y agoThe leading models have been very good at this for over a year now. Try copying one your existing C++ header files into GPT-5 or Claude 4 or Gemini 2.5 as an experiment and see how they do.
- shakna 1y agoThey certainly invent new functions whenever I try.
- AirMax98 1y agoI really disagree with this too, especially given the article's next line: > ...You’ll be forever tweaking individual lines of code, asking for a .reduce instead of a .map.filter, bikeshedding function names, and so on. At the same time, you’ll miss the opportunity to guide the AI away from architectural dead ends. I think a good review will often do both, and understand that code happens at the line level and also the structural level. It implies a philosophy of coding that I have seen be incredibly destructive firsthand — committing a bunch of shit that no one on a team understands and no one knows how to reuse.
- tossandthrow 1y ago> for a .reduce instead of a .map.filter... This is distinctly not the api, but an implementation detail. Personally, i can ask colleagues to change function names, rework hierarchy, etc. But leave this exact example be, as it does not have any material difference difference - regardless of my personal preference.
- 000ooo000 1y agoThis blog gets posted often but the content is usually lousy. Lots of specious assertions about the nature of software development that really give off a "I totally have this figured out" vibe. I can't help but feel that anyone who feels so about this young industry that changes so rapidly and is so badly performed at so many places, is yet to summit Mt. Stupid.
- jffhn 1y agoAgreed. A program is made of names, these names are of the utmost importance. For understanding, and also for searchability. I do a lot of code reviews, and one of the main things I ask for, after bug fixes, is renaming things for readers to understand at first read unambiguously and to match the various conventions we use throughout the codebase. Ex: new dev wrote "updateFoo()" for a method converting a domain thing "foo" from its type in layer "a" to its type in layer "b", so I asked him to use "convertFoo_aToB()" instead.
- lapcat 1y agoIf you are good at code review, you will also be good at not using AI agents.
- fhd2 1y agoThis. Having had the pleasure to review the work and fix the bugs of agent jockeys (generally capable developers that fell in love with Claude Code et al), I'm rather sceptical. The code often looks as if they were on mushrooms. They cannot reason about it whatsoever, like they weren't even involved, when I know they weren't completely hands off. I really believe there are people out there that produce good code with these things, but all I've seen so far has been tragic. Luckily, I've witnessed a few snap out of it and care again. Literally looks to me as if they had a substance abuse problem for a couple of months. If you take a critical look at what comes out of contemporary agentic workflows, I think the conclusion must be that it's not there. So yeah, if you're a good reviewer, you would perhaps come to that conclusion much sooner.
- carlmr 1y ago>The code often looks as if they were on mushrooms. They cannot reason about it whatsoever Interesting comparison, why not weed or alcohol?
- fhd2 1y agoNever tried psychedelic mushrooms, so that part is speculation. But no amount of weed or alcohol could get me even close to writing code that unhinged.
- bluefirebrand 1y ago> I really believe there are people out there that produce good code with these things, but all I've seen so far has been tragic I don't believe this at all, because all I've seen so far is tragic I would need to see any evidence of good quality work coming from AI assisted devs before I start to entertain the idea myself. So far all I see is low effort low quality code that the dev themself is unable to reason about
- QuantumNoodle 1y agoCode review is part of the job, but one of the least enjoyable parts. Developers like _writing_ and that gives the most job satisfaction. AI tools are helpful, but inherently increases the amount of code we have to review with more scrutiny than my colleagues because of how unpredictable - yet convincing - it can be. Why did we create tools that do the fun part and increase the non-fun part? Where are the "code-review" agents at?
- cmrdporcupine 1y agoIf you have a paid Copilot membership and a Github project you can request a code review from Copilot. And it doesn't do a terrible job, actually.
- sublinear 1y agoI will second this. I believe code review agents and search summaries are the way forward for coding with LLMs. The ability to ignore AI and focus on solving the problems has little to do with "fun". If anything it leaves a human-auditable trail to review later and hold accountable devs who have gone off the rails and routinely ignored the sometimes genuinely good advice that comes out of AI. If humans don't have to helicopter over developers, that's a much bigger productivity boost than letting AI take the wheel. This is a nuance missed by almost everyone who doesn't write code or care about its quality.
- phito 1y agoBecause the goal of "AI" is not to have fun, it's to solve problems and increase productivity. I have fun programming too, but you have to realize the world isn't optimizing make things more fun.
- fhd2 1y agoI hear you, but without any enjoyment in the process, quality and productivity go down the drain real fast. The Ironies of Automation paper is something I mention a lot, the core thesis is that making humans review / rubber stamp automation reduces their work quality. People just aren't wired to do boring stuff well.
- sublinear 1y ago> If you’re a nitpicky code reviewer, I think you will struggle to use AI tooling effectively. [...] Likewise, if you’re a rubber-stamp code reviewer, you’re probably going to put too much trust in the AI tooling. So in other words, if you are good at code review you are also good enough at writing code that you will be better off writing it yourself for projects you will be responsible for maintaining long term. This is true for almost all of them if you work at a sane place or actually care about your personal projects. Writing code for you is not a chore and you can write it as fluently and quickly as anything else. Your time "using AI" is much better spent filling in the blanks when you're unfamiliar with a certain tool or need to discover a new one. In short, you just need a few google searches a day... just like it ever was. I will admit that modern LLMs have made life easier here. AI summaries on search engines have indeed improved to the point where I almost always get my answer and I no longer get hung up meat-parsing poorly written docs or get nerd-sniped pondering irrelevant information.
- rollulus 1y agoI have received a few LLM produced PRs from peers from adjacent teams, in good faith but not familiar with the project, and they increasingly infuriate me. They were all garbage, but there’s a great asymmetry: it costs my peers nothing to generate them, it costs me precious time to refute them. And what can I do really? Saying “it’s irreparable garbage because the syntax might be right but it’s conceptually nonsense” but that’s not the most constructive take.
- sothatsit 1y agoThis feels like a culture problem. I have seen higher-quality PRs as people use AI to review their work before pushing it. This means less silly typos and obvious small bugs.
- esperent 1y agoYou could use an LLM to give you advice on how to present that take in a more constructive manner. Partially sarcastic but I do personally use LLMs to guide my communication in very limited cases: 1. It's purely business related, and 2. I'm feeling too emotionally invested (or more likely, royally pissed off) and don't trust myself to write in a professional manner, and 3. I genuinely want the message to sound cold, corporate, and unemotional Number 3 would fit you here. These people are not being respectful to you in presenting code for review that respects your time. Why should you take the time to write back personally? It should be noted that this accounts for maybe 5% of my business communications, and I'm careful not to let that number grow.
- walleeee 1y ago> Why should you take the time to write back personally? Because it's 3 sentences, if you want to be way more polite and verbose than necessary. "I will close PRs if they appear to be largely LLM-generated. I am always happy to review something with care and attention if it shows the same qualities. Thanks!" The idea is to get your coworkers to stop sending you AI slop, send them AI slop in retaliation?
- 1y ago
- rsynnott 1y agoThis idea that you can get good results from a bad process as long as you have good quality control seems… dubious, to say the least. “Sure, it’ll produce endless broken nonsense, but as long as someone is checking, it’s fine.” This, generally, doesn’t really work. You see people _try_ it in industry a bit; have a process which produces a high rate of failures, catch them in QA, rework (the US car industry used to be notorious for this). I don’t know of any case where it has really worked out. Imagine that your boss came to you, the tech lead of a small team, and said “okay, instead of having five competent people, your team will now have 25 complete idiots. We expect that their random flailing will sometimes produce stuff that kinda works, and it will be your job to review it all.” Now, you would, of course, think that your boss had gone crazy. No-one would expect this to produce good results. But somehow, stick ‘AI’ on this scenario, and a lot of people start to think “hey, maybe that could work.”
- ChrisMarshallNY 1y agoThat depends. If the engineer, doing the implementation is top-shelf, you can get very good results from a “flawed” process (in quotes, because it’s not actually “bad.” It’s just a process that depends on the engineer being that particular one). Silicon Valley is obsessed with process over people, manifesting “magical thinking” that a “perfect” process eliminates the need for good people. I have found the truth to be in-between. I worked for a company that had overwhelming Process, but that process depended on good people, so it hired top graduates, and invested huge amounts of money and time into training and retention.
- rhetocj23 1y agoSteve Jobs said this decades ago. Its the content that matters, not the process.
- marklubi 1y agoSaid a little more crass/simply: A people hire A people. B people hire C people. The first is phenomenal until someone makes a mistake and brings in a manager or supervisor from the C category that talks the talk but doesn't walk the walk. If you accidentally end up in one that turns out to be the later. It's maddening trying to get anything accomplished if the task involves anyone else. Hire slow, fire fast.
- AshamedCaptain 1y agoNo. The failure conditions of "AI agents" are not even close to classical human mistakes (the only one ones where code review has anything more than an infinitesimal chance to catch). There is absolutely no skill transfer and it is a poor excuse anyway since review was never going to catch anything anyway.
- habibur 1y agoYou review the code and found it broken. Then what? - Rewrite it yourself? - Tell AI to generate it again? — will lead to worse code than the first. - Write the long prompt (like 6 page) even longer and hope it works this time?
- simonw 1y agoIn my experience you tell the AI how to fix it and get better code based on your instructions.
- sothatsit 1y agoGetting AI to produce a bunch of code and then you having to filter through it all is a massive waste of time. The focus should be on getting AI to produce better code in the first place (e.g., using detailed plans), rather than on the volume of code you can produce... I have only had real advantages with AI for helping me plan changes, and for it helping me to review my code. Getting it to write code for me has been somewhat helpful, but only for simple tedious changes or first drafts. But it is definitely not something I want to leverage by getting AI to produce more and more code that I then have to filter through and review. No thank you. I feel like this is really the wrong focus for implementing AI into your workflows.
- fleischhauf 1y agocan I ask what language you are using AI for, there is also a difference in performance for AI in different languages
- sothatsit 1y agoTypeScript with NextJS. I've also used AI tools with C and Zig, and AI is much better at writing TS. But even though TS works much better, it's still not that great. This is largely because the quality of the code that AI writes is not good enough, so then I have to spend a decent chunk of time fixing it. Everyone I know trying to use AI in large codebases has had similar experiences. AI is not good enough at following the rules of your codebase yet (i.e., following structure, code style, library usage, re-using code, refactoring, etc...). This makes it far less useful for writing code changes and additions. It can still be useful for small changes, or for writing first drafts of functions/classes/interfaces, but for more meaningful changes it often fails. That is why I believe that right now, if you want to maintain a large codebase, and maintain a high bar for quality, AI tools are just not good enough at writing most code for you yet. The solution to this is not to get AI to write even more code for you to review and throw out and iterate upon in a frustrating cycle. Instead, I believe it is to notice where AI is helpful and focus on those use-cases, and avoid it when it is not. That said, AI labs seem to be focusing a lot of effort on improving AI for coding right now, so I expect a lot of progress will be made on these issues in the next few years.
- simianparrot 1y agoOr: As long as you have a good editor with endless time, a thousand monkeys with typewriters will reproduce Shakespeare.
- praptak 1y agoUsername checks out.
- brap 1y agoMy process is basically 1. Give it requirements 2. Tell it to ask me clarifying questions 3. When no more questions, ask it to explain the requirements back to me in a formal PRD 4. I criticize it 5. Tell it to come up with 2 alternative high level designs 6. I pick one and criticize it 7. Tell it to come up with 2 alternative detailed TODO lists 8. I pick one and criticize it 9. Tell it to come up with 2 alternative implementations of one of the TODOs 10. I pick one and criticize it 11. Back to 9 I usually “snapshot” outputs along the way and return to them to reduce useless context. This is what produces the most decent results for me, which aren’t spectacular but at the very least can be a baseline for my own implementation. It’s very time consuming and 80% of the time I end up wondering if it would’ve been quicker to just do it all by myself right from the start.
- stavros 1y agoI have a similar, though not as detailed, process. I do the same as you up to the PRD, then give it the PRD and tell it the high level architecture, and ask it to implement components how I want them. It's still time-consuming, and it probably would be faster for me to do it myself, but I can't be bothered manually writing lines of code any more. I maybe should switch to writing code with the LLM function by function, though.
- bluefirebrand 1y ago> but I can't be bothered manually writing lines of code any more. I maybe should switch to writing code with the LLM function by function, though. Maybe you should consider a change of career :/
- stavros 1y agoWhy?
- scuff3d 1y agoThat's like a chef saying they can't be bothered to cook...
- notachatbot123 1y agoI love doing code review for colleagues since I know that it bolsters our shared knowledge, experience and standards. Code review for an external, stubborn, uncooperative AI? No thanks, that sounds like burnout.
- deleted 1y ago[deleted]
- swaptr 1y agoAI-generated code can be useful in the early stages of a project, but it raises concerns in mature ones. Recently, a 280kloc+ Postgres parser was merged into Multigres (https://github.com/multigres/multigres/pull/109 https://github.com/multigres/multigres/pull/109) with no public code review. In open source, this is worrying. Many people rely on these projects for learning and reference. Without proper review, AI-generated code weakens their value as teaching tools, and more importantly the trust in pulling as dependencies. Code review isn’t just about bugs, it’s how contributors learn, understand design choices, and build shared knowledge. The issue isn’t speed of building software (although corporations may seem to disagree), but how knowledge is passed on. Edit: Reference to the time it took to open the PR: https://www.linkedin.com/posts/sougou_the-largest-multigres-pr-ever-submitted-activity-7374107384559923200-oid- https://www.linkedin.com/posts/sougou_the-largest-multigres-...
- sougou 1y agoI oversaw this work, and I'm open to feedback on how things can be improved. There are some factors that make this particular situation different: This was an LLM assisted translation of the C parser from Postgres, not something from the ground up. For work of this magnitude, you cannot review line by line. The only thing we could do was to establish a process to ensure correctness. We did control the process carefully. It was a daily toil. This is why it took two months. We've ported most of the tests from Postgres. Enough to be confident that it works correctly. Also, we are in the early stages for Multigres. We intend to do more bulk copies and bulk translations like this from other projects, especially Vitess. We'll incorporate any possible improvements here. The author is working on a blog post explaining the entire process and its pitfalls. Please be on the lookout. I was personally amazed at how much we could achieve using LLM. Of course, this wouldn't have been possible without a certain level of skill. This person exceeds all expectations listed here: https://github.com/multigres/multigres/discussions/78 https://github.com/multigres/multigres/discussions/78.
- wg002 1y ago"We intend to do more bulk copies and bulk translations like this from other projects" Supabase’s playbook is to replicate existing products and open source projects, release them under open source, and monetize the adoption. They’ve repeated this approach across multiple offerings. With AI, the replication process becomes even faster, though it risks producing low-quality imitations that alienate the broader community and people will resent the stealing of their work.
- jmull 1y ago> In my view, the best code review is structural. It brings in context from parts of the codebase that the diff didn’t mention. That may be true for AI code. But it would be pretty terrible for human-written code to bring this up after the code is written, wasting hours/days effort for lack of a little up-front communication on design. AI makes routine code generation cheap -- only seconds/minutes and cents are being wasted -- but you essentially still need that design session.
- harimau777 1y agoI think that I review code much differently than the author. When I'm reviewing code, my assumption is that the person writing it has already verified that it works. I am primarily looking for readability and code smells. In an ideal world I'd probably be looking more at the actual logic of the code. However, everywhere I've worked it's a full time job just despirately trying to fight ballooning complexity from people who prioritize quick turn around over quality code.
- conartist6 1y agoI am good at code review, sure, but I don't like doing it. It's about as strong an engineering technique as coding at a whiteboard. I know I'm at a tiny fraction of my potential without debugging tools and for that reason code review on github is usually a waste of my time. I'll just write code thanks and I'll move the needle on quality by developing. As a reviewer I'll scan for smells but I assume that you too would be most effective if I left you make and clean up your own messes so long as they aren't egregious
- rco8786 1y agoUnfortunately code review is like the least fun part of software engineering.
- HarHarVeryFunny 1y agoThe title of this article seems way too glib. Code review isn't the same as design review, nor are these the only type of things (coding and design) that someone may be trying to use AI for. If you are going to use AI, and catch it's mistakes, then you need to have expertise in whatever it is you are using the AI for. Even if we limit the discussion just to coding, then being a good code reviewer isn't enough - you'd need to have skill at whatever you are asking the AI to do. One of the valuable things AI can do is help you code using languages and frameworks you are not familiar with, which then of course means you are not going to be competent to review the output, other than in most generic fashion. A bit off topic, but it's weird to me to see the term "coding" make a comeback in this AI/LLM era. I guess it is useful as a way to describe what AI is good at - coding vs more general software developer, but how many companies nowadays hire coders as opposed to software developers (I know it used to be a thing with some big companies like IBM)? Rather than compartmentalized roles, it seems the direction nowadays is more expecting developers to be able to do everything from business analysis and helping develop requirements, to architecture/design and then full-stack development, and subsequent production support.
- karmakaze 1y agoSeems so. > Using AI agents correctly is a process of reviewing code. [...] > Why is that? Large language models are good at producing a lot of code, but they don’t yet have the depth of judgement of a competent software engineer. Left unsupervised, they will spend a lot of time committing to bad design decisions. Obviously you want to make course corrections sooner than later. Same as I would do with less experienced devs, talk through the high level operations, then the design/composition. Reviewing a large volume of unguided code is like waiting for 100k tokens to be written only to correct the premise in the first 100 and start over.
- scuff3d 1y agoMy official title is "Software Engineer", in the last five years I have.. 1. Stood up and managed my own Kubernetes clusters for my team 2. Docker, just so so much Docker 3. Developed CI/CD pipelines 4. Done more integration and integration testing then I care to think about 5. Written god knows how many requirements and produced and endless stream of diagrams and graphs for systems engineering teams 6. Don't a bunch of random IT crap because our infrastructure team can't be bothered 7. Wrote some code once in a while
- jaredcwhite 1y agoSorry, this is not the profession of programming, and people in the near future will be looking back at this era and laughing their asses off. But not me, because I will never touch an agentic tool. And believe me, I have a big smile on my face. Life is good! =D
- ath3nd 1y ago[dead]
- moltar 1y agoI’ve been thinking about this a lot lately. What’s the best way to review AI code? I wish I had a local, GitHub PR review-like experience where I can leave comments for the agent.
- em-bee 1y agocode review can be almost as much effort as writing the code, especially when the code is not up to the expectations of the reviewer. this is fine, because you want two people (the original author, and the reviewer) on the code. when reviewing AI code, not only will the effort needed by the reviewer increase, you also lose the second person (the author) looking at the code, because AI can't do that. it can produce code but not reason about or reflect on it like humans can.
- eowyn 1y agoWhat does this mean for juniors? A few companies are now introducing expectations that all engineers will use coding agents including juniors and grads. If they haven't yet learnt what good looks like through experience how are they going to review code produced by AI agents?
- dawnerd 1y agoAs someone that basically does code reviews for a living, last thing I want to do is code review agents. I want to reduce how much review I’m doing, not hand hold some ai agents.
- dearilos 1y agoI'm building something to do exactly that - just reduce and automate the boring parts of code review like enforcing standards.
- tdeck 1y agoI think I'm good at code review, but we've all seen parts of the codebase where it's written by one teammate with specific domain knowledge and your option is to approve something you don't fully understand or to learn the background necessary to understand it. In my experience, not having to learn the background is the biggest time saver provided by LLM coding (e.g. not having to read through API docs or confirm details of a file format or understand some algorithm). So in a way I feel like there is a fundamental tension.
- insane_dreamer 1y agoTwo observations: - If I had to iterate as much with a Jr dev as CC on not highly difficult stuff ("of course, I'll just do X!" then X doesn't work, then "of course, the answer is Y!" then Y doesn't work, etc.) I probably would have fired them by now or just say "never mind, I'll do it myself" . - On the other hand a Jr dev will (hopefully) learn as they go, get better each time, so a month from now they're not making the same mistakes. An LLM can't learn so until there's a new model they keep making the same mistakes (yes, within a session they can learn -- if the session doesn't get too long -- but not across sessions). Also, the Jr dev can test their solution (which may require more than just running unit tests) and iterate on it so that they only come to me when it works and/or they're stuck. Just yesterday, on a rather simple matter, I wasted so much time telling the LLM "that didn't work, try again".
- vrighter 1y agoIf I am good at the most boring part of my job, I get to do that and only that from here on out? No thank you. Also, the article is wrong, it's always better for a bug not to be in there in the first place, than to be there and possibly be missed.
- egberts1 1y agoAsking AI to stay true to my requested parameters is hard, THEY ALL DRIFT AWAY, RANDOMLY When working on nftables syntax highlighters, I have 230 tokens, 2,500 state, and 50,000+ state transitions. Some firm guidelines given to AI agents are: 1. Fully-deterministic LL(1) full syntax tree. 2. No use of Vim 'syntax keyword' statement 3. Use long group names in snake_case whose naming starts with 'nft_' prefix (avoids collision with other Vim namespaces) 4. For parts of the group names, use only nftables/src/parser_bison.y semantic action and token names as-is. 5. For each traversal down the syntax tree, append that non-terminal node name from parser_bison.y to its group names before using it. With those 5 "simple" user-requested requirements, all AI agents drift away from at least each of the rules at seemingly random interval. At the moment, it is dubiois to even trust the bit-length of each packet field. Never mind their inability to construct a simple Vimscript. I use AI agents mainly as documentation. On the bright side, they are getting good at breaking down 'rule', 'chain_block stmt', and 'map_stmt_expr' (that '.' period we see at chaining header expressions together; just use the quoted words and paste in one of your nft rule statements.
- legacynl 1y agoIm a dev (who knows nothing about nftables) and I don't understand your instructions. I think maybe you could improve your situation by formulating them as "when creating new groupnames use the semantic actions and token names as defined in parser_bison.y" I.e. with if conditions so that the correct rules apply to the correct situations. Because your rules are written as if to apply to every line of code, it might unnecessarily try to incorporate context even when it's not applicable.