12 ms·
There is an AI code review bubble
- seanmccann 8mo agoAs Claude Code (and Opus) improves, Greptile is finding fewer issues in my code reviews.
- personjerry 8mo agoI don't really understand how this differentiates against the competition. > Independence Any "agent" running against code review instead of code generation is "independent"? > Autonomy Most other code review tools can also be automated and integrated. > Loops You can also ping other code review tools for more reviews... I feel like this article actually works against you by presenting the problem and inadequately solving them.
- dakshgupta 8mo ago> Independence It is, but when a model/harness/tools/system prompts are the same/similar in the generator and reviewer fail in similar ways. Question: Would you trust a Cursor review of Claude-written code more, less, or the same as a Cursor review of Cursor-written code? > Autonomy Plenty of tools have invested heavily in AI-assisted review - creating great UIs to help human reviewers understand and check diffs. Our view is that code validation will be completely autonomous in the medium term, and so our system is designed to make all human intervention optional. This is possibly a unpopular opinion, and we respect the camp that might say people will always review AI-generated code. It's just not the future we want for this profession, nor the one we predict. > Loops You can invest in UX and tooling that makes this easier or harder. Our first step towards making this easier is a native Claude Code plugin in the `/plugins` command that let's Claude code do a plan, write, commit, get review comments, plan, write loop.
- liamconnell 8mo ago> It is, but when a model/harness/tools/system prompts are the same/similar in the generator and reviewer fail in similar ways. Is there empirical evidence for that? Where is it on an epistemic meter between (1) “it sounds good when I say it”, and (10) “someone ran evaluation and got significant support.” “Vibes” (2/3 on scale) are ok, just honestly curious.
- sdenton4 8mo agoIndependence is ridiculous - the underlying llm models are too similar on their training days and methodologies to be anything like independent. Trying different models may somewhat reduce the dependency, but all have read stack overflow, Reddit, and GitHub in their training. It might be an interesting time to double down on automatically building and checking deterministic models of code which were previously too much of a pain to bother with. Eg, adding type checking to lazy python code. These types of checks really are model independent, and using agents to build and manage them might bring a lot of value.
- rohansood15 8mo ago> Would you trust a Cursor review of Claude-written code more, less, or the same as a Cursor review of Cursor-written code? You're assuming models/prompts insist on a previous iteration of their work being right. They don't. Models try to follow instructions, so if you ask them to find issues, they will. 'Trust' is a human problem, not a model/harness problem. > Our view is that code validation will be completely autonomous in the medium term. If reviews are going to be autonomous, they'd be part of the coding agent. Nobody would see it as an independent activity, you mentioned above. > Our first step towards making this easier is a native Claude Code plugin. Claude can review code based on a specific set of instructions/context in an MD file. An additional plugin is unnecessary. My view is that to operate in this space, you gotta build a coding agent or get acquired by one. The writing was on the wall a year ago.
- geooff_ 8mo agoThis article has a catchy headline, but there's really no content to it. This is content marketing without content. It seems like every week on Hacker News, there's a dozen of these. All seemingly code reviewers, too. Keep it to LinkedIn.
- MichaelRo 8mo ago[flagged]
- trjordan 8mo ago1. I absolutely agree there's a bubble. Everybody is shipping a code review agent. 2. What on earth is this defense of their product? I could see so many arguments for why their code reviewer is the best, and this contains none of them. More broadly, though, if you've gotten to the point where you're relying on AI code review to catch bugs, you've lost the plot. The point of a PR is to share knowledge and to catch structural gaps. Bug-finding is a bonus. Catching bugs, automated self-review, structuring your code to be sensible: that's _your_ job. Write the code to be as sensible as possible, either by yourself or with an AI. Get the review because you work on a team, not in a vacuum.
- dakshgupta 8mo ago2. There is plenty of evidence for this elsewhere on the site, and we do encourage people to try it because like with a lot of AI tools, YMMV. You're totally right that PR reviews go a lot farther than catching issues and enforcing standard. Knowledge sharing is a very important part of it. However, there are processes you can create to enable better knowledge sharing and let AI handle the issue-catching (maybe not fully yet, but in time). Blocking code from merging because knowledge isn't shared yet seems unnecessary.
- ahmadyan 8mo ago> 2. What on earth is this defense of their product? i think the distribution channel is the only defensive moat in low-to-mid-complexity fast-to-implement features like code-review agents. So in case of linear and cursor-bugbot it make a lot of sense. I wonder when Github/Gitlab/Atlassian or Xcode will release their own review agent.
- lenerdenator 8mo ago> More broadly, though, if you've gotten to the point where you're relying on AI code review to catch bugs, you've lost the plot. > The point of a PR is to share knowledge and to catch structural gaps. Well, it was to share knowledge and to catch structural gaps. Now you have an idea, for better or for worse, that software needs to be developed AI-first. That's great for the creation of new code but as we all know, it's almost guaranteed that you'll get some bad output from the AI that you used to generate the code, and since it can generate code very fast, you have a lot of it to go through, especially if you're working on a monorepo that wasn't architected particularly well when it was written years ago. PRs seem like an almost natural place to do this. The alternative is the industry finding a more appropriate place to do this sort of thing in the SDLC, which is gonna take time, seeing as how agentic loop software development is so new.
- candiddevmike 8mo agoNone of these tools perform particularly well and all lack context to actually provide a meaningful review beyond what a linter would find, IMO. The SOTA isn't capable of using a code diff as a jumping off point. Also the system prompts for some of them are kinda funny in a hopelessly naive aspirational way. We should all aspire to live and breathe the code review system prompt on a daily basis.
- dakshgupta 8mo agoI agree that none perform _super_ well. I would argue they go far beyond linters now, which was perhaps not true even nine months ago. To the degree you consider this to be evidence, in the last 7 days, the authors of a PR has replied to a Greptile comment with "great catch", "good catch", etc. 9,078 times.
- onedognight 8mo agoI fully agree. Claude’s review comments have been 50% useful, which is great. For comparison I have almost never found a useful TeamScale comment (classic static analyzer). Even more important, half of Claude’s good finds are orthogonal to those found by other human reviewers on our team. I.e. it points out things human reviewers miss consistently and v.v.
- Sharlin 8mo agoTBH that sounds like TeamScale just has too verbose default settings. On the other hand, people generally find almost all of the lints in Clippy's [1] default set useful, but if you enable "pedantic" lints, the signal-to-noise ratio starts getting worse – those generally require a more fine-grained setup, disabling and enabling individual lints to suit your needs. [1] https://doc.rust-lang.org/stable/clippy/ https://doc.rust-lang.org/stable/clippy/
- tadfisher 8mo agoNot trying to sidetrack, but a figure like that is data, not evidence. At the very minimum you need context which allows for interpretation; 9,078 positive author comments would be less impressive if Greptile made 1,000,000 comments in that time period, for example.
- quanwinn 8mo agoI liked that the post is self-aware that it's promoting its own product. But the writing seemed more focus on the philosophy behind code reviews and the impact of AI, and less on the mechanics of how greptile differs from competitors. I was hoping to see more on the latter.
- dakshgupta 8mo agoThanks! We go over that on many other pages. Here are some: https://www.greptile.com/benchmarks https://www.greptile.com/benchmarks https://www.greptile.com/greptile-vs-coderabbit https://www.greptile.com/greptile-vs-coderabbit https://www.greptile.com/greptile-vs-bugbot https://www.greptile.com/greptile-vs-bugbot
- ahmadyan 8mo agoProblem with Code Review is it is quite straightforward to just prompt it, and the frontier models, whether Opus or GPT5.2Codex do a great job at code-reviews. I don't need second subscription or API call when the first one i already have and focus on integration works well out of the box. In our case, agentastic.dev, we just baked the code-review right into our IDE. It just packages the diff for the agent, with some prompt, and sends it out to different agent choice (whether claude, codex) in parallel. The reason our users like it so much is because they don't need to pay extra for code-review anymore. Hard to beat free add-on, and cherry on top is you don't need to read a freaking poems.
- zurfer 8mo agowe use codex review. it's working really well for us. but i don't agree that it's straightforward. moving the number of bugs catched and signal to noise ratio a few percentage points is a compounding advantage. it's a valuable problem to solve, amplified by the fact that ai coding produces much more code. that being said, i think it's damn hard to compete with openai or anthropic directly on a core product offering in the long run. they know that it's an important problem and will invest accordingly.
- sastraxi 8mo agoContrary to some of the other anecdotes in this thread, I've found automated code review to discover some tricky stuff that humans missed. We use https://www.cubic.dev/ https://www.cubic.dev/
- pomarie 8mo agoFounder of cubic here, thanks for the shoutout!
- aurareturn 8mo agoBefore I push any code, I always ask 2 different frontier LLMs to review the changes for any potential issues. Saved my ass a few times before pushing to production.
- taude 8mo agoIt's not terribly hard to write a Copilot GHA that does this yourself for your specific teams needs. Not sure why you'd been to bring a vendor on for this.... What do the vendors provide? I looked at a couple which were pretty snazzy at first glance, but now that I know more about how copilot agents work and such, I'm pretty sure in a few hours, I could have the foundation for my team to build on that would take care of a lot of our PR review needs....
- jackconsidine 8mo ago> Only once would you have X write a PR, then have X approve and merge it to realize the absurdity of what you just did. I get the idea. I'll still throw out that having a single X go through the full workflow could still be useful in that there's an audit log, undo features (reverting a PR), notifications what have you. It's not equivalent to "human writes ticket, code deployed live" for that reason
- TuringTest 8mo ago>A human rubber-stamping code being validated by a super intelligent machine is the equivalent of a human sitting silently in the driver's seat of a self-driving car, "supervising". So, absolutely necessary and essential? In order to get the machine out of trouble when the unavoidable strange situation happens that didn't appear during training, and requires some judgement based on ethics or logical reasoning. For that case, you need a human in charge.
- pavan_panto 8mo ago[dead]
- pawelduda 8mo agoGood code reviews are part of team's culture and it's hard to just patch it with an agent. With millions of tools it will be arms race between which one is louder about as many things as possible because: - it will have higher chance at convincing the author that the issue was important by throwing more darts - something that a human wouldn't do because it takes real mental effort to go through an authentic review, - it will sometimes find real big issue which reinforces the bias that it's useful - there will always be tendency towards more feedback (not higher quality) because if it's too silent, is it even doing anything? So I believe it will just add more round of back and forth of prompting between more people, but not sure if net positive Plus PRs are a good reality check if your code makes sense, when another person reviews it. A final safeguard before maintainability miss, or a disaster waiting to be deployed.
- deleted 8mo ago[deleted]
- pnathan 8mo agoClaude code's code review is _sufficient_ imo. still need HITL, but the human is shifted right and can do other things rather than grinding through fiddly details.
- themafia 8mo ago> Unfortunately, code review performance is ephemeral and subjective > Today's agents are better than the median human code reviewer Which is it? You cannot have it both ways.
- maxverse 8mo ago> Today's agents are better than the median human code reviewer "...at catching issues and enforcing standards, and they're only getting better". I took this to mean what good code review is is subjective. But if you clearly define standards and patterns for your code, your linter/automated tools/ AI code reviewer will always catch more than humans.
- mohsen1 8mo agoSo far I've been pretty happy with Greptile. Tried Copilot and Cubic.dev but landed on Greptile
- disillusionist 8mo agoMy company just finished a several week review period of Greptile. Devs were split over the usefulness of the tool (compared to our current solution, Cursor). While Greptile did occasionally offer better insights than Cursor, it also exhibited strange behavior such as entirely overwriting PR descriptions with its own text and occasionally arguing with itself in the comments. In the end we decided to NOT purchase Greptile as there were enough "not quite there" issues that made it more trouble than worthwhile. I am certain, though, that the Greptile team will resolve all those problems and I wish them the best of luck!
- dcreater 8mo agoReminder that this comes from from the founder that got rightly lambasted for his comments about work life balance and then doubled down when called out.
- dcreater 8mo agoThere is an AI bubble. Can drop the extra words
- maxverse 8mo agoMaybe I'm buying into the cool-aid, but I actually really liked the self-aware tone of this post. > Based on our benchmarks, we are uniquely good at catching bugs. However, if all company blogs are to be trusted, this is something we have in common with every other AI code review product. One just has to try a few, and pick the one that feels the best.
- h1fra 8mo agoone more ai code review please, I promise it will fix everything this time, please just one more
- cbovis 8mo agoI've also noticed this explosion of code review tools and felt that there's some misplaced focus going on for companies. Two that stood out to me are Sentry and Vercel. Both have released code review tools recently and both feel misplaced. I can definitely see why they thought they could expand with that type of product offering but I just don't see a benefit over their competition. We have GH copilot natively available on all our PRs, it does a great job, integrates very well with the PR comment system, and is cheap (free with our current usage patterns). GH and other source control services are well placed to have first-class code review functionality baked into their PR tooling. It's not really clear to me what Sentry/Vercel are offering beyond what copilot does and in my brief testing of them didn't see noticeable difference in quality or DX. Feels like they're fighting an uphill battle from day one with the product choice and are ultimately limited on DX by how deeply GH and other source control service allow them to integrate. What I would love to see from Vercel, which they feel very well placed to offer, is AI powered QA. They already control the preview environments being deployed to for each PR, they have a feedback system in place with their Vercel toolbar comments, so they "just" need to tie those together with an agentic QA system. A much loftier goal of course but a differentiator and something I'm sure a lot of teams would pay top dollar for if it works well.
- heliumtera 8mo agoNo shit. What is the point of using an llm model to review code produced by an llm model? Code review pressupose a different perspective, which no platform can offer at the moment because they are just as sophisticated as the model they wrap. Claude generated the code, and Claude was asked if the code was good enough, and now you want to be in the middle to ask Claude again but with more emphasis, I guess? If I want more emphasis I can ask Claude myself. Or Qwen. I can't even begin to understand this rationale.
- kaishin 8mo agoWe used Greptile where I work and it was so bad we decided to switch to Claude. And even Claude isn’t nearly as good at reviewing as an experienced programmer with domain knowledge.
- cmrdporcupine 8mo agoMy experience is that Claude or others are good at pointing out things I will want to look at and then I can go review more thoroughly. So it's helped to some degree. But like everything else with it, it tries to do too much. What I want is a review "wizard" agent -- something that identifies the pieces I should look at, and takes me through them diff by diff asking me to read them, while offering its commentary ("this appears to be XX....") and letting me make my own.
- rrhjm53270 8mo agoWhy not let AI write the code and then have it reviewed by humans? If you use AI to review my code, then you can't stop me from using another AI to refute it: this only foreshadows the beginning of internal friction.
- tfarias 8mo agoMy experience with code review tools has been dreadful. In most cases I can remember the reviews are inaccurate, "you are absolutely right" sycophantic garbage, or missing the big picture. The worst feature of all is the "PR summary" which is usually pure slop lacking the context around why a PR was made. Thankfully that can be turned off. I have to be fair and say that yes, occasionally, some bug slips past the humans and is caught by the robot. But these bugs are usually also caught by automated unit/integration tests or by linters. All in all, you have to balance the occasional bug with all the time lost "reviewing the code review" to make sure the robot didn't just hallucinate something.
- rushingcreek 8mo agoGreptile is a great product and I hope you succeed. However, I disagree that independence is a competitive advantage. If it’s true that having a “firewall” between the coding agent and review agent leads to better code, I don’t see why a company like Cursor can’t create full independence between their coding and review products but still bundle them together for distribution. Furthermore, there might well be benefits to not being fully independent. Imagine if an external auditor was brought in to review every decision made inside your company. There would likely be many things they simply don’t understand. Many decisions in code might seem irrational to an external standalone entity but make sense in the broader context of the organization’s goals. In this sense, I’m concerned that fully independent code review might miss the forest for the trees relative to a bundled product. Again, I’m rooting for you guys. But I think this is food for thought.
- cmrdporcupine 8mo ago"While some other products have built out great UIs for humans to review code in an AI-assisted paradigm, we have chosen to build for what we consider to be an inevitable future - one where code validation requires vanishingly little human participation." Ok good, now I know not to bother reading through any of their marketing literature, because while the product at first interested me, now I know it's exactly not what I want for my team. The actual "bubble" we have right now is a situation where people can produce and publish code they don't understand, and where engineers working on a system no longer are forced to reckon with and learn the intricacies of their system, and even senior engineers don't gain literacy into the very thing they're working on, and so are somewhat powerless to assess quality and deal with crisis when it hits. The agentic coding tools and review tools I want my team (and myself) to have access to are ones that ones that force an explicit knowledge interview & acquisition process during authoring and involve the engineer more intricately in the whole flow. What we got instead with claude code & friends is a thing way too eager to take over the whole thing. And while it can produce some good results it doesn't produce understandable systems. To be clear, it's been a long time since writing code has been the hard part of the job? in many many domains. The hard part is systems & architecture and while these tools can help with that, there's nothing more potentially terrifying pthan a team full of people who have agentically produced a codebase that they cannot holistically understand the nuances of. So, yeah, I want review tools for that scenario. Since these people have marketed themselves off the table... what is out there?
- jacobegold 8mo agoYep. We see this future and are working on exactly what you're talking about (Graphite)
- cmrdporcupine 8mo agoYou just completely contradicted yourself then.
- jacobegold 8mo agoNot sure how? Meant this: > The agentic coding tools and review tools I want my team (and myself) to have access to are ones that ones that force an explicit knowledge interview & acquisition process during authoring and involve the engineer more intricately in the whole flow.
- zmmmmm 8mo agoMy experience with using AI tools for code review is that they do find critical bugs (from my retrospective analysis, maybe 80% of the time), but the signal to noise ratio is poor. It's really hard to get it not to tell you 20 highly speculative reasons why the code is problematic along with the one critical error. And in almost all cases, sufficient human attention would also have identified the critical bug - so human attention is the primary bottleneck here. Thus poor signal to noise ratio isn't a side issue, it's one of the core issues. As a result, I'm mostly using this selectively so far, and I wouldn't want it turned on by default for every PR.
- Quarrelsome 8mo ago> but the signal to noise ratio is poor Nail on the head. Every time I've seen it applied, its awful at this. However this is the one thing I loathe in human reviews as well, where people are leaving twenty comments about naming and then the actual FUNCTIONAL issue is just inside all of that mess. A good code reviewer knows how to just drop all the things that irk them and hyperfocus on what matters, if there's a functional issue with the code. I wonder if AI is ever gonna be able to conquer that one as its quite nuanced. If they do though, then I feel the industry as it is today, is kinda toast for a lot of developers, because outside of agency, this is the one thing we were sorta holding out on being not very automatable.
- zenolijo 8mo agoNaming comments can be very useful in code that gets read by a lot of people. It can make the process of understanding the code much quicker. On the other hand, if it's less important code or the renaming is not clearly an improvement it can be quite useless. But I've met some developers who has the opinion of reviews as pointless and just say "this works, just approve it already" which can be very frustrating when it's a codebase with a lot of collaboration.
- Quarrelsome 8mo ago> Naming comments can be very useful in code that gets read by a lot of people. It can make the process of understanding the code much quicker. yes but it can be severely diminishing returns. Like lets step back a second and ask ourselves if: var itemCount = items.Count; vs var numberOfItems = items.Count; is ever worth spending the time discussing, versus how much of a soft improvement it makes to the code base. I've literally been in a meeting room with three other senior engineers killing 30 minutes discussing this and I just think that's a complete waste of time. They're not wrong, the latter is clearer, but if you have a PR that improves the repo and you're holding it back because of something like this, then I don't think you have your priorities straight.
- Fervicus 8mo agoLLMs writing code, and then LLMs reviewing the code. And when customers run into a problem with the buggy slop you just churned out, they can talk to a LLM chat bot. Isn't it just swell?
- dullcrisp 8mo agoJust let the support chat bot submit, review, and deploy code changes and there are no longer any customer problems!
- segmondy 8mo agoIf you give LLM a hammer everything looks like a nail, you give it a saw everything looks like wood. You ask LLM to find issues, it will find "issues" At the end of the day, you will have to fix those issues, if you decide to have another LLM fix those issues, by the time you are done with that cycle, you are going to end up with code that will be thoroughly over engineered.
- heliumtera 8mo agoIf by engineering you mean doing whatever vibes you feel, than yeah, over engineering. If by engineering you mean using the engineering design process than it would not be engineered at all, let alone over engineered.
- randusername 8mo agoThis article surprised me. I would have expected it would be about how _human_ code review is unsustainable in the face of AI-enhanced velocity. I would be interested to hear of some specific use-cases for LLMs in code review. With static analysis, tests, and formatters I thought code review was mostly interpersonal at this point. Mentorship, ensuring a chain of liability in approvals, negotiating comfort levels among peers with the shared responsibility of maintaining the code, that kind of thing.
- raincole 8mo agoI still think any business that is based on someone else's model is worthless. I know I'm sounding like the 'dropbox is just FTP' guy, but it really feels like that any good idea will just be copied by OpenAI and Anthropic. If AI code review is proven a good idea is there any reason to expect Codex or Claude Code to not implement some commands to do code review?
- bluGill 8mo agoThe shakiest business model is one where you have no competition - if nobody else had the idea already: you are probably wrong - they did but it was a bad idea so they failed. The real question is how can you compete. There are lots of answers here, but something new and good is rare.
- sthuck 8mo agoVery very strictly speaking relying on models in it's essence is not the problem I think. There is enough "meat" there you can build a nice small profitable company. Those tools are better than vanilla agents by dedicating expensive human time on evaluating and fine tuning models. You can also build various integration, management and reporting features to add value. If you freeze model progress today, or 12 months ago when most of those companies started, it's a viable business I think. But any gains you make on the first part will be lost to newer models, and the 2nd part is not as valuable when llms allow people to build fairly complicated features quickly. I don't if worthless but all those companies have very limited time to gather customers and at least make themselves valuable for an acquisition
- vrighter 8mo agoYou cannot be profitable unless the service you rely on is also profitable. You might make some profits during their honeymoon period, but then they will squeeze you by the balls pretty soon and force you to enshittify as well. WinRAR is more profitable than OpenAI...
- tokioyoyo 8mo agoWe do our review through Claude github actions. Works well.
- the__alchemist 8mo agoWe have Code Rabbit at work, and it's made PRs unreadable. The Bun pollutes the comments and code diffs with noise.
- kxbnb 8mo ago[dead]
- alittletooraph2 8mo agoEither become a platform or get swallowed up by one (e.g. Cursor acquiring Graphite to become more of a platform). Trying to prove out that your code review agent is marginally better than others when the capability is being included in every single solution is a losing strategy. They can just give the capability away for free. Also, the idea that code review will scale dramatically in importance as more code is written by agents is not new.
- sidgarimella 8mo agowhere we draw the line on agent "identity" when the models being orchestrated are generally the same 3 frontier intelligences is an interesting question indeed I would think this idea of creating a third-party to verify things likely centers more around liability/safety cover for a steroidal increase in velocity (i.e. --dangerously-skip-permissions) rather than anything particularly pragmatic or technical (but still poised to capture a ton of value)
- dzonga 8mo agoor stick with known frameworks documented - so you don't have to pay for this nonsense since they're likely telling you things you know if you test and write your own code. oh - writing your own code is a thing of the past - a.i writes, a.i then finds bugs
- simbleau 8mo agoAfter testing several bots in our org, specifically Devin, Graphite, and Cursor, I’ve noticed Cursor is the better bug bot out there right now.
- coopykins 8mo agoSame here, tested a bunch and cursor has been given little noise and usually decent suggestions. In this case its on a react app, so other projects might not find it as good.
- Yizahi 8mo ago> This might seem far-fetched but the counterfactual is Kafkaesque. > As the proprietors of an, er, AI code review tool suddenly beset by an avalanche of competition, we're asking ourselves: what makes us different? > Human engineers should be focused only on two things - coming up with brilliant ideas for what should exist, and expressing their vision and taste to agents that do the cruft of turning it all into clean, performant code. > If there is ambiguity at any point, the agents Slack the human to clarify. Was this LLM advertisement generated by an LLM? Feels so at least.
- lifetimerubyist 8mo agoHaven’t used a single one that was any good. Basically a 50/50 crapshoot if what they are saying makes any sense at all, let alone it being considered “good” comments. Basically no different than random chance.
- ex-aws-dude 8mo agoI find a lot of times with co-pilot it calls out issues where if the AI had more context of the whole codebase it would realize that scenario can’t actually occur. Or it won’t understand some invariant that you know but is not explicit anywhere
- nickitolas 8mo ago> In addition, success is generally pretty well-defined. Everyone wants correct, performant, bug-free, secure code. I feel like these are often not well defined? "Its not a bug it's a feature", "premature optimization is the root of all evil", etc In different contexts, "performant enough" means different things. Similarly, many times I've seen different teams within a company have differing opinions on "correctness"
- 0xbadcafebee 8mo agoHot take: Code review is an anti-pattern. We spend a ton of time looking at the code and blocking merges, and the end result is still full of bugs. AI code review only provides a minor improvement. The only reason we do code review at all is humans don't trust that the code works. Know another way to tell if code works? Running it. If our code is so utterly inconceivable that we can't make tests that can accurately assess if the code works, then either our code design is too complicated, or our tests suck. OTOH, if the reason you're doing code review is to ensure the code "is beautiful" or "is maintainable", again, this is a human concern; the AI doesn't care. In fact, it's becoming apparent that it's easier to replace entire sections of code with new AI generated code than to edit it.
- insin 8mo agoTests can't tell you if the design of the code is fit for purpose, or about requirements you completely missed or punted on, or that a core new piece that's going to be built upon next is barely-coherent, poorly-performing slop that "works" but is going to need to be actually designed while being rewritten by the next person instead, or that you skipped trying to understand how the feature should work or thinking about the performance characteristics of the solution before you started and just let the LLM drive, so you never designed anything, arriving at something which "works" on your machine and passes the tests which were generated for it, but will hammer production under production loads. Neither will running it on your own machine or in Dev. No amount of telling the LLM to "Dig up! Make no mistakes!" will help with non-designed slop code actively poisoning the context, but you have to admire the attempt when you see comments added while removing code, referring to the code that's being removed. It's weird to see tickets now effectively go from "ready for PR" to 0% progress, but at least you're helping that person meet whatever the secret AI* usage quota is for their performance review this year.
- 0xbadcafebee 8mo ago> Tests can't tell you if the design of the code is fit for purpose, or about requirements you completely missed or punted on This is what acceptance tests are for. Does it do the thing you wanted it to do? Design a test that makes it do the thing, and check the result matches what you expect. If it's not in the test, don't expect it to work anywhere else. Obviously this isn't easy, but that's why we either need a different design or different tests. Before that would have been a tremendous amount of work, but now it's not. (Making this work requires learning how to make it work right. This is a skill with brand-new techniques which 99.999% of people will need over year to learn) > or that a core new piece that's going to be built upon next is barely-coherent, poorly-performing slop that "works" but is going to need to be actually designed while being rewritten by the next person instead This is the "human" part I mentioned being irrelevant now. AI does not care if the code is slop or maintainable. AI can just rewrite the entire thing in an hour. And if the tests pass, it doesn't matter either. Take the human out of the loop. (Concerned about it "rewriting tests" to pass them? You need independent agents, quality gates, determinism, feedback loops, etc. New skills and methods designed to keep the AI on the rails, like a psychotic idiot savant that can build a spaceship if you can keep it from setting fire to it) > or that you skipped trying to understand how the feature should work or thinking about the performance characteristics of the solution before you started and just let the LLM drive, so you never designed anything This is not how AI driven coding works. You have to give the AI very specific design instructions. If you do it right, it will make what you want. Sadly, this means most programmers today will be irrelevant because they can't design their way out of a wet paper bag. (You know how agile eschews planning and documentation, telling developers and product people to just build "whatever works right now" and keep rewriting it indefinitely as they meet blockers they never planned for? AI now encourages the planning and documentation.)
- iblaine 8mo agoI had a bad experience with greptile due to what seemed to be excessive noise and nit comments. I have been using cursorbot for a year and really like it.
- ottah 8mo ago> Today's agents are better than the median human code reviewer at catching issues Not my experience > A human rubber-stamping code being validated by a super intelligent machine What? I dunno how they define intelligence, but LLMS are absolutely not super intelligent. > If agents are approving code, it would be quite absurd and perhaps non-compliant to have the agent that wrote the code also approve the code. It's all the same frontier models under the hood. Who are you kidding.
- deleted 8mo ago[deleted]
- anon7000 8mo agoI’ve found only one good code review bot, and that’s Unblocked. It doesn’t always leave a comment, and when it does, it’s often found 1-2 real bugs in the code crossing multiple files (even like “hey you forgot to update this reference in this other file not edited in the PR”). Things you’d expect someone with a deeper knowledge of the code to know. You do get a handful of false positives, especially if what it reports is technically correct, but we’re just handling the issue in a sort of weird/undocumented way. But it’s only one one comment that’s easy to dismiss, and it’s fairly rare. It’s not like huge amounts of AI vomit all over PRs. It’s a lot more focused.
- bofadeez 8mo ago[flagged]
- dang 8mo ago"Don't be curmudgeonly. Thoughtful criticism is fine, but please don't be rigidly or generically negative." https://news.ycombinator.com/newsguidelines.html https://news.ycombinator.com/newsguidelines.html Edit: Could you please stop posting unsubstantive comments and flamebait? You've unfortunately been doing it repeatedly. It's not what this site is for, and destroys what it is for.
- pavan_panto 8mo ago[dead]
- las3r 8mo agoI would suggest you check out your Greptile discord and/or answer your messages on X where people are trying to reach you with problems and questions about your service. Unless that no longer matters.
- sebra 8mo agoI've tried Greptile and it's pretty much pure noise. I ran it for 3 PRs and then gave up. Here are three examples of things it wasted my time on in those 3 PRs: * Suggested to silence exception instead of crash and burn for "style" (the potential exception was handled earlier in code but it did not manage to catch that context). When I commented that silencing the exception could lead to uncaught bugs it replies "You're absolutely right, remove the try-catch" which I of course never added * Us using python 3.14 is a logic error as "python 3.14 does not exist yet" * "Review the async/await patterns Heavy use of async in model validation might indicate these should be application services instead." whatever this vague sentence means. Not sure if it is suggesting us changing the design pattern used in our entire code base. Also the "confidence" score added to each PR being 4/5 or something due to these irrelevant comments was a really annoying feature IMO. In general AI tools giving a rating when they're wrong feels like a big productivity loss as then the human reviewer will see that number and think something is wrong with the PR. -- Before this we were running Coderabbit which worked really well and caught a lot of bugs / implementation gotchas. It also had "learnings" which it referenced frequently so it seems like it actually did not repeat commenting on intentional things in our code base. With Coderabbit I found myself wanting to read the low confidence comments as well since they were often useful (so too quiet instead of too noisy). Unfortunately our entire Coderabbit integration just stopped working one day and since then we've been in a long back and forth with their support. -- I'm not sure what the secret sauce is but it feels like Greptile was GPT 3.5-tier and Coderabbit was Sonnet 4.5-tier.
- bjackman 8mo agoMy experience is that basic generic agents are useless but an agent with extensive prompting about your usecase is extremely valuable. In my case using these prompts: https://github.com/masoncl/review-prompts https://github.com/masoncl/review-prompts Took things from "pure noise" to a world where, if you say there's a bug in your patch, people's first question will be "has the AI looked at it?" FWIW in my case the AI has never yet found _the_ bug I was hunting for but it has found several _other_ significant bugs. I also ran it against old commits that were already reviewed by excellent engineers and running in prod. It found a major bug that wasn't spotted in human review. Most of the "noise" I get now just leads me to say "yeah I need to add more context to the commit message". E.g the model will say "you forgot to do X" when X is out of scope for the patch and I'm doing it in a later one. So ideally the commit messages should mention this anyway.
- clarus 8mo agoWhat should be added, I think, to code reviewing is that it can get really complex, for example if we add formal verification in the mix to catch very subtle bugs. So in the end I think there will still be some disappointment, as one would expect it should be fully automated and only about reading the code, like this article suggests. In reality, I think it is harder than writing code.
- EGREF 8mo ago4GVFDGDGFFFFEGRFEDS
- Manfred 8mo agoFuzzy automated reviews should always run in an interactive loop with a developer on their workstation and contain enough context to quickly assess if they are valid or not. When developers create a PR, they already feel they are "done", and they have likely already shifted their focus on another task. False positive are horrible at this point, especially when they keep changing with each push of commits.
- __0x01 8mo agoIs "AI code review" a correct term? A code review requires reasoning and understanding, things that to my knowledge a generative model cannot do. Surely the most an AI code review ever could be is something that looks like a code review.
- veunes 8mo agoThe main problem with current AI reviewers isn't catching bugs, it's shutting up when there is no bug. Humans have an intuitive filter like "this code is weird, but it works and won't break prod, so I'll let it slide". LLMs lack this, they generate 20 comments about variable naming and 1 comment about a critical race condition. As a result the developer gets fatigue and ignores everything. Until AI learns to understand the context of importance, not just code context, it will remain an expensive linter
- kachapopopow 8mo agoI can't get over how every single code rabbit ad was some incorrectly classified bug / completely wrong to begin with or pointless at best.
- hathym 8mo agoarticle by greptile, the AI code reviewer :D
- jv22222 8mo agoWe built an internal code review tool at the day job and are getting pretty good results with it (CLI tool). Here's a summary of the top-level ideas behind it. Hope it's helpful! Core Philosophy - "Advisor, not gatekeeper" - Every issue includes a "Could be wrong if..." caveat because context matters and AI can't see everything. Developers make the final call. (Just this idea makes it less annoying and stops devs going down rabbit holes because it it pretty good at thinking why it might be wrong) - Prompt it to be critical but not pedantic - Focus on REAL problems that matter (bugs, security, performance), not style nitpicks that linters handle. - Get the team to run it on the command line just before each commit. Small, focused reviews not after batching 10 commits. Small diffs get better feedback. Smart Context Gathering - Full file contents, not just diffs - The tool reads complete changed files plus 1-level-deep imports to understand how changed code interacts with the codebase. Prompt Engineering - Diff-first, context-second - The diff is marked as "REVIEW THIS" while context files are explicitly marked "DO NOT REVIEW - FOR UNDERSTANDING ONLY" to prevent false positives on unchanged code. BUT that extra context makes a huge difference in correctness. - Structured output format - Emoji-prefixed bullets ( Critical, Major, Minor), max 3 issues per section, no fluff or praise. - Explicit "Do NOT" list - Prevents common AI review mistakes: don't flag formatting (Prettier handles it), don't flag TypeScript errors (IDE shows them), don't repeat issues across files, don't guess line numbers. Final note - Also plugged it in to a github action for last pass, but again non blocking.
- just6979 8mo agoIf you need the AI to indicate "could be wrong" on everything it writes to prevent your devs from blindly following everything it says, you're doing it so wrong. That should be the default mindset. Of course it could be wrong.
- jv22222 8mo agoNot quite. The could be wrong part is very helpful because it has (on multiple occasions) dug up something that was long-lost-to-lore about why something should work in a non conventional way. Without that, the advice looks perfectly sensible and would send devs down a Rabbit hole, because the AI recommendation "looks right".
- m3kw9 8mo ago"review my code for edge cases" should pop it
- bp93592203 8mo agototally agree. Looks like the most common problem with the bubble is the terrible signal to noise ratio. Has anyone found a solution that works well? I see augment code is claiming their review agent is the best in terms of signal to noise ratio https://www.augmentcode.com/blog/we-benchmarked-7-ai-code-review-tools-on-real-world-prs-here-are-the-results https://www.augmentcode.com/blog/we-benchmarked-7-ai-code-re... has anyone tried it?
- pavan_panto 8mo ago[dead]
- AnViF 8mo ago[dead]
- AnViF 8mo ago[dead]
- DavidYoussef 8mo agoThe article nails the core issue but I think misdiagnoses the solution space. The problem isn't that AI code review exists - it's that current tools are solving the wrong problem. They review code that humans wrote. The actual crisis is reviewing code that AI wrote. When AI increases code volume by 10x but reviewer count stays flat, you don't need better review tools. You need risk triage. Not every PR deserves the same attention: - Typo fix to a README? L0. Auto-approve with an evidence log. - New utility function with tests? L1. One model scans it, posts findings. - Changes to auth middleware or payment flow? L3. Three models have to reach consensus before a human even looks at it. - Production deployment config? L4. Models + mandatory human sign-off. We've been building this (codeguard-action on GitHub, MIT licensed) - a GitHub Action that classifies PR risk, runs multi-model review proportional to that risk, and produces a cryptographic evidence bundle proving what was checked. The evidence is hash-chained and independently verifiable offline with a separate tool. The point isn't to replace human reviewers. It's to stop burning them out on L0-L1 changes so they have capacity for the L3-L4 ones that actually matter. The 786-PR-backlog problem mentioned upthread isn't a review problem. It's a triage problem.
- atomicnature 8mo agoAI code review has genuinely helpful - especially when we generate code with copilot, etc. Many times, these GenAI tools can delete/modify code mistakenly. I use LiveReview's git precommit features - so the review happens right before I commit code automatically. And it has saved me many (100s of) times. Give LiveReview's Precommit checks a try.