5 ms·
> Each model’s responses are ranked by a high-performing judge model — typically OpenAI’s o3 — which compares outputs for quality, relevance, and clarity. These
by comex 1y ago
> Each model’s responses are ranked by a high-performing judge model — typically OpenAI’s o3 — which compares outputs for quality, relevance, and clarity. These rankings are then aggregated to produce a performance score.
So there's no ground truth; they're just benchmarking how impressive an LLM's code review sounds to a different LLM. Hard to tell what to make of that.
- ImageXav 1y agoYes, especially as models are known to have a preference towards outputs of models in the same family. I suspect this leaderboard would change dramatically with different models as the judge.
- spiderfarmer 1y agoThey are different models already but yes, I already let ChatGPT judge Claude's work for the same reason.
- jacquesm 1y agoI don't care about either method. The ground truth should be what a human would do, not what a model does.
- mirekrusin 1y agoThere may be different/better solutions for almost all those kind of tasks. I wouldn’t be surprised if optimal answer to some of them would be refusal/defer ask, refactor first, then solve it properly.
- jacquesm 1y agoThat response is quite in line with the typical human based PR response on a first draft. There is a possibility that machine based PR reviews are better: for instance because they are not prejudiced based on who is the initiator of the PR and because they don't take other environmental factors into account. You'd expect a machine to be more neutral, so on that front the machine should and possibly could score better. But until the models consistently outperform the humans in impartially scored quality vs a baseline of human results it is the humans that should call this, not the machines.
- jeltz 1y agoI wouldn't necessarily expect a machine to be more neutral. Machines can easily be biased too.
- jacquesm 1y agoOn something like a PR review I would. But on anything that would involve private information such as the background, gender, photographs and/or video as well as other writings by the subject I think you'd be right. It's just that it is fairly trivial to present a PR to a machine in such a way that it can only comment on the differences in the code. I would find it surprising if that somehow led to a bias about the author. Can you give an example of how you think that would creep into such an interaction?
- eviks 1y agoWhy is it hard to ignore an attempt to assess reality that is not grounded in reality?
- jtrn 1y agoThat's an extremely dense question :) (Not pejorative, but conceptual dense). I had some fun trying to answer it, ignoring fixating on whether or not the premise is true, for argument's sake. My answer is: I would think "attempting to assess reality that is not grounded in reality" is hard to ignore due to a combination of "it's what is available," being easy to understand, and seeming useful (decoupled from whether it's really so). As a result, it's hard to ignore because it's what is mostly available to us for consumption and is easy to make "consumable." I think there is a LARGE overlap in this topic with my pet peeve and hatred of mock tests in development. They are not completely useless, but their obvious flaws and vulnerabilities seem to me to be in the same area: "Not grounded in reality." Said another way: Because it's what's easy to make, and thus there is a lot of it, creating a positive feedback loop of mere-exposure effect. Then it becomes hard to ignore because it's what's shoved in our face.
- raincole 1y agoThat's how 99% of 'LLM benchmark numbers' circulating on the internet work.
- Lionga 1y ago[flagged]
- qsort 1y agoNo, they aren't. Most benchmarks use ground truth, not evaluation by another LLM. Using another LLM as verifier, aside from the obvious "quis custodiet custodes ipsos", opens an entire can of worms, such as the fact that there could be systematic biases in the evaluation. This is not in and of itself disqualifying but it should be addressed, and the article doesn't even say anything.
- sigmoid10 1y agoGround truth evaluation is not that simple unless you are doing multiple-choice-style tests or something similar where the correctness of an answer can be determined by a simple process. Open ended natural language tasks like this one are incredibly difficult to evaluate and using LLMs as judge is not just the current standard, it is basically the only way to do it at scale economically.
- qsort 1y agoThe original comment was this: > So there's no ground truth; they're just benchmarking how impressive an LLM's code review sounds to a different LLM. Hard to tell what to make of that. The comment I replied to was: > That's how 99% of 'LLM benchmark numbers' circulating on the internet work. And that's just false. SWE-Bench verified isn't like this. Aider Polyglot isn't like this. SWE-Lancer Diamond isn't like this. The new internal benchmarks used by OpenAI in GPT-5's model card aren't like this. Maybe this benchmark is a special snowflake and needs LLM-as-a-judge, but this doesn't invalidate the original concern: setting up a benchmark this way runs into a series of problems and is prone to show performance differences that might not be there with a different setups. Benchmarks are already hard to trust, I'm not sure how this is any more indicative than the rest.
- with 1y agoIt’s a widely accepted eval technique and it’s called “llm as a judge”
- magicalhippo 1y agoShouldn't one review the ratings of say a random 1% to ensure it's performing as expected?
- jacquesm 1y agoAccepted does not mean correct. It's like using a rubber yardstick as the means to figure out who won the pumpkin growing competition.
- ben_w 1y agoI'd say it's worse than that, a rubber ruler still has a definite length when not under tension etc. This might be more like asking amateur painters to each paint a picture of a different one of the pumpkins, then judging each other's paintings without seeing the actual pumpkin that painting was based on.
- jacquesm 1y agoOk, that is indeed better. For a further improvement we should let the previous generation of paintings judge the new one.
- sensanaty 1y agoAccepted by whom, the people shoving AI down our throats?
- deleted 1y ago[deleted]
- kingstnap 1y agoIt's widely accepted because it's cheap, but LLMs aren't really good judges. It's supposed to leverage a "generate vs. critique" gap in skill level as a form of self-improvement. It's easier to judge how good food is vs. make it. But here's the thing. When it comes to code review, you need to be effectively as skilled as the person who wrote it. There isn't really a gap. And then the real clincher is this. LLMs naturally have a skill gap between their judgement and generation skills as is. The reason is that they have superhuman pattern matching and memorization ability. They can use their memorized patterns as a massive crutch for their actual reasoning skills, but they can't do the same for judgement calls in code review.
- shikon7 1y agoAlso, using an OpenAI model to judge the performance of an OpenAI model seems prone to all kinds of biases.
- mirekrusin 1y agoExactly, they should at least compare with judges as best models from others, ideally verified by human/ground truth/tests.
- LauraMedia 1y agoAm I missing something? If LLM-1 is supposed to judge LLM-2, doesn't LLM-1 have to be better than LLM-2? If LLM-1 is only 40% as good at coding as LLM-2, why would you trust the LLM with the lesser knowledge?
- BlindEyeHalo 1y agoAt the heart of the P vs NP problem lies the observation that solution verification seems to be much easier than solution generation. If that applies in this context is another question but I think it is not unreasonable to assume that the judge needs to be less powerful than the performer. Or in other words, I don't need to be a chef myself to decide if a meal is good or not.
- rowanG077 1y agoThat really doesn't hold for all problems. You can imagine any number of problems where a valid solution is easier, complexity wise, to generate than it is to validate. A trivial example is semiprime factorization. Easy to generate any semiprime, hard to factor.
- jama211 1y agoPretty sure they know that, their point still stands
- deleted 1y ago
- kruxigt 1y ago[dead]
- croes 1y agoIt undermines the private benchmark approach if the evaluation is done that way.
- dvfjsdhgfv 1y ago> Hard to tell what to make of that. It's not hard. You are visiting a website with an .ai domain. You already know what the conclusions will be.
- andrepd 1y agoIt's almost too on the nose to be satire, yet here we are.