3 ms·
How? The agent does something incredibly difficult, and there is no human code review. Why do you think it doesn't match ?
by bonjourjoel 2mo ago
How? The agent does something incredibly difficult, and there is no human code review. Why do you think it doesn't match ?
- ben_w 2mo ago"Here's a case study where it worked once" != "the end of human code review". Don't get me wrong, this is impressive and capabilities do still seem to be on a generally upward trend with nothing "hitting a wall" despite all the parroting of that phrase, but there's a huge gap between the first time a machine manages something impressive enough to document, and that machine becoming so good at that task that humans need not apply.
- bonjourjoel 2mo agoAgreed. But also this inequality: "Is this the end of human code review?" != "the end of human code review" Now since the problem in this study is more complex than 90 to 99.5% of what a normal ticket is, the question makes sense. The agent does not have knowledge of the project, at the begining of each session, and manages to handle something more complicated than virtually anything a developper has to do. The question holds IMO.
- kubb 2mo agoLiterally every top 10 LLM can tell you why. Maybe we’re in „this is the end of making other people spell it out for you” era?
- bonjourjoel 2mo agoLLMs are leaning toward Corporate first order. They will judge anything based on how much corporation would like it said. I call this "corporate RLHF" and they are ALL into it. The fact is, the agent did something amazingly difficult, that zero human on earth could do. Read the spec if you believe someone could do it. It did it clean, everything is public, and it works perfectly. Zero code review. https://aisovereignlabs.ai/docs/case-study/liveSession/logs/liveSession-logs-phase2-specification-en.pdf https://aisovereignlabs.ai/docs/case-study/liveSession/logs/... 99.9% of the IT dev tickets are easier than this, excluding off topics like discrete math, computer vision algorithms, etc, which are irrelevant here. Do you think there's anything you've done, with human code review, in the past 30 years, matching the difficulty of this ?