4 ms·
You and the AI agree, but how about you and the team who will eventually read and do code review. Cognitive burden increases marginally with AI assisted coding
by prisonguard 25d ago
You and the AI agree, but how about you and the team who will eventually read and do code review.
Cognitive burden increases marginally with AI assisted coding.
This is why we haven't seen big projects(think browsers and browser engines) spawning in the past year.
- trhaynes 25d agoIndeed. I'm curious what the security team has to say about that approach, for example.
- brabel 25d agoIf you get an AI to review the code especially for security, it does a very good job st finding issues. Better than any human reviewers I have worked with, and getting better. As someone who works in security , I feel much less worried about security bugs on code reviewed by a AI for security issues, be it written by AI or human.
- yehoshuapw 25d agoextra reviews never hurt. But trusting only the AI to both code and review security-wise? not even close to usable
- a2ff6eeb0 25d ago> but how about you and the team who will eventually read and do code review. Why are you reviewing AI code in detail? Do you also review the assembly output of GCC line by line?
- rogerrogerr 25d agoI can’t remember the last time GCC emitted code that just flat out called the wrong function. If it did that occasionally, I would review it.
- electric_toucan 25d agoCompilers like GCC are deterministic and the source code already fully defines the behavior. LLMs are non-deterministic and will accept ambiguity, filling in details where you haven’t. These sorts of comparisons aren’t really fair. In the case of writing, it’s like hiring someone to write a book for you vs. hiring someone to translate a book you wrote into another language. In the first case, you didn’t really define the message for readers, whereas in the second case you did, and the translator is converting that same message for another audience to consume.
- a2ff6eeb0 25d agoSure, but I don't know what GCC's behavior is, and I don't vet behavior differences between compiler upgrades. As long as the output works, why does it matter that the black box is deterministic?
- rogerrogerr 25d agoBecause _someone_ has vetted the output of GCC. It’s used in flight-critical stuff. The closest thing we have to vetting LLMs is “whoa look, it escaped this sandbox, that’s prolly not great but it’s so cool!”
- a2ff6eeb0 25d agoSure, I manually test the output of the LLM. Manual testing is actually the main role for humans doing software engineering these days. I wouldn't use it for flight control software yet, at least not without careful review, but most software isn't exactly critical. At the same time, I wouldn't trust flight control software that was only reviewed by humans, since AI is so much better at debugging. We'll probably need humans in the loop for safety critical software for at least a year or two, before AI fully outpaces humans at generating correct code.
- gopher_space 25d agoCan’t imagine a client allowing me to pass the buck like this.
- lelanthran 25d agoCome on, this is a take we expect from a 1st grader! Gcc makes maybe 1 mistake ever 2 billion emissions. LLMs make 1 mistake ever 3rd emission.
- cozzyd 25d agoFor instructions you really care about, yes of course you review the assembly output! Usually when you're doing SIMD or want to check atomics are doing what you expect.
- a2ff6eeb0 25d agoYou can do that with LLMs for the parts you really care about too. The LLMs aren't regenerating the codebase from scratch every time, so the results stick around.
- bigstrat2003 25d agoNo, because GCC doesn't randomly fuck up the assembly generation (much less on a fairly frequent basis the way LLMs do). If it did, you bet I'd be reviewing the assembly line by line, or decline to use such a poorly performing tool (as I have with LLMs).
- zdragnar 25d agoIt's anthropic, of course the entire team is also using AI to do the code reviews.