4 ms·
> Those are not code problems. They are evaluation problems. > Code becomes precious when it is the only place knowledge lives. Reading AI code all day is _ag
by trjordan 4mo ago
> Those are not code problems. They are evaluation problems.
> Code becomes precious when it is the only place knowledge lives.
Reading AI code all day is _agonizing_. Just, a horrible way to live, and it melts people's brains at the moment you need them to be the most capable.
Manual programming has this really productive and gratifying feedback loop, where you read the code, write the code, and fix it until it compiles/runs/does what you want. AI code not only does half that for you, but it makes the "click" at the end uninspiring because you're never sure if it's cheated a bit to get to that moment.
Trying to operate with AI-generated code as the only durable artifact of programming is a dead end for the industry. Charity points to (and correct discards) architecture diagrams/specs as an interesting space to work in. My suspicion is that it's closer to the thing that's hand-written: prompts, markdown plans, and other nudges. Focus on the thing that you, as a human, produce, and that's the basis for both the core loop of "did the AI follow my instructions" and it's higher-leverage when you go to code review.
By the time you get to the PR, you've probably typed enough to Claude that you can regenerate the code, but the current industry default is to just throw away all those sessions and ship the code. That's backwards!
- mooreds 4mo agoAre there any products out there that are capturing the prompts/sessions? I imagine you could do it in an adhoc way, asking Claude to write up a summary of the session as part of the commit message. But is there anything else that's more structured/higher level?
- trjordan 4mo agoWe're working on it, thought it's all early. I'd love feedback: https://tern.sh https://tern.sh First product compares the code to the prompts and highlights places the agent made decisions you weren't involved in: https://tern.sh/docs/tours/ https://tern.sh/docs/tours/
- latentsea 4mo agoWe just have hook that runs on git push that instructs Claude to ensure the PR description is up to date. Works well enough for us.
- sdesol 4mo agoI am working on solving the AI Code Provenance problem and I believe my repos may be the first that provides AI code provenance. See the following example: https://github.com/gitsense/gsc-cli/blob/main/internal/cli/root.go https://github.com/gitsense/gsc-cli/blob/main/internal/cli/r... Notice how the code block header attributes the model. The UUID can be traced to the conversation so everybody can tell exactly how the code came about. For this to work though, you need to use my chat app as it ensures you can't tamper with things if you are truly serious about AI code provenance. I also have a lot more human-focused method which is part of my CLI tool. https://github.com/gitsense/gsc-cli https://github.com/gitsense/gsc-cli I am currently looking at making pi (https://github.com/earendil-works/pi https://github.com/earendil-works/pi) support AI code provenance, but for now if you want a more structured way to capture what you have done in an agent session that can be used in code reviews and be carried forward as knowledge that lives inside your repository, I have gsc lessons The basic idea is, after you have finished chatting/working with the agent, you would work with it to identify lessons worth carrying forward. You can store your session if you want, but really, the lessons should be something that can help you review code better and to prevent future mistakes. I have a real working example at https://github.com/gitsense/smart-ripgrep https://github.com/gitsense/smart-ripgrep This is a fork of the BurntSushi/ripgrep repository. It shows how you can use lessons to learn from past design decisions.
- mplanchard 4mo agoSo many. The sibling comments, plus GitAI, plus empathic, plus many others
- philbo 4mo agoIf a coworker dumped a 5k-line code review on you, you'd tell them to come back when it's broken down into smaller, reviewable chunks. Large dumps of code are basically unreviewable by humans, but it seems like a lot of people have forgotten about that when it comes to LLMs.
- win311fwg 4mo agoIt is not so much forgetting as much as it is acceptance that when welcoming AI into a codebase, the code can no longer matter; that all that matters is that the properties of the system are validated. That isn't a change that comes free, so nobody should be expecting magic, it is a different set of tradeoffs. There is no such thing as a panacea.
- ChrisLTD 4mo agoHow can the code no longer matter? It literally is the logic (not to mention performance, and reliability) of the software.
- win311fwg 4mo agoYou might say in the same way that machine code stopped mattering when programming languages gained in popularity. Almost nobody will ever review machine code. I anticipate 90% of all programmers today wouldn't even know how. The move again is towards a higher level of abstraction; this time validation. Instead of describing how the program is to function, you define the properties of the system and let the fancy compiler figure out what the code should look like. If that means something that a human would call spaghetti, oh well.
- ChrisLTD 4mo agoI can be 99.99999999% certain when I write an if statement like "if (x > 1) do y" that the compiler will turn that into the equivalent machine code. So, yes, unless I hit some crazy performance bottleneck, I'm not concerned about reviewing the machine code. However, LLM outputs change with slight re-wording of prompts and with each new model release. I could hand write a test that says if x > 1 make sure y happens, but then what productivity was gained?
- keybored 4mo agoFlintstone Engineering is applying Space Age synthetic intelligence (in a metaphorical sense) technology with code generation. Babysitting, version controlling, etc. generated code should be a thing of the past. But that is what GenAI is. At the very least apply it at a higher level: specification, proofs, anything but generating Rust/Java/C and then letting yourself or an agent babysit it.
- gavinh 4mo agoI agree that reading AI code all day is agonizing. We're relying on code review to develop parts of our mental model of the system that were previously developed through coding. We're having more difficulty comprehending and recall details of the system. This is probably unsurprising; people recall information better that they "generated" than information they read. I am applying some lessons from pedagogy to extend code review. If this resonates with you, I would like to talk.
- trjordan 4mo agoWould love to chat -- ping me tr at tern dot sh
- arsmoriendi 4mo agoI also would like to know more about those pedagogy lessons you're applying.
- agumonkey 4mo agothe act, eval, adjust loop is probably neurologically important.. reading about things you didn't dive into is really a dread depending on your industry, you might be able to ship half-slop and then fix some bugs downstream though
- deaton 4mo agoIt almost seems like the juice might not be worth the squeeze. If you want verifiable code that conforms well to a well-designed plan, you have to basically write pseudocode and have the AI translate it for you. At that point why use the AI to write the code at all? And then, personally, I find that I just have more fun planning, writing, and debugging myself. I think its kinda the part of programming that I fell in love with in the first place.
- pydry 4mo agoThis is the core of the insanity nobody who vibe codes seems to be able to grasp. It's not even that its more fun, AI can spew endless slop boilerplate but it simply can't handle boiling an application down to its component essence in a way that makes it straightforward to maintain and bug free.
- vjvjvjvjghv 4mo ago"Reading AI code all day is _agonizing_. Just, a horrible way to live, and it melts people's brains at the moment you need them to be the most capable." I think it's very similar to dealing with large offshore teams. Every day you get a huge pile of code to review. It's really exhausting. I prefer dealing with AI because at least it tends to follow rules once I write them down. Not so much with a lot of offshore guys. Same mistakes every day. I guess my company needs to hire better offshore devs....
- tonyedgecombe 4mo agoI do wonder what all this means for offshore development. What's the point of sending the work to another continent with different time zones and languages when your AI can do the same job directly.
- theshrike79 4mo agoOffshoring and contracting bulk work out is decreasing. AI can give you 80-90% of the quality, but the feedback loop is hours or days, not weeks. This means the inevitable iteration of "no make that a bit greener, move that there, that is the wrong style for this scene" can be done faster, which means cheaper.