6 ms·
I have a weird issue with using AI for coding. I can code something entirely by myself at my baseline speed; call it 1x. Or I can use Claude to do it, and it do
by Xcelerate 2mo ago
I have a weird issue with using AI for coding. I can code something entirely by myself at my baseline speed; call it 1x. Or I can use Claude to do it, and it does it in 1/10 - 1/4 of the time. The problem, however, is that to review Claude’s code properly takes 2-3x the amount of time it would have taken me to write it all by hand.
So my two choices are basically “YOLO, LGTM” and hope I can revert if it breaks something, or to just write all the code by hand from the start. With the increased pressure for output, I’ve noticed both myself and coworkers tending more toward “commit and hope it works” over time. It’s sort of perverse incentives in a way...
- taberiand 2mo agoThe models are good enough these days that I see it like managing a team of juniors. The LLMs need guidance and oversight, and about the same amount of time I'd spend reviewing code from a junior I spend on the LLM output. I think generally it's best practice to work in a way that "commit and hope it works (because it passes all the CI/CD automated testing and verification, and it's a change gated behind feature flags and has an identified and minimal blast radius, etc)" is possible
- vekker 2mo agoI feel like these days it's more like managing a team of seniors who are technically very competent, but may not always have the full domain context of the problem they're solving. And/or they forget parts, if that context is too large. And/or they have a poor temporal dimension (deprecated badly maintained information polluting the context). They also tend to overengineer solutions if not guided well. If you think about it, in that respect it's not that different from managing actual human senior engineers. I think this is why focussing one's limited human attention more the input (defining clear requirements) as well as on validating the output (good CI/CD including end-to-end automated tests) is far more important than manual code reviews and micro-managing the development process.
- jurgenburgen 2mo agoReviewing LLM code is quite different from junior code. With juniors you tread carefully and give feedback only on important points to encourage growth. With LLMs you channel the inner sailor and nitpick so much that even a senior would start to cry.
- t-writescode 2mo agoThis is what I've seen. There has been a middle-ground though, that I've seen. I was using Jetbrains' AI Assistant, or whatever it's called, during an interview for the (nearly) first time, and I only used it as auto-complete, but it was consistently, surprisingly, very accurate. I literally did write it faster. I scanned the literal next line, in the context of what I was writing, with all that context in my brain, and I saw what it wrote and went "yeah, that looks good". I'm a very fast typist (~> 100 wpm) and I still found it beneficial for writing out work fast. Further, since everything is still in my mental context and I can evaluate it at the same time, I feel like I can trust that type of code more. But maybe that's just me.
- stephenr 2mo agoHow long ago was this? I tried their "local ai autocomplete" thing a few years ago for a little while and it was hot garbage. It guessed the right/acceptable completion about 30% of the time at best. I haven't looked at it since.
- t-writescode 2mo agoWithin the last, like, ... 3 months, I think.
- encyclopedism 2mo agoEconomics drives behaviour more than people think. Your employer cares about results. It's 'YOLO, LGTM' all the way down now.
- mrbnprck 2mo agonot to forget the cognitive strain of the constant back and forth between LLM and you. One challenge is when knowing how the code should look like, the LLM solution always looks weird, and one tries to maunally steer against it, so accepting a bit of "good enough" is unavoidable to gain some productivity
- devinparadise 2mo agoThose aren’t the only choices though. What you can do now is have a separate agent review the diff before you create the PR, or to leave feedback on the PR for the agent that wrote the code to fix. You write the feature with one agent, and have different agents do the review, each with their own rules and context. Each review agent can even focus on different things, such as security, or adherence to your stylistic or architectural preferences. This can find all kinds bugs and edge cases the first agent missed, and that you might never have thought of yourself, even with careful human review.