4 ms·
Exactly. That would mean going back to square one, and I also personally don't review the code anymore but I am more focused on asking the model to demonstrate
by menaerus 1mo ago
Exactly. That would mean going back to square one, and I also personally don't review the code anymore but I am more focused on asking the model to demonstrate the value it created through benchmarks, workload-generators, and e2e tests.
Mostly it proves as a valid approach, barring the bugs the model can introduce to value-demonstrating benchmarks which can of course skew the evidence on hypotheses, and thus code trajectory the model opts to go with.
The problem I see with this is really that I am not anymore under the control but I am not sure I see other alternative. I am becoming more and more like a system observer with surface-level understanding of the system rather than the engineer with zoomed-in level of understanding of how the code actually behaves. Perhaps we're transitioning into a QA roles present.
- fortzi 1mo agoAre you working on complex systems that serve many users in production? Do you actually save time?
- menaerus 1mo agoYes, I can't say exactly what but the backbone of cloud computing and/or infra, think databases, distributed storage, filesystems, ...
- t-writescode 1mo agoI mean, AWS did just recently go down due to a vibe-coding / agent-run-amuck incident, didn’t it?
- menaerus 1mo agoI don't know which one exactly but can we now count how many incidents there were in the pre-agent era?
- rasz 1mo ago>asking the model to demonstrate the value it created through benchmarks, workload-generators, and e2e tests. LLMs are fantastic at faking those