4 ms·
I don't know. I find that I'm moving up a level and improving my product-management skills while delegating most of the code to the agents. I'm still very much
by paulmooreparks 4mo ago
I don't know. I find that I'm moving up a level and improving my product-management skills while delegating most of the code to the agents. I'm still very much hands-on with the design and requirements, and I'm asking questions like, "What's our security story for XYZ?", "Are we accounting for colour-blindness?", etc. Not being down in the code allows me to prairie-dog a bit more and see the landscape better.
- tamaz_entalpa 4mo ago[flagged]
- ctdinjeu7 4mo ago[flagged]
- bluGill 4mo agoI'm about 50% that way. However when the AI is done coding I then step back and review to find places the code quality is unacceptable. I also have to stop the AI once in a while because it forgets the point and does something stupid. Junior engineer learn, AI does not.
- seunosewa 4mo agoUnless you log its mistakes and how they were solved in decisions.log
- deleted 4mo ago[deleted]
- dpoloncsak 4mo ago> Junior engineer learn, AI does not. This is technically true, but lets not act like we haven't seen immense improvement of both models are harnesses for these models in the past years. They may not be learning, but they are getting better
- nyrikki 4mo agoThey are getting better at historical data, not at the fundamental issue. As a recent example, I recently had to abandon the multiple LLM reviewer/verifier model I was using because zig 0.16 was released with major changes. I actually reverted back to full self hosted because the foundation models we’re trying too hard to revert to the older versions of the language. It is going to be a balancing act and there is fundamentally no way for LLMs to get around this. We will have to develop methods to do so, most likely by focusing agents on problems that are more static.
- askonomm 4mo agoI find great success in not relying on LLM's built-in knowledge, but giving it links to necessary docs/manuals and have it read that before doing anything.
- embedding-shape 4mo agoAlso, add "no assumptions or guesses" and if you use a model with really strong prompt adherence (most SOTA models), they'll figure out the right version first, then look up docs, then implement.
- nyrikki 4mo agoCurrently, with zig 0.16 the agent has to have access to the zig compiler and std library to even produce code that will compile. If you have zig installed, you can run ‘zig std’ to see that. You still have the limitations of attention etc… Even zed’s agent will leverage that built in tarball, but it doesn’t solve the problem, especially as some of the languages killer features are unavailable in C and other languages.
- smj-edison 4mo agoQuestion for you, since I also use Zig 0.16: how do you get it to use Zig idioms? I use Kimi 2.6, and I feel like whenever I try to get my agent to write modern Zig based on a C reference it decides to start writing everything in a C style (doesn't use defer, doesn't use opaque enums even when I explicitly tell it to, doesn't use Zig's error unions, swallows errors instead of asserting, and some more). It's quite frustrating, and a lot of catchable errors crop up until I've beat modern practices into it.
- paulmooreparks 4mo agoI don't abandon the code to the agent entirely. I have my own... I wouldn't call it a harness as such, but rather a shared Kanban board, and it'll be the subject of a "Show HN" soon. It suffices to say that I define Kanban cards for each feature or bug, and I have clearly defined review points for each card, post-spec and post-code, where I step in. On top of that, after my review, there is an agentic review, and agents can and do catch things that I missed. The quality of the software has improved quite a bit since I instituted that flow.
- gameshot911 4mo agoI think the right comparison is AI models versions, not intra-AI-model growth (although even that can 'learn' with persistent memory & contexts).
- lanstin 4mo agoI find that is the case for production code that will be running 24x7 unattended, but also Claude lets me build a lot more highly specific dashboards or visualization tools that I really don’t give a fig what the code is, as long as the numbers sum up and the links work. So my batch job I am careful with, the dashboard I check every morning to see what batches and lambdas failed eh I can wait the two minutes it takes to populate all the data; better to have time to top off coffee than having to understand modern JavaScript, canvas, D3 etc and web frameworks. I do force it to use python and flask for the web serving and SQLite for caching/ memoization, but everything else carte blanche.
- xantronix 4mo agoOne thing I've noticed is that LLMs have allowed middle managers trapped inside the role of a developer to finally self actualise.
- slopinthebag 4mo ago> What's our security story for XYZ? lmao I hope I never use your products with anything sensitive ever
- paulmooreparks 4mo agoI think you missed the point. I don't abandon security to whatever the agent decides to write.
- willXare 4mo ago[flagged]