6 ms·
> Nowadays, it is better to write prompts Very big doubt. AI can help for a few very specific tasks, but the hallucinations still happen, and making things up
by Disposal8433 1y ago
> Nowadays, it is better to write prompts
Very big doubt. AI can help for a few very specific tasks, but the hallucinations still happen, and making things up (especially APIs) is unacceptable.
- NitpickLawyer 1y ago> but the hallucinations still happen, and making things up (especially APIs) is unacceptable. The new models are much better at reading the codebase first, and sticking to "use the APIs / libraries already included". Also, for new libraries there's context7 that brings in up-to-date docs. Again, newer models know how to use it (even gpt5-mini works fine with it).
- sigseg1v 1y agoWhat size of codebases are we talking here? I've had a lot of issues trying to do pretty much anything across a 1.7 million LOC codebase and generally found it faster to use traditional IDE functionalities. I've had much more success with things under 20k LOC but that isn't the stuff that I really need any assistance with.
- wongarsu 1y agoIn languages with strong compile-time checks (like say rust) the obvious problems can mostly be solved by having the agent try to compile the program as a last step, and most agents now do that on their own. In cases where that doesn't work (more permissive languages like python, or http APIs) you can have the AI write tests and execute them. Or ask the AI to prototype and test features separately before adding them to the codebase. Adding MCP servers with documentation also helps a ton. The real issues I'm struggling with are more subtle, like unnecessary code duplication, code that seems useful but is never called, doing the right work but in the wrong place, security issues, performance issues, not implementing the prompt correctly when it's not straight forward, implementing the prompt verbatim when a closer inspection of the libraries and technologies used reveals a much better way, etc. Mostly things you will catch in code review if you really pay attention. But whether that's faster than doing the task yourself greatly depends on the task at hand
- Disposal8433 1y ago> the obvious problems can mostly be solved by having the agent try to compile the program The famous "It compiles on my machine." Is that where engineering is going? Spending $billions to get the same result as the laziest developer ever?
- wongarsu 1y agoIf it compiles on my machine then the library and all called methods exist and are not hallucinated. If it runs on my machine then the called external APIs exist and are not hallucinated That obviously does not mean that it's good software. That's why the rest of my comment exists. But "AI is hallucinating libraries/APIs" is something that can be trivially solved with good software practices from the 00s, and that the AI can resolve by itself using those techniques. It's annoying for autocomplete AI, but for agents it's a non-issue
- lsaferite 1y agoThe subtle bugs are horrible. We used Claude Code the other day to add a new record type to an API and it was mostly right. CC decided (for some weird reason) to use a slightly different return shape on a list endpoint than the entire rest of the API. It changed two field names (count/items became total_count/data). This divergence was missed until the code was released because it 'worked' and had full tests and everything. But when the standard client lib code was used to access the API it failed on the list endpoint. Didn't take long to discover the issue. Luckily, it was a new feature so nothing broke, but it was a very clear reminder that you have to be very thorough when reviewing coding agent PRs. FWIW, I use CC frequently and have mostly positive things to say about it as a tool.
- deleted 1y ago[deleted]
- salomonk_mur 1y agoHard disagree. LLMs are now incredibly good for any coding task (with popular languages).
- Disposal8433 1y agoYou can't disagree with facts. Every time I try to give a chance to all those LLMs, they always use old APIs, APIs that don't exist, or mix things up. I'll still try that once a month to see how it evolves, but I have never been amazed by the capabilities of those things. > with popular languages Don't know, don't care. I write C++ code and that's all I need. JS and React can die a painful death for all I care as they have injected the worst practices across all the CS field. As for Python, I don't need help with that thanks to uv, but that's another story.
- dingnuts 1y agoIf you want them to not make shit up, you have to load up the context with exactly the docs and code references that the request needs. This is not a trivial process and ime it can take just as long as doing stuff manually a lot of the time, but tools are improving to aid this process and if the immediate context contains everything the model needs it won't hallucinate any worse than I do when I manually enter code (but when I do it, I call it a typo) there is a learning curve, it reminds me of learning to use Google a long time ago
- th0ma5 1y agoSo, I've done this, I've pasted in the headers and pleaded with it to not imagine ABIs that don't exist, and multiple models just want to make it work however they can. People shouldn't be so quick to reply like this, many people have tried all this advice... It also doesn't help that there is no independent test that can describe these issues, so all there is anecdote to use a different vendor or that the person must be doing something wrong? How can we talk about these things with these rhetorical reflexes?
- 1y ago
- mg 1y agoDo others here encounter that problem? I never do. I can't remember the last time I saw a hallucination in a commit. Maybe it's because the libraries I use are made from small files which easily fit into the context window.
- brulard 1y agoSame here, very low hallucination rate and it can pretty quickly correct itself (Claude Code). To force it to use recent versions of libraries instead of old ones, it's good to have it specifically required in CLAUDE.md and also having docs MCP (like context7) can help.
- verdverm 1y agoWriting prompts makes these issues way less significant and makes the agents way more capable. Prompt / context engineering is still an underrated and underutilized activity (imo)
- tomjen3 1y agoIts surprisingly fine, as long as you allow the AI to iterate on its work. It will discover that it doesn't compile, and then maybe lookup the API and then it will most often fix it and move on. AI is no more capable of reliably one shotting solutions that you are.