3 ms·
I am constantly correcting the AI code it gives me, and all I get for it is "oh your right! here is the corrected code" then it gives me more hallucinations c
by aspbee555 1y ago
I am constantly correcting the AI code it gives me, and all I get for it is "oh your right! here is the corrected code"
then it gives me more hallucinations
correcting the latest hallucination results in it telling me the first hallucination
- Trasmatta 1y agoI have this same experience. Vibe coding is literally hell.
- dwringer 1y agoIME it is rarely productive to ask an LLM to fix code it has just given you as part of the same session context. It can work but I find that the second version often introduces at least as many errors as it fixes, or at least changes unrelated bits of code for no apparent reason. Therefore I tend to work on a one-shot prompt, and restart the session entirely each time, making tweaks to the prompt based on each output hoping to get a better result (I've found it helpful to point out the AI's past errors as "common mistakes to be avoided"). Doing the prompting in this way also vastly reduces the context size sent with individual requests (asking it to fix something it just made in conversation tends to resubmit a huge chunk of context and use up allowance quotas). Then, if there are bits the AI never quite got correct, I'll go in bit by bit and ask it to fix an individual function or two, with a new session and heavily pruned context.
- Jcampuzano2 1y agoI agree with this, you will almost always get better results by simply undoing and rewording you prompt vs trying to coerce it to fix something it already did. Most of the time when I do use it, I almost always use just a couple prompts before starting a completely new one because it just falls off a cliff in terms of reliability after the first couple messages. At that point you're better off fixing it yourself than trying to get it to do it a way you'll accept.
- aspbee555 1y agothis is also what I started doing, sometimes it will give an actual correct answer but it usually easy to just start a new session. I can even ask the exact same question and get a correct answer with a new session
- ta_09867534567 1y ago100% this is my main workflow for a few years now when interacting with any LLM. Whenever I see people claim they struggle with using LLM as a part of their workflow, I ask them show me how you are solving a small problem with the AI, and subsequently, they show me this very sub-optimal workflow like GP is describing. 1. Ask a question / present a problem, but usually without enough context to the problem and solution space they want to zero in on. 2. The AI does an honest job given the context, but is off alignment in some specific way that the user did not clarify initially up front. 3. Asks the AI to correct for this, along with some multiple other requests for changes toward the solution they want. 4, 5, 6 Loop. They get a response, like the corrections (sometimes) and continue to make changes, in back-and-forth-conversation like interaction, only copying out corrected code blocks and copying in specific code chunks for correction. 7+. The output gets progressively worse and worse, undoing corrections/changes/modifications that were previously discussed. At this point I try to interrupt the spiraling death loop and ask the user: - (rhetorical) why are you talking to the AI like a human being? - What is in your context window at this point in the conversation? If they can answer the context window question, AND understand how the AI ingests input and produces output, usually its a lightbulb moment. If they don't quite realize that they are polluting their context window, then I try to get them to be aware that everything in the context window is statistically weighted and will affect the output. If a tainted input is provided, the chances of an untainted output are lower than otherwise. You want to provide high quality context window input, ideally fully control it. That means, you do NOT want to have a conversation with the AI for real work; you need to embrace `zero shotting` everything you ask. This approach maximizes exactly what the AI are best trained for, trained on, how they are trained, and how they `understand` things. This requires a lot more hand holding and curating prompting, ie prompt engineering, than people will honestly realize/admit to. Prompt engineering isn't black magic, its intelligent contextualization that plays into the strengths of the implicit knowledge AI has. Worst things for a LLM super user? - copy-paste tedium (doing it by hand) - RAG auto-compression (letting an algorithm determine critical context decisions) - opaque context window systems (how is the conversation stored and presented to the LLM each turn?) - system prompt inaccessibility in certain online providers (system prompt is still super critical for driving) - general `magic` behavior exhibited when using a plain/simple chat interface (this is usually unraveled ONLY by understanding the full context window) The only LLM that has been SUCCESSFUL at conversing with me and maintaining state through the flowing conversation has been the newest Gemini 2.5 Pro offering, and ONLY up to 100K out of 1M context window. I have had (very minor) forgetting after 100K, and I deep dove into the conversation at that point to understand what was going on, and it appears that the conceptual conversation compression is in some way, lossy losing some conversation bits. Every other LLM has had the facade of maintaining conversation state, but only Gemini 2.5 Pro Preview has actually held that up (with firm limitations!). I suspect that large context window optimization/compression is to blame, some providers are aggressive with it.
- Jcampuzano2 1y agoI find it's only really useful in terms of writing entire features if you're building something fairly simple, on top of using the most well known frameworks and libraries. If you happen to like using less popular frameworks, libraries, packages etc it's like fighting an uphill battle because it will constantly try to inject what it interprets as the most common way to do things. I do find it useful for smaller parts of features or writing things like small utilities or things at a scale where it's easy to manage/track where it's going and intervene But full on vibe coding auto accept everything is madness whenever I see it.
- mrweasel 1y agoSame thing happens to me. The LLM will make up some reasonably sounding answer, I correct it, three, four, five time, and then it circles back to the original answer... which is still just as wrong. Either they don't retain previous information, or they are so desperate to give you any answer that they'd prefer the wrong answer. Why is it that an LLM can't go: Yeah, I don't know.