7 ms·
This very much aligns with my experience — I had a case yesterday where opus was trying to do something with a library, and it encountered a build error. Rather
by mikeocool 1y ago
This very much aligns with my experience — I had a case yesterday where opus was trying to do something with a library, and it encountered a build error. Rather than fix the error, it decided to switch to another library. It then encountered another error and decided to switch back to the first library.
I don’t think I’ve encountered a case where I’ve just let the LLM churn for more than a few minutes and gotten a good result. If it doesn’t solve an issue on the first or second pass, it seems to rapidly start making things up, make totally unrelated changes claiming they’ll fix the issue, or trying the same thing over and over.
- the__alchemist 1y agoThis is consistent with my experience as well.
- fcatalan 1y agoI brought over the source of the Dear imgui library to a toy project and Cline/Gemini2.5 hallucinated the interface and when the compilation failed started editing the library to conform with it. I was all like: Nono no no no stop.
- BoiledCabbage 1y agoOh man that's good - next step create a PR to push it up stream! Everyone can benefit from its fixes.
- qazxcvbnmlp 1y agowhen this happens I do thew following 1) switch to a more expensive llm and ask it to debug: add debugging statements, reason about what's going on, try small tasks, etc 2) find issue 3) ask it to summarize what was wrong and what to do differently next time 4) copy and paste that recommendation to a small text document 5) revert to the original state and ask the llm to make the change with the recommendation as context
- rurp 1y agoThis honestly sounds slower than just doing it myself, and with more potential for bugs or non-standard code. I've had the same experience as parent where LLMs are great for simple tasks but still fall down surprisingly quickly on anything complex and sometimes make simple problems complex. Just a few days ago I asked Claude how to do something with a library and rather than give me the simple answer it suggested I rewrite a large chunk of that library instead, in a way that I highly doubt was bug-free. Fortunately I figured there would be a much simpler answer but mistakes like that could easily slip through.
- qazxcvbnmlp 1y agoit is slower than doing it yourself in the following scenarios 1) working in a language you are familiar in but has speed advantages in other areas 1) working in languages and or frameworks you are not familiar in 2) by documenting where the llm went wrong, you can append that to its rules and avoid the issue next time
- ziml77 1y agoYeah if it gets stuck and can't easily get itself unstuck, that's when I step in to do the work for it. Otherwise it will continue to make more and more of a mess as it iterates on its own code.
- nico 1y ago> 1) switch to a more expensive llm and ask it to debug You might not even need to switch A lot of times, just asking the model to debug an issue, instead of fixing it, helps to get the model unstuck (and also helps providing better context)
- enraged_camel 1y agoI've actually thought about this extensively, and experimented with various approaches. What I found is that the quality of results I get, and whether the AI gets stuck in the type of loop you describe, depends on two things: how detailed and thorough I am with what I tell it to do, and how robust the guard rails I put around it are. To get the best results, I make sure to give detailed specs of both the current situation (background context, what I've tried so far, etc.) and also what criteria the solution needs to satisfy. So long as I do that, there's a high chance that the answer is at least satisfying if not a perfect solution. If I don't, the AI takes a lot of liberties (such as switching to completely different approaches, or rewriting entire modules, etc.) to try to reach what it thinks is the solution.
- prmph 1y agoBut don't they keep forgetting the instructions after enough time have passed? How do you get around that? Do you add an instruction that after every action it should go back and read the instructions gain?
- enraged_camel 1y agoThey do start "drifting" after a while, at which point I export the chat (using Cursor), then start a new chat and add the exported file and say "here's the previous conversation, let's continue where we left off". I find that it deals with the transition pretty well. It's not often that I have to do this. As I mentioned in my post above, if I start the interaction with thorough instructions/specs, then the conversation concludes before the drift starts to happen.
- skerit 1y ago> I don’t think I’ve encountered a case where I’ve just let the LLM churn for more than a few minutes and gotten a good result. Is this with something like Aider or CLine? I've been using Claude-Code (with a Max plan, so I don't have to worry about it wasting tokens), and I've had it successfully handle tasks that take over an hour. But getting there isn't super easy, that's true. The instructions/CLAUDE.md file need to be perfect.
- nico 1y ago> I've had it successfully handle tasks that take over an hour What kind of tasks take over an hour?
- aprilthird2021 1y agoYou have to give us more about your example of a task that takes over an hour with very detailed instruction. That's very intriguing
- onlyrealcuzzo 1y ago> If it doesn’t solve an issue on the first or second pass, it seems to rapidly start making things up, make totally unrelated changes claiming they’ll fix the issue, or trying the same thing over and over. Sounds like a lot of employees I know. Changing out the entire library is quite amusing, though. Just imagine: I couldn't fix this build error, so I migrated our entire database from Postgres to MongoDB...
- butterknife 1y agoThanks for the laughs.
- mathattack 1y agoIt may be doing the wrong thing like an employee, but at least it's doing it automatically and faster. :)
- didgeoridoo 1y agoProbably had “MongoDB is web scale” in the training set.
- KronisLV 1y agoThat is amusing but I remember HikariCP in an older Java project having issues with DB connections against an Oracle instance. Thing is, no settings or debugging could really (easily) narrow it down, whereas switching to DBCP2 both fixed whatever was the stability issue in that particular pairing, as well as has nice abandoned connection tracking too. Definitely the quick and dirty solution that still has good results.
- PaulHoule 1y ago[flagged]
- Workaccount2 1y agoThey poison their own context. Maybe you can call it context rot, where as context grows and especially if it grows with lots of distractions and dead ends, the output quality falls off rapidly. Even with good context the rot will start to become apparent around 100k tokens (with Gemini 2.5). They really need to figure out a way to delete or "forget" prior context, so the user or even the model can go back and prune poisonous tokens. Right now I work around it by regularly making summaries of instances, and then spinning up a new instance with fresh context and feed in the summary of the previous instance.
- joshstrange 1y agoGreat term, I never put it into words but I feel this deeply. I rarely go back and forth more than 2-3 times with an LLM before ejecting to a new conversation. I’ve just been burned so much by old context informing the conversation later to my detriment. As soon as it gets something wrong I know the “rot” has set in and I need to start over (bringing over the best parts). But, OpenAI and friends should let me purge my questions and, more importantly, the LLM response from the chat. More often than not, it’s poisoning itself with bad ideas, flip-flopping, etc. I hate having to pick up and move to a new chat but if I don’t the conversation will only go downhill.
- codeflo 1y agoI wonder to what extent this might be a case where the base model (the pure token prediction model without RLHF) is "taking over". This is a bit tongue-in-cheek, but if you see a chat protocol where an assistant makes 15 random wrong suggestions, the most likely continuation has to be yet another wrong suggestion. People have also been reporting that ChatGPT's new "memory" feature is poisoning their context. But context is also useful. I think AI companies will have to put a lot of engineering effort into keeping those LLMs on the happy path even with larger and larger contexts.
- potatolicious 1y agoI think this is at least somewhat true anecdotally. We do know that as context length increases, adherence to the system prompt decreases. Whether that de-adherence is reversion to the base model or not I'm not really qualified to say, but it certainly feels that way from observing the outputs. Pure speculation on my part but it feels like this may be a major component of the recent stories of people being driven mad by ChatGPT - they have extremely long conversations with the chatbot where the outputs start seeming more like the "spicy autocomplete" fever dream creative writing of pre-RLHF models, which feeds and reinforces the user's delusions. Many journalists have complained that they can't seem to replicate this kind of behavior in their own attempts, but maybe they just need a sufficiently long context window?
- peacebeard 1y agoVery common to see in comments some people saying “it can’t do that” and others saying “here is how I make it work.” Maybe there is a knack to it, sure, but I’m inclined to say the difference between the problems people are trying to use it on may explain a lot of the difference as well. People are not usually being too specific about what they were trying to do. The same goes for a lot of programming discussion of course.
- heyitsguay 1y agoI've noticed this a lot, too, in HN LLM discourse. (Context: Working in applied AI R&D for 10 years, daily user of Claude for boilerplate coding stuff and as an HTML coding assistant) Lots of "with some tweaks i got it to work" or "we're using an agent at my company", rarely details about what's working or why, or what these production-grade agents are doing.
- alganet 1y ago> People are not usually being too specific about what they were trying to do. The same goes for a lot of programming discussion of course. In programming, I already have a very good tool to follow specific steps: _the programming language_. It is designed to run algorithms. If I need to be specific, that's the tool to use. It does exactly what I ask it to do. When it fails, it's my fault. Some humans require algorithmic-like instructions too. Like cooking a recipe. However, those instructions can be very vague and a lot of humans can still follow it. LLMs stand on this weird place where we don't have a clue in which occasions we can be vague or not. Sometimes you can be vague, sometimes you can't. Sometimes high level steps are enough, sometimes you need fine-grained instructions. It's basically trial and error. Can you really blame someone for not being specific enough in a system that only provides you with a text box that offers anthropomorphic conversation? I'd say no, you can't. If you want to talk about how specific you need to prompt an LLM, there must be a well-defined treshold. The other option is "whatever you can expect from a human". Most discussions seem to juggle between those two. LLMs are praised when they accept vague instructions, but the user is blamed when they fail. Very convenient.
- peacebeard 1y ago
- accrual 1y agoI've had some similar experiences. While I find agents very useful and able to complete many tasks on its own, it does hit roadblocks sometimes and its chosen solution can be unusual/silly. For example, the other day I was converting models but was running out of disk space. The agent decided to change the quantization to save space when I'd prefer it ask "hey, I need some more disk space". I just paused it, cleared some space, then asked the agent to try the original command again.
- vadansky 1y agoI had a particularly hard parsing problem so I setup a bunch of tests and let the LLM churn for a while and did something else. When I came back all the tests were passing! But as I ran it live a lot of cases were still failing. Turns out the LLM hardcoded the test values as “if (‘test value’) return ‘correct value’;”!
- bluefirebrand 1y agoThis is the most accurate Junior Engineer behavior I've heard LLMs doing yet
- ffsm8 1y agoMissed opportunity for the LLM, could've just switched to Volkswagen CI https://github.com/auchenberg/volkswagen https://github.com/auchenberg/volkswagen
- EGreg 1y agoThis is gold lol
- artursapek 1y agolmfao
- vunderba 1y agoI've definitely seen this happen before too. Test-driven development isn't all that effective if the LLM's only stated goal is to pass the tests without thinking about the problem in a more holistic/contextual manner.
- matsemann 1y agoReminds me of trying to train a small neural net to play Robocode ~10+ years ago. Tried to "punish" it for hitting walls, so next morning I had evolved a tanks that just stood still... Then punished it for standing still, ended up with a tanks just vibrating, alternating moving back and forth quickly, etc.
- nico 1y ago> Rather than fix the error, it decided to switch to another library I’ve had a similar experience, where instead of trying to fix the error, it added a try/catch around it with a log message, just so execution could continue
- mtalantikite 1y agoI had Claude Code deep inside a change it was trying to make, struggling with a test that kept failing, and then decided to delete the test to make the test suite pass. We've all been there! I generally treat all my sessions with it as a pairing session, and like in any pairing session, sometimes we have to stop going down whatever failing path we're on, step all the way back to the beginning, and start again.
- nojs 1y ago> decided to delete the test to make the test suite pass At least that’s easy to catch. It’s often more insidious like “if len(custom_objects) > 10:” or “if object_name == ‘abc’” buried deep in the function, for the sole purpose of making one stubborn test pass.
- akomtu 1y agoClaude Doctor will hopefully do better.
- Aeolun 1y agoHah, I got “All this async stuff looks really hard, let’s just replace it with some synchronous calls”. “Claude, this is a web server!” “My apologies… etc.”
- dylan604 1y ago> it encountered a build error. does this mean that even AI gets stuck in dependency hell?
- hhh 1y agoof course, i’ve even had them actively say they’re giving up after 30 turns
- reactordev 1y agoThis happens in real life too when a dev builds something using too much copy pasta and encounters a build error. Stackoverflow was founded on this.
- EnPissant 1y agoWhere you using Claude Code or something else? I've had very good luck with Claude Code not doing what you described.
- Wowfunhappy 1y ago> I don’t think I’ve encountered a case where I’ve just let the LLM churn for more than a few minutes and gotten a good result. I absolutely have, for what it's worth. Particularly when the LLM has some sort of test to validate against, such as a test suite or simply fixing compilation errors until a project builds successfully. It will just keep chugging away until it gets it, often with good overall results in the end. I'll add that until the AI succeeds, its errors can be excessively dumb, to the point where it can be frustrating to watch.
- civilian 1y agoYeah, and I have a similar experience watching junior devs try to get things working-- their errors can be excessively dumb :D
- insane_dreamer 1y agoI have the same experience. When it starts going down the wrong path, it won't switch paths. I have to intervene, put my thinking cap on, and tell the agent to start over from scratch and explore another path (which I usually have to define to get them started in the right direction). In the end, I'm not sure how much time I've actually saved over doing it myself.
- smeeth 1y agoI suspect its something to do with the following: When humans get stuck solving problems they often go out to acquire new information so they can better address the barrier they encountered. This is hard to replicate in a training environment, I bet its hard to let an agent search google without contaminating your training sample.
- Aeolun 1y agoI feel like they might eventually arrive at the right solution, but generally, interrupting it before it goes off on a wild tangent saves you quite a bit of time.
- ninetyninenine 1y agoI got an idea. Context compression. Once the context reaches a certain size threshold have the LLM summarize it into bulletpoints then start a new session with that summary as the context. Humans as well don’t remember the entire context either. For your case the summary already says tried library A and B and it didn’t work, it’s unlikely the LLM will repeat library A given that the summary explicitly said it was attempted. I think what happens is that if the context gets to large the LLM sort of starts rambling or imitating rambling styles it finds online. The training does not focus on not rambling and regurgitation so the LLM is not watching too hard for that once the context gets past a certain length. People ramble too and we repeat shit a lot.