3 ms·
This has been my experience with a recent try to guide the LLM to a complete implementation of a small internal tool. I had in an hour what would have taken me
by fcatalan 2y ago
This has been my experience with a recent try to guide the LLM to a complete implementation of a small internal tool.
I had in an hour what would have taken me 4 or 5 to write.
But after that, it was an endless loop of the LLM adding logging code to find some bug and failing to fix it, only to add more logging code and ineffectual changes and so on.
The problem is that even after it's lost at sea, it's still answering in a completely confident and self assured tone, so when you decide to take matters in your hands you might be too far gone from sanity and have an unfixable mess in your hands.
I guess I can go back to where it strayed and retake it from there, but by now the experiment seems to be a failure.
- morsecodist 2y agoAt least in my experience as soon as something goes a little wrong it just gets worse from there. The more of it's confusion and contradictory information are in the chat history the worse it gets. It also has to make changes to the code so you accumulate these spurious changes and the problem gets more confusing. I've had some luck starting over with a new chat asking what is wrong but if that doesn't work I just assume I'm on my own.
- diggan 2y agoI've found that quality degrades really quickly after just the first reply, for some reason. They all seem heavily biased towards one-shot correct answers, and as you say, they go down the wrong path really quickly if you even get the first message slightly wrong. I tend to restart chats from the beginning pretty much all the time, because of this.
- iamflimflam1 2y agoI’ve also found this to be the case. Starting a new chat or in Cursor composer session puts things back on the right track. Also, prompting is really important. A lot of people just seem to think they have some kind of oracle - “fix the bug” - how is anything supposed to work from that?
- fragmede 2y agoyou're not graded on getting the LLM to output perfect code, the point is to get the code in git and PR'd. If your LLM tooling doesn't automatically commit to git so you can trivially go back to "where it strayed" you need to find a better tool. (My current favorite is aider) It's a tool not a person. When was the last time you got mad at a hammer for being smug?
- sebastiennight 2y agoTo drive the point home, hammers are quite smug. Mine always thinks it nailed it on the first try, and it's pretty hard-headed when you point out mistakes. If you can't work around those limitations, you're screwed.
- williamcotton 2y agoPart of the skill in using these tools is recognizing when it spins off the rails and backtracking immediately. Most of the time something can be gleaned from that wrong approach which can then guide further attempts.
- joshstrange 2y agoThis is my experience with Aider. When I first started using it, I turned off the auto git commits, but I’ve since turned them back on because they serve as perfect rollback points. My personal style is only commit once I have a feature fully working but with Aider it's best to have it commit after each exchange. I've gone 2-6 steps down a path before realizing this isn't going to work or the LLM is stuck in a loop. I just hard reset back to the first commit in that chain and either approach the task differently or skip it if it wasn't really that important.
- KronisLV 2y ago> But after that, it was an endless loop of the LLM adding logging code to find some bug and failing to fix it, only to add more logging code and ineffectual changes and so on. The problem is that even after it's lost at sea, it's still answering in a completely confident and self assured tone, so when you decide to take matters in your hands you might be too far gone from sanity and have an unfixable mess in your hands. I wonder how much better or worse things would get, if we took the human factor out of the loop. Give the LLM the ability to run tests and see the results, then iterate on its own output and branch off with different approaches, gradually increase the temperature etc. Maybe it’d turn out that you need 10 LLMs running in parallel for an hour to fix something, or perhaps even a 100 would never stumble upon a solution for a particular type of problem. And even then I wonder, whether it’d get better if you fed it your entire codebase or the codebases of the entire libraries or frameworks that you use (though at that point you’re either training it yourself or are selectively finding and feeding the correct bits not to exceed the context).
- renewedrebecca 2y agoBut why? What is there to be gained in all of this work around the inherent limitations of this technology?
- KronisLV 2y agoExploration of what’s possible and what’s not, identifying whether the weaknesses can or cannot be addressed. A bit like traditional autocomplete can help streamline familiarising oneself with various libraries, a clear step ahead when compared to just needing to dig through documentation as much. Maybe there’s a class of code problems that LLMs can be decent at solving, given the ability to iterate, verify solutions and what works or doesn’t, perhaps with 10x more compute than is utilized in the typical chat mode of interaction though.
- whattheheckheck 2y agoGet more people into computer science. Knuth said early on in his career he thought he needed to make the computer faster or cheaper but really it was about getting more users. Anyone can program. Or try to then learn about computer science
- almog 2y agoBack in early 2023 I tried to write a tool to do my taxes based on my broker CVS files. Since I wasn't familiar with how the data was structured, I let the LLM lead me while building this in incremental steps. The result was not just buggy, it simply failed to detect the relationships in the data (multiple somewhat implicitly embedded tables that needed to be joined). Even after I pointed this out, it failed to handle it, getting stuck in the same kind of loop you described. To this day, no LLM that I tried passed this task of leading the development while detecting the underlying structure of the data.