7 ms·
Can't think of anything an LLM is good enough at to let them do on their own in a loop for more than a few iterations before I need to reign it back in.
by datpuz 1y ago
Can't think of anything an LLM is good enough at to let them do on their own in a loop for more than a few iterations before I need to reign it back in.
- CuriouslyC 1y agoThe main problem with agents is that they aren't reflecting on their own performance and pausing their own execution to ask a human for help aggressively enough. Agents can run on for 20+ iterations in many cases successfully, but also will need hand holding after every iteration in some cases. They're a lot like a human in that regard, but we haven't been building that reflection and self awareness into them so far, so it's like a junior that doesn't realize when they're over their depth and should get help.
- ariwilson 1y agoIs there value in adding an overseer LLM that measures the progress between n steps and if it's too low stops and calls out to a human?
- solumunus 1y agoAnd how does it effectively measure progress?
- NotMichaelBay 1y agoIt can behave just like a senior role would - produce the set of steps for the junior to follow, and assess if the junior appears stuck at any particular step.
- chongli 1y agoProducing the set of steps is the hard part. If you can do that, you don’t need a junior to follow it, you have a program to execute.
- abletonlive 1y agoIf this is true then we wouldn't have senior engineers that delegate. My suggestion is to think a couple more cycles before hitting that reply button. It'll save us all from reading obviously and confidently wrong statements.
- guappa 1y agoAI aren't real people… You do that with real people because you can't just rsync their knowledge. Only on this website of completely reality detached individuals such an obvious comment would be needed.
- abletonlive 1y agoSo...you don't think you can give LLMs more knowledge ?? You're the one operating in detached reality. The reality is that a ton of engineers are finding LLMs useful, such as the author. Maybe consider if you don't find it useful you're working on problems that it's not good at, or even more likely, you just suck at using the tools. Anybody that finds value out of LLMs has a hard time understanding how one would conclude they are useless and you can't "give it instructions because that's that hard part" but it's actually really easy to understand. The folks that think this are just bad at it. We aren't living in some detached reality. The reality is that some people are just better than others
- guappa 1y agoIf you have a dumb AI and a good AI, why not just use the good AI?
- TeMPOraL 1y agoSenior engineers delegate in part because they're coaxed into a faux-management role (all of the responsibilities, none of the privileges). Coding is done by juniors; by the time anyone gains enough experience to finally begin to know what they're doing, they're relegated to "mentoring" and "training" new cohort of fresh juniors. Explains a lot about software quality these days.
- CuriouslyC 1y agoI have actually had great success with agentic coding by sitting down with a LLM to tell it what I'm trying to build and have it be socratic with me, really trying to ask as many questions as it can think of to help tease out my requirements. While it's doing this, it's updating the project readme to outline this vision and create a "planned work" section that is basically a roadmap for an agent to follow. Once I'm happy that the readme accurately reflects what I want to build and all the architectural/technical/usage challenges have been addressed, I let the agent rip, instructing it to build one thing at a time, then typecheck, lint and test the code to ensure correctness, fixing any errors it finds (and re-running automated checks) before moving on to the next task. Given this workflow I've built complex software using agents with basically no intervention needed, with the exception of rare cases where its testing strategy is flakey in a way that makes it hard to get the tests passing.
- Xevion 1y ago>I have actually had great success with agentic coding by sitting down with a LLM to tell it what I'm trying to build and have it be socratic with me, really trying to ask as many questions as it can think of to help tease out my requirements. Just curious, could you expand on the precise tools or way you do this? For example, do you use the same well-crafted prompt in Claude or Gemini and use their in-house document curation features, or do you use a file in VS Code with Copilot Chat and just say "assist me in writing the requirements for this project in my README, ask questions, perform a socratic discussion with me, build a roadmap"? You said you had 'great success' and I've found AI to be somewhat underwhelming at times, and I've been wondering if it's because of my choice of models, my very simple prompt engineering, or if my inputs are just insufficient/too complex.
- CuriouslyC 1y agoI use Aider with a very tuned STYLEGUIDE.md and AI rules document that basically outlines this whole process so I don't have to instruct it every time. My preferred model is Gemini 2.5 Pro, which is definitely by far the best model for this sort of thing (Claude can one shot some stuff about as well but for following an engineering process and responding to test errors, it's vastly inferior)
- CuriouslyC 1y agoI don't think you need an overseer for this, you can just have the agent self-assess at each step whether it's making material progress or if it's caught in a loop, and if it's caught in a loop to pause and emit a prompt for help from a human. This would probably require a bit of tuning, and the agents need to be setup with a blocking "ask for help" function, but it's totally doable.
- p_v_doom 1y agoBruh, we're inventing robot PMs for our robot developers now? We're so fucked
- suninsight 1y agoYes it works really well. We do something like that at NonBioS.ai - longer post below. The agent self reflects if it is stuck or confused and calls out the human for help.
- vendiddy 1y agoI think they are capable of doing it, but it requires prompting. I constantly have to instruct them: - Go step by step, don't skip ahead until we're done with a step - Don't make assumptions, if you're unsure ask questions to clarify And they mostly do this. But this needs to be default behavior! I'm surprised that, unless prompted, LLMs never seem to ask follow-up questions as a smart coworker might.
- mkagenius 1y agoI built android-use[1] using LLM. It is pretty good at self healing due to the "loop", it constantly checks if the current step is actually a progress or regress and then determines next step. And the thing is nothing is explicitly coded, just a nudge in the prompts. 1. clickclickclick - A framework to let local LLMs control your android phone (https://github.com/BandarLabs/clickclickclick https://github.com/BandarLabs/clickclickclick)
- Groxx 1y agoThey're extremely good at burning through budgets, and get even better when unattended
- _kb 1y agoMaximising paperclip production too.
- mycall 1y agoIs that really true? I though there free models and $200 all you can eat models.
- nsomaru 1y agoThese tools require API calls which usually aren’t priced like the consumer plans
- adastra22 1y agoYeah they’re cheaper. I’ve written whole apps for $0.20 in API calls.
- monsieurbanana 1y agoWith which agent? What kind of apps? Without more information I'm very skeptical that you had e.g. Claude Code create a whole app (so more than a simple script) with 20 cents. Unless it was able to one-shot it, but at that point you don't need an agent anyway.
- adastra22 1y agoAider, Claude 3.7.
- datpuz 1y agoI've "written" whole apps by going to GitHub, cloning a repo, right clicking, and renaming it to "MyApp." Impressed?
- hbbio 1y agoThat's why in practice you need more than this simple loop! Pretty much WIP, but I am experimenting with simple sequence-based workflows that are designed to frequently reset the conversation [2] This goes well with Microsoft paper "LLMs Get Lost In Multi-Turn Conversation " that was published Friday [1]. - [1]: https://arxiv.org/abs/2505.06120 https://arxiv.org/abs/2505.06120 - [2]: https://github.com/hbbio/nanoagent/blob/main/src/workflow.ts https://github.com/hbbio/nanoagent/blob/main/src/workflow.ts
- eru 1y agoThe hope is that the ground truth from calling out to tools (like compilers or test runs) will eventually be enough keep them on track. Just like humans and human organisations also tend to experience drift, unless anchored in reality.
- loa_in_ 1y agoYou don't have to. Most of the appeal is automatically applying fixes like "touch file; make" after spotting a trivial mistake. Just let it at it.
- vidarh 1y agoThey've written most of the recent iterations of X11 bindings for Ruby, including a complete, working example of a systray for me. They also added the first pass of multi-monitor support for my WM while I was using it (restarted it repeatedly while Claude Code worked, in the same X session the terminal it was working in was running). You do need to reign them back in, sure, but they can often go multiple iterations before they're ready to make changes to your files once you've approved safe tool uses etc.
- JeremyNT 1y agoDefinitely true currently, which is why there's so much focus on using them to write real code that humans have to actually commit and put their names on. Longer term, I don't think this holds due to the nature of capitalism. If given a choice between paying for an LLM to do something that's mostly correct versus paying for a human developer, businesses are going to choose the former, even if it results in accelerated enshittification. It's all in service of reducing headcount and taking control of the means of production away from workers.