5 ms·
The gap between "useful chatbot" and "useful agent" is way bigger than people realize. A chatbot can be wrong 10% of the time and still help you. An agent that'
by vishalkundar 3mo ago
The gap between "useful chatbot" and "useful agent" is way bigger than people realize. A chatbot can be wrong 10% of the time and still help you. An agent that's wrong 10% of the time is sending bad emails and making wrong API calls with no one checking.
- stbenjam 3mo agoThe gulf is bridgeable. The problem is that a lot of people are building agents without strong enough judgment layers around them. Work that can be verified with reasonable accuracy are the sweet spot right now.
- Avicebron 3mo agoHow many of these layers are just trying to rediscover/rebuild the idempotence of code?
- mikebs1 3mo ago[flagged]
- ben_w 3mo ago> The gulf is bridgeable. Only with an LLM that's actually at agent-quality. If "useful chatbot" and "useful agent" are two rungs on a ladder, the rung before them is "useful autocomplete". Autocomplete that only gets the next token right 90% of the time won't give you compiling code.
- wonnage 3mo agoThis is much harder than it sounds. Most techniques I’ve seen end up using separate agents to do the planning, implementation, and judging. The elaborate workarounds you have to build to help an agent which fundamentally doesn’t know what it’s doing reminds me of this old blog post about TDD: https://pindancing.blogspot.com/2009/09/sudoku-in-coders-at-work.html?m=1 https://pindancing.blogspot.com/2009/09/sudoku-in-coders-at-... IMO present technology is tailored for an experienced developer to give agents manageable tasks that can be one-shot. The marketing right now reminds me of the 90s when AskJeeves promised natural language search when the technology was fundamentally still stuck in keyword search, and learning to craft a search query for Google is today’s prompt engineering
- csomar 3mo agoThe problem is that with text/code, judgement is hard. Here is what it looks like for physical activity: https://www.youtube.com/shorts/lK7TjujKQLw https://www.youtube.com/shorts/lK7TjujKQLw It's hard to see how that it's not useful at best and could be a disaster for any unsupervised use.
- mikebs1 3mo ago[flagged]
- skybrian 3mo agoI see this as the gap between an general-purpose agent and a coding agent. A coding agent can imagine something to be true, test it, discover that it's wrong, and recover. But if you go beyond what can be tested easily, asking the agent to do real work rather than writing a patch, imagining things to be true is a problem.
- steveBK123 3mo agoThis to me is the big leap from being good at coding to being good at many other tasks. Coding could be treated as a low stakes (time & money consequences for retries) closed loop system where most other tasks cannot. If it screws up booking your flight/hotel room, how does the agent verify this, and even if it verifies.. there is an actual cost to changes/cancellations. Similar with agentic e-commerce, lots of ability to screw that up and just seems ripe for fraud / being picked off by bad actors.
- steveBK123 3mo agoTo reply to myself here.. I can STILL replicate this behavior in Google AI summaries 10% of the time: "is <SOMEPLANT> ok for cats" to which it replies: "Yes, <SOMEPLANT LONG SCIENTIFIC NAME VERBOSE PHRASING> is toxic for cats" The other one going around this weekend: "how long hot dogs on grill" Summary: "The hot dogs on your grill are likely around 5-6 inches long .. " So scale this category of error to unsupervised agents with access to your credit card.
- argee 3mo agoThis repros nearly 100% of the time on most LLMs, even the most advanced ones: https://share.gemini.google/u9NwYu7lbgxe https://share.gemini.google/u9NwYu7lbgxe
- imajoredinecon 3mo agon=1 but I gave this to Sonnet 5 medium effort (free model) and it had no trouble with it