3 ms·
I’ve observed exactly these patterns with Opus and Fable as well - for example, forgetting that they can edit files and instead use python scripts as a patching
by buildbot 24d ago
I’ve observed exactly these patterns with Opus and Fable as well - for example, forgetting that they can edit files and instead use python scripts as a patching tool…
- jaapz 24d agoWhy would you use a constrained edit tool when you are also allowed to use the complete power of python?
- oblio 24d agoWhy even offer the edit tool in that case? Also, what kind of editing could they possible do what wouldn't be possible with POSIX ed?
- arcanemachiner 24d agoYou can chain a lot more commands together with this technique than with a single Edit tool call.
- oblio 24d agoThe funny thing is that... POSIX ed is composable :-) You can do a gazillion edits with it in one shot. Of course, LLM edit tools are probably small bits of their custom code, I just find it funny. I wonder if it's a desire for certain technical characteristics that require custom code or just a lack of info on basic tools. Heck, if it's about platform availability, using an LLM to port ed to Windows (for example) should be trivial[1]. * * * [1] And there are probably a million existing ports. Also, sed, ex, vi, whatever.
- zarzavat 24d agoBecause the complete power of Python also includes the power to fuck things up.
- rootlocus 24d agoSimple is better than complex Complex is better than complicated Or something, I don't remember...
- Burebista 24d ago... simply the best, better than all the rest (Tina Turner)
- oblio 24d agoUșor, Burebista :-)
- thewhitetulip 24d agoThat was pre LLM. Now everything is a prompt that you type into AI lol
- oefrha 24d agoHave you ever counted the number of times Claude fucked up quoting/escaping and had to issue a corrected tool call? Or get stuck in some tricky quoting situation for two minutes, throwing a couple piles of shit at the wall to see what sticks. IIRC I’ve even seen it eventually using the edit tool out of frustration once.
- dools 24d agoHaving an agent edit 100 files means the job will definitely get done correctly. When it writes a script to bulk edit things it fucks up and spends ages debugging their script.
- d5lt5 24d agoSounds like you don't have enough experience with coding agents. Deterministic scripts must always be preferred instead of LLM tool calls. In fact, you should instruct your agents to write code to execute instead of letting them call tools.
- weird-eye-issue 24d agoSounds like you completely lack all reading comprehension ability LLMs sometimes like to execute one-off Python scripts to make edits to files rather than just calling the edit tool directly. Both are tool calls so saying that you should have it write code instead of doing tool calls makes no sense because writing code is a tool call for it...
- d5lt5 24d ago[flagged]
- arcanemachiner 24d agoThey're talking about writing a file with a harness-native Edit tool. They're saying the agents aren't doing that, but are using ad-hoc methods of writing the files. (My agents seem to prefer see these days.)
- d5lt5 24d agoWhy do you think your agents prefer to create scripts instead of doing tool calls these days? I wonder, is it easier to modify a script that agent wrote before to satisfy your prompt, or is it easier to write a new one from scratch each time a retry happens? Are input tokens more expensive than output tokens?
- weird-eye-issue 24d agoMy god. They are not reusing the scripts. They are adhoc, inline Python scripts just used to make a single edit. You seem to fundamentally not understand what everyone else is talking about
- llama-for3ver 24d agothis is intentional, afaik agents do better with python and alike than the harness tooling.
- exceptione 24d agoUsing python or any other stone-age approach for search and replace is stupid when your language provides you with a complete, fully typed AST, like .NET does.
- whstl 24d agoThis is an instruction by the harness. It re-injects the prompt every other message, so that's why it "forgets" to use the Edit tool.
- nvch 24d agoThe harness instructs them to behave this way. Also this approach saves tokens. The scripts allow to edit files in bulk, and most of the session cost is in cache reads (e.g. for 300K context each command costs the same as 30K input tokens).
- dools 24d agoMy harness forbids it, they end up spending time debugging their scripts
- TeMPOraL 24d agoAlso it's the only way that makes sense when you need to work with big files, or large amount of files, or documents that look small when fetched through a RAG tool, but then you read one and get hit with couple megabytes of base64-encoded binary data you didn't expect because RAG tool stripped out embedded images... Ask me how I know. Or don't. I have a standing rule for all agents warning about that failure mode (and related, doing `ls` in `/tmp` and few other directories that like to accumulate files by the hundreds..)
- lelanthran 24d ago> The harness instructs them to behave this way. Also this approach saves tokens. The scripts allow to edit files in bulk, and most of the session cost is in cache reads (e.g. for 300K context each command costs the same as 30K input tokens). I understand the reasoning, but at that point wouldn't the LLM be better off creating `sed` commands and executing those? I mean, if it's already executing Python, it can literally do anything to the environment, so using `sed` is at least as safe, with a bonus that it (or a subagent, or a human) can double-check the intention with the sed script and flag incorrect or missing changes.
- chickensong 24d agoI've experimented quite a bit with giving agents python vs sed + awk. They make mistakes with both, a lot. The only thing that has stood out is that agents reach for python too quickly if it's available, and that awk causes the least problems, while sed might take several attempts to get results, similar to python.
- ZeWaka 24d agoI use AST replacers, much more reliable.