6 ms·
Looking at the prompts op has shared, I'd recommend more aggressively managing/trimming the context. In general you don't give the agent a new task without /cle
by friggeri 1y ago
Looking at the prompts op has shared, I'd recommend more aggressively managing/trimming the context. In general you don't give the agent a new task without /clearing the context before. This will enable the agent to be more focused on the new task, and decrease its bias (if eg. reviewing changes it has made previously).
The overall approach I now have for medium sized task is roughly:
- Ask the agent to research a particular area of the codebase that is relevant to the task at hand, listing all relevant/important files, functions, and putting all of this in a "research.md" markdown file.
- Clear the context window
- Ask the agent to put together a project plan, informed by the previously generated markdown file. Store that project plan in a new "project.md" markdown file. Depending on complexity I'll generally do multiple revs of this.
- Clear the context window
- Ask the agent to create a step by step implementation plan, leveraging the previously generated research & project files, put that in a plan.md file.
- Clear the context window
- While there are unfinished steps in plan.md:
-- While the current step needs more work
--- Ask the agent to work on the current step
--- Clear the context window
--- Ask the agent to review the changes
--- Clear the context window
-- Ask the agent to update the plan with their changes and make a commit
-- Clear the context window
I also recommend to have specialized sub agents for each of those phases (research, architecture, planning, implementation, review). Less so in terms of telling the agent what to do, but as a way to add guardrails and structure to the way they synthesize/serialize back to markdown.
- lukaslalinsky 1y agoEven better approach, in my experience is to ask CC to do research, then plan work, then let it implement step 1, then double escape, move back to the plan, tell it that step 1 was done and continue with step 2.
- nadis 1y agoThis is a really interesting approach as well, and one I'll need to try! Previously, I've almost never moved back to the plan using double escape unless things go wrong. This is a clever way to better use that functionality. Thanks for sharing!
- righthand 1y agoYou just convinced me that Llms are a pay-to-play management sim.
- bubblyworld 1y agoHeh, there are at least as many different ways to use LLMs as there are pithy comments disparaging people who do.
- righthand 1y agoI’m not disparaging it, just actualizing it and sharing that thought. If you don’t understand that most modern “tools” and “services” are gamified, then yes I suppose I seem like a huge jerk. The author literally talks about managing a team of multiple agents and Llm services requiring purchase of “tokens” is similar to popping a token into an arcade machine.
- hiAndrewQuinn 1y agoElectricity prices are also token-based, in a sense, yet most people broadly agree this is the best way to price them.
- righthand 1y agoI’m not following your point. What is the disagreement about Llm pricing?
- deleted 1y ago[deleted]
- dgfitz 1y agoI read a quote on here somewhere: "Hacker culture never took root in the AI gold rush because the LLM 'coders' saw themselves not as hackers and explorers, but as temporarily understaffed middle-managers"
- 1y ago
- shafyy 1y agoDude, why not just do it yourself if you have to micromanage the LLM this hardcore?
- nadis 1y agoOP here, this is great advice. Thanks for sharing. Clearing context more often between tasks is something I've started to do more recently, although definitely still a WIP to remember to do so. I haven't had a lot of success with the .md files leading to better results yet, but have only experimented with them occasionally. Could be a prompting issue though, and I like the structure you suggested. Looking forward to trying! I didn't mention it in the blog post but actually experimented a bit with using Claude Code to create specialized agents such as an expert-in-Figma-and-frontend "Design Engineer", but in general found the results worse than just using Claude Code as-is. This also could be a prompting issue though and it was my first attempt at creating my own agents, so likely a lot of room to learn and improve.
- deleted 1y ago[deleted]
- giancarlostoro 1y ago> Looking at the prompts op has shared, I'd recommend more aggressively managing/trimming the context. In general you don't give the agent a new task without /clearing the context before. This will enable the agent to be more focused on the new task, and decrease its bias (if eg. reviewing changes it has made previously). My workflow for any IDE, including Visual Studio 2022 w/ CoPilot, JetBrains AI, and now Zed w/ Claude Code baked in is to start a new convo altogether when I'm doing something different, or changing up my initial instructions. It works way better. People are used to keeping a window until the model loses its mind on apps like ChatGPT, but for code, the context Window gets packed a lot sooner (remember the tools are sending some code over too), so you need to start over or it starts getting confused much sooner.
- nadis 1y agoI've been meaning to try Zed but haven't gotten into it yet; it felt hard to justify switching IDEs when I just got into a working flow with VS Code + Claude Code CLI. How are you finding it? I'm assuming positive if that's your core IDE now but would love to hear more about the experience you've had so far.
- lukaslalinsky 1y agoIf you are Claude Code user, you will likely not enjoy the version integrated into Zed. Many things are missing, for example, no slash commands. I use Zed, but still run Claude Code in the terminal. As an editor, Zed is excellent, especially as a Vim replacement.
- nadis 1y agoOh interesting, that’s good to know. Thank you. I might try that combination - I like using the Claude Code CLI so hopefully less of a painful transition.
- giancarlostoro 1y agoI guess I am a weirdo because I never used Claude Code until Zed added it. At least Zed has a built in terminal emulator to boot.
- jngiam1 1y agoI also ask the agent to keep track of what we're working on in a another md file which it save/loads between clears.
- jimbo808 1y agoThis seems like a bit of overkill for most tasks, from my experience.
- dingnuts 1y agoit just seems like a lot of work when you could just write the code yourself, just a lot less typing to go ahead and make the edits you want instead of trying to guide the autocorrect to eventually predict what you want from guidelines you also have to generate to save time like I'm sorry but when I see how much work the advocates are putting into their prompts the METR paper comes to mind.. you're doing more work than coding the "old fashioned way"
- lo5 1y agoit depends on the codebase. if there's adequate test coverage, and the tests emit informative failures, coding agents can be used as constraint-solvers to iterate and make changes, provided you stage your prompts properly, much like staging PRs. claude code is really good at this.
- devingould 1y agoI pretty much never clear my context window unless I'm switching to entirely different work, seems to work fine with copilot summarizing the convo every once in a while. I'm probably at 95% code written by an llm. I actually think it works better that way, the agent doesn't have to spend as much time rereading code it had previously just read. I do have several "agents" like you mention, but I just use them one by one in the same chat so they share context. They all write to markdown in case I do want to start fresh if things do go the wrong direction, but that doesn't happen very often.
- loudmax 1y agoI wouldn't take it for granted that Claude isn't re-reading your entire context each time it runs. When you run llama.cpp on your home computer, it holds onto the key-value cache from previous runs in memory. Presumably Claude does something analogous, though on a much larger scale. Maybe Claude holds onto that key-value cache indefinitely, but my naive expectation would be that it only holds onto it for however long it expects you to keep the context going. If you walk away from your computer and resume the context the next day, I'd expect Claude to re-read your entire context all over again. At best, you're getting some performance benefit keeping this context going, but you are subjecting yourself to context rot. Someone familiar with running Claude or industrial-strength SOTA models might have more insight.
- kookamamie 1y agoCC absolutely does not read the context again during each run. For example, if you ask it to do something, then revert its changes, it will think the changes are still there leading to bad times.
- catlifeonmars 1y agoWhen you say “revert its changes” do you mean undo the changes outside of CC? Does CC watch the filesystem?
- 1y ago
- raducu 1y agoIn 2025, does it make any difference to tel the LLM "you're an expert/experienced engineer?"
- adastra22 1y agoYes. The fundamental reason why that works hasn’t changed.
- hirako2000 1y agoHow has it ever worked. I have thousands of threads with various LLMs, none have that role play cue, yet the responses always sound authoritative and similar to what one would find in literature written by experts in the field. What does work is to provide clues for the agent to impersonate a clueless idiot on a subject, or a bad writer. It will at least sound like it in the responses. Those models have been heavily trained with RLHF, if anything today's LLMs are even more likely to throw authoritative predictions, if not in accuracy, at least in tone.
- pwython 1y agoI also don't tell CC to think like expert engineer, but I do tell it to think like a marketer when it's helping me build out things like landing pages that should be optimized for conversions, not beauty. It'll throw in some good ideas I may miss. Also when I'm hesitant to give something complex to CC, I tell that silly SOB to ultrathink.
- Fuzzwah 1y agoEvery LLM has role play cues in their system prompts: https://github.com/0xeb/TheBigPromptLibrary/tree/main/SystemPrompts https://github.com/0xeb/TheBigPromptLibrary/tree/main/System...
- adastra22 1y agoEvery LLM you have ever used has this role play baked into its system prompt.
- enraged_camel 1y agoThis is overkill. I know because I'm on the opposite end of the spectrum: each of my chat sessions goes on for days. The main reason I start over is because Cursor slows down and starts to stutter after a while, which gets annoying.
- fizx 1y agoCursor when not in "MAX" mode does its own silent context pruning in the background.
- philipp-gayret 1y agoSince reading I can --continue I do the same. If I find it's missing something after compressing context I'll just make it re-read a plan or CLAUDE.md
- nadis 1y agoThat’s a solid approach and one I hadn’t thought of myself. Thank you!
- nadis 1y agoClaude auto-condenses context, which is both good/bad. Good in that it doesn't usually get super slow, bad in that sometimes it does this in the middle of a todo and then ends up (I suspect) producing something less on-task as a result.
- antihero 1y agoThis sounds like more effort than just writing the code.
- Jweb_Guru 1y agoIt is.
- dotancohen 1y agoUsually, managing a development team is more work than just writing the code oneself. However, managing a development team (even if that team consists of a single LLM and yourself) means that more work can be done in a shorter period of time. It also provides much better structure for ensuring that tests are written, and that documentation is written if that is important. And in my experience, though not everybody's experience I understand, it helps ensure a clean, useful git history.
- R0m41nJosh 1y agoI have been reluctant to use AI as a coding assistant though I have installed claude code and bought a bunch of credits. When I see comments like this I genuinely asking what's the point. Are you sure that going through all of these manipulation instead of directly editing the source code makes you more productive? In which way? Not trolling, true question.
- yomismoaqui 1y agoTime is the answer here, if this dance is 2 hours and implementing it by hand is 8 hours you have won. Also while Claude Code is crunching floats you can do other things (maybe direct another agent instance)
- lukaslalinsky 1y agoYears ago, I was joking with my colleagues that I'm living two weeks ahead, writing the present day code is a chore, thinking about the next problems is more important, so that when the time comes to implement them, I know how. I don't have much time to code these days, but I still have the ability to think. Instead of doing the chore myself, I now delegate it to Claude Code. I still do coding occasionally, usually when it's something hard that I know AI will mess up, but in those instances, I enjoy it.
- 1oooqooq 1y agoand it will completely ignoring the instructions because user input cannot afect it, but it will waste even more context space to fool you that it did.