6 ms·
What really burns tokens is sub agents. I once gave Claude Code a pretty big task, and it immediately launched 7 sub agents which burned through my budget befor
by mcv 3mo ago
What really burns tokens is sub agents. I once gave Claude Code a pretty big task, and it immediately launched 7 sub agents which burned through my budget before even one of them was finished. Tried again 5 hours later: same result.
If I let the main agent do the same task sequentially, it was no problem at all. I don't know if it's really just communication and orchestration that makes sub agents so inefficient, or if Anthropic figured that most people using sub agents pay per token on a big corporate account, so this is an easy way to make more money from tokenmaxxers.
- thejazzman 3mo agofor subagents to be cheap/effective, you have to specify the size of those subagents; i.e. right now by default 5.6-sol spawns many 5.6-sol subagents. 5.4-mini as subagent saves me tons of tokens. 5.6-sol audits the work before accepting it, so there's not really a quality issue.
- qpricjalcbeu 3mo agoAnd in my experience the sub agent performance is usually worse than just a single agent.
- tudelo 3mo agoI find it useful for code reviews (spawn a subagent with minimal/no context to review X commit). Of course, this is more or less a shortcut that could be done with a seperate agent. Another use is multiple reviews at once if tokens are not an issue, with seperate "personas" or focuses. As far as implementation goes I have not seen any major usecase.
- ryanchants 3mo agoYeah, my personal workflow has different reviewers for codebase(patterns, code cleanliness, etc), frontend, security, product fit, etc. So they spawn as separate subagents. Both so that they stay limited to their role, and so they don't have preconceived notions about the implementations. It's a bit heavy-handed but works for me.
- retired 3mo agoDid it deploy five AWS m8g.12xlarge instances?
- a_c 3mo agoEvery subagent send the same ~30k system prompts. If you are using fable/opus, that's easily 30% of a 5-hour window for 7 subagent, before doing any work
- megous 3mo agoIf it's always the same prompt, can't they have it pre-cached globally for all?
- erikus 3mo agoI'm pretty sure the system instructions are a function of your environment and not the same universally. That said, there should be a finite number of branches so still cacheable.
- megous 3mo agoSystem specific stuff is probably quite limited, it can be a short dynamic segment at the end of the system prompt, perhaps.
- a_c 3mo agoThe system behaviour is totally up to anthropic's discretion. Its current behaviour is verifiable. In claude code, spawn a subagent with 1. Agent("Test") 2. look at your token usage 3. Repeat a few times I didn't check again as I type this message but am somewhat sure subagent doesn't cache system prompt as of maybe last week
- micw 3mo agoI recently did a few tests. And always the same prompt has been cached properly.
- ricardobeat 3mo agoCache is usually not shared between agents - they can have different base prompts, tools, and be an entirely different model.
- btown 3mo agoAs a counterpoint: in a complex project, Fable's "curiosity" may be exactly what you want for an exploration and planning stage - not just for the orchestrator that turns your prompt into different angles with which to explore, but for each subagent whose task is to search the codebase for one of those "angles." If you truly want no stone unturned, letting those subagents spawn their own discoveries, and recursively grow the surface area of the inquiry, then it's quite reasonable to want Fable throughout. That said, if your project is "do this well-planned thing on a bunch of things in parallel" then you should absolutely be instructing to have subagents "step down" to less curious models. Their output may well be more cohesive as a result!
- adastra22 3mo agoThe curiosity is inefficient though. So many times I have to stop the agent and tell it to just fucking write the code and try compiling it. Otherwise it will fill its entire context tracing through the program logic to derive from the code itself whether the thing it is about to do would work. It completely fails to notice it can just… try.
- ACCount37 3mo agoIt's tuned for the kinds of tasks where "just try" doesn't get good results. A major complaint with AI code was that AIs struggle with complex codebases, don't respect existing conventions, reinvent functionality multiple times over, etc. So, newer high end AIs are tuned with the "explore/exploit" dial turned towards "explore". You could probably get it to do things "quick and dirty" with prompting, but that, of course, requires prompting for it.
- giovannibonetti 3mo agoPerhaps what is missing is a better memory/caching layer to avoid doing the same for explorations over and over again.
- hoppp 3mo ago
- wongarsu 3mo agoSub agents each have to read part of your code base again to get enough context for the task. And if they take too long, your orchestrator's context is no longer in cache so you pay full price for that again once the subagents finish If you do it sequentially you only read those files approximately once, and everything hits the same prefix cache
- EMM_386 3mo agoYes but one of the key things about subagents is they keep all of their tool calls and exploration out of the parent context. If you plan on continuing on in the parent, and aren't going to necessarily be touching the systems the other agents are exploring, it can be worth it. It's useful in certain situations where the parent context may need the "10,000 foot" view of something without going back in there. But subsystem-specific AGENTS.md/CLAUDE.md files are still superior and accomplish the same thing. The problem with those is they can become stale.
- hidelooktropic 3mo agoThey are just making the point that it makes sense that subagents would use more tokens because they have none of the parent's context.
- drob518 3mo agoRight, so it’s a trade off between contexts. There are two reasons to use subagents, parallelism and tailoring of context. For the second, there is the “personality” of the subagents as well as how much context is injected from the main agent. Ignoring the personality, you ideally want the injected context to be small and focused on a single task so the subagent doesn’t get distracted. You want the main agent to be orchestrating all the subagents, but not reading all the same files they are reading, otherwise you’ll be paying for the same tokens in multiple contexts. IMO, this is where prompt engineering comes in, to be able to guide the main agent as to where subagents are desired and where not.
- hedgehog 3mo ago
- ValentineC 3mo ago> What really burns tokens is sub agents. I once gave Claude Code a pretty big task, and it immediately launched 7 sub agents which burned through my budget before even one of them was finished. Tried again 5 hours later: same result. Probably because the general purpose subagents inherit the parent model. I tell Claude explicitly to use Explore subagents, which use Haiku only, now.
- CjHuber 3mo ago> Probably because the general purpose subagents inherit the parent model only if you don't specify which model should be used
- killix 3mo ago[flagged]
- joshcartme 3mo agoThey changed it with the release on July 1 Explore now inherits the model, it isn't always haiku. https://code.claude.com/docs/en/changelog#2-1-198 https://code.claude.com/docs/en/changelog#2-1-198 > The built-in Explore agent now inherits the main session’s model (capped at opus) instead of running on haiku
- ValentineC 3mo agoUrgh, thanks for the heads up. I guess I need to be explicit about my choice of model now.
- joshcartme 3mo agoYeah, I was surprised. I had Claude make a skill to extract the explore agent from Claude Code, but set the model back to haiku. Here it is if it's helpful: https://gist.github.com/joshcartme/dd71df7b4c51c356760b28d7f383dff2 https://gist.github.com/joshcartme/dd71df7b4c51c356760b28d7f...
- onlyrealcuzzo 3mo agoThis is why I happily use Codex. I run it basically 24/7 on a ~500k line repo, and only rarely run out of quota before the end of the week. My experience with Claude Code was very good until about 2.5 months ago, and then it suddenly turned unbelievably terrible for me. I have not and will hopefully never look back. I still have PTSD from how ungodly terrible it was that last week of using it.
- peterlk 3mo agoCan you be more specific about what “unbelievably terrible” means?
- rendx 3mo ago> I still have PTSD from how ungodly terrible it was Please, for the sake of everyone suffering from actual PTSD: Don't. It's hard enough already for victims to communicate what difficulties they are facing without people watering down terminology like that.
- adaml_623 3mo agoThey have Coder PTSD or CPTSD.... Is that a better acronym??? Sorry just teasing.
- searealist 3mo agoPlease don't act as the hyperbole police. People exaggerate all the time (I'm starving, etc). It's normal, and you are being a jerk to call them out.
- rendx 3mo agoI am asking them to reconsider and reflect on what that kind of language use does. You're the one reading it as "calling them out". How else are we supposed to learn from each other, voice our opinions, point out our mistakes to each other? For me, this is communication. And currently 8 upvotes seem to agree with me and my request. Feel free to ignore it, or consider it, for your own use of language. But, sorry, to me, you're the one acting like a jerk and trying to "police", not me.
- adastra22 3mo ago--disallowedTools Task
- beezlewax 3mo agoSpawning a bunch of agents seems to happen randomly. I almost never want this.
- kadoban 3mo agoI think there's some setting to restrict the number of them, or maybe turn them off. Doesn't happen for me ~ever and it's not my $$ (work) so I haven't really looked at it much.
- alansaber 3mo agoSuch is the nature of tool use
- mcv 3mo agoIn my CLAUDE.md I put: > CRITICAL: Do NOT spawn sub-agents for any reason. Perform all work in the main session. If a task is too large, ask me to break it down manually. > This is a big task, and can easily get too large. However, sub-agents make the situation worse, and eat through our token budget way too fast. Do not use them. > Take on manageable tasks. Don't try to do everything at once. When you start on a big task, break it down into smaller tasks, and make sure you finish each task before starting on the next one. Or actually Claude put it there for me. Maybe it's a bit much, but it seems to work.
- nomel 3mo agoIf there's some "find the file" task, using full context for that isn't ideal.
- reinitctxoffset 3mo agoSubagents with a fat tailed latency distribution completely masks the trough filling that puts the most downwards pressure on per-token COGS. This is why the subscription plans are forced through the harness (the "OpenClaw Wars"): it creates a false equivalence in the minds of many customers between API tokens (latency sensitive, easy to measure) and Claude Code tokens (remnant backfill to stay to the right of the roofline, marginal cost often zero). Selling sausage as sirloin is a great business if people go for it. And there's nothing inherently wrong with spot pricing, as long as you're honest about it...
- leptons 3mo agoIt's in the best interest for AI companies to gobble up tokens. I feel like every new release - Fable, etc - is just a way to extract more tokens/money.
- drob518 3mo agoOf course it is. How could it be anything different? Clearly, that’s how these companies make money.
- twelve40 3mo agoit's a very handwavey way to "explain" anything. Yes, they make money. But they have competition. And if someone runs out of tokens and switches to deepseek or just goes for a friggin hike in the woods, that does not benefit them. If they get a public image of a ripoff that burns all shit on trivial tasks, that does not do them good either. So there is a limit to this "companies make money" thing.
- drob518 3mo agoSure, fair enough. Clearly, if they increase costs by too much, people will go to their competitors, but those competitors also make money selling tokens, so the whole industry is incentivized to inflate token consumption up to the point of driving people to the competition. And nobody is incentivized to reduce token count. In fact, the one model with great price/performance is Deepseek v4 Flash and I suspect that they are subsidizing it deeply to get access to everyone’s prompts for training. We may find that they raise prices on the next version (v5) after they’ve mined the user data.
- leptons 3mo agoAny AI service that people (and to some extent companies) can afford to pay for today is being heavily subsidized. Will that last forever? I really don't know how those economics work, but I know that bubbles do burst having lived through the dot com burst in 2000. And I know this current one is going to hurt if/when it bursts.
- xhrpost 3mo agoFor a while everyone was saying sub agents is how you save tokens, use lower quality models with limited context to do simple parts of the job after a smart planning agent has put it all in place. Is that no longer true or is this just the result of sub agent being used at the wrong time?
- duxup 3mo agoIt’s funny too because I’ll ask fairly simple things and it’s fine, similarly simple question might spin up a bunch of sub agents and I don’t know why…. I feel like maybe it could have asked for clarification or something rather than go and try to calculate all the digits of pi all of a sudden.
- joshcartme 3mo agoThey did recently change it so the default explorer agent inherits the session agent (capped at Opus). Before Explore was always haiku. I had Claude write a skill that extracts the built in Explorer agent skill, and then writes an identical Explore agent that uses Haiku
- viccis 3mo agoSame for me. I never use them. I use Fable on highest effort to plan things and then record the plan in tickets. I use Kata, which is CLI and agent oriented, but I suppose Jira or other systems would work too. I tell it to put enough context in each ticket to on-board a fresh coding agent to implement it. Then I just do /goal, telling to to run `kata ready` to get new tickets to work and continue until they're all closed according to acceptance criteria or until they're blocked on actions from me. I need to play around with getting it to switch to smaller models (or spawning 1 subagent) to do ticket implementation and then auto compact after each. Either way, it results in really easy workflows and uses very few tokens compared to the built in subagent flows that doing this completely avoids.
- Scrapemist 3mo agoVery interesting approach. Thanks for sharing.
- nvch 3mo agoI had learnt that trick, so now I explicitly disallow Fable subagents. Yesterday, I wanted to review a complex piece after a large refactoring, and requested a review plan beforehand. The first step was 8 agents + one more to verify the findings (all Fable). Looks good, approved. The verification step turned into an attempt to throw a party with 41 Fable verifiers. It will find a way.
- greenavocado 3mo agoDon't do that; limit concurrency
- lemagedurage 3mo ago"LGTM" That'll be $50 — please.
- deleted 3mo ago[deleted]
- hoppp 3mo agoProbably both. The default subagent orchestration is designed for infinite pockets. Maybe when they realize there is need to change this they come up with a more configurable interface for us mere mortals who can't afford to gamble their house on a pay as you go subscription.
- brianwawok 3mo agolol I asked fable to help me estimate my TAM and it launched 102 agents and blew my $120 quota in 6 minutes. I do realize I can limit the agent count , hah
- hinkley 3mo agoThere is a negative incentive to fix problems that result in customers picking a more expensive plan to work around it. There are probably several engineers who have ideas about fixing this and they get apathy from many people and obstruction from a few, and sometimes active hostility by a manager somewhere in the chain. The best you can do in such an environment is seek to introduce new features at the top tier, and then pull old features down the stack as the cost of those features has been amortized out, or to hurt your competitors by raising the ladder.
- avinoth 3mo agoI’ve had similar experiences. I now have an explicit line in AGENTS.md to not use subagents unless explicitly requested. It also helps that for the tasks that are big enough to benefit from subagents are also the ones with high chances of going off-rails and/or a poor review phase. I’d rather do the orchestrator role and that way I can split up the review phase in a much more manageable chunk.
- vinnymac 3mo agoYesterday I gave Claude Fable a difficult task. It then proceeded to spawn 415 agents. It got it done, but damn was it expensive.
- deleted 3mo ago[deleted]
- nihsett 3mo agoThey optimized it to burn more token in the recent months I feel. I made a small ~100 line change to a codebase by hand and threw claude at it to review. It spawned several sub-agents and burnt a ton of tokens. I guess the word 'review' now triggers some sort of in-built skill or something. It's absurd how rapid enshittification is taking over.
- k3liutZu 3mo agoIndeed it feels like I do the same work, ask the same questions, get the same result. But somehow the cost has doubled in the last few months.
- lemagedurage 3mo agoTrue. For Claude Code, I disabled explore subagents globally by adding this to ~/.claude/settings.json: "permissions": { "deny": [ "Task(Explore)" ] }
- mcv 3mo agoIs Explore the only thing subagents are ever used for?
- derintegrative 3mo agoShould be "Agent(Explore)"
- lemagedurage 3mo agoYou're right, looks like they changed it, though Task should still work. > In version 2.1.63, the Task tool was renamed to Agent. Existing Task(...) references in settings and agent definitions still work as aliases. https://code.claude.com/docs/en/sub-agents https://code.claude.com/docs/en/sub-agents
- Computer0 3mo agoSubagents are quite inefficient and the lossy context transfer between them does lead to more cost and more waiting. However I have found it to produce more reliable output, whether that is worth it for a given task has been a consideration.
- hgoel 3mo agoI like to use subagents a lot, but I find them to be most useful when explicitly specified. E.g. "assign these tasks to 2 Sonnet, 2 Opus and 1 Fable subagent". Helps keep allocation consumption under control.
- vitorgrs 3mo agoNot only a Claude Code issue. Started using OMP with GPT 5.6, and gave up, it loves to use subagents, and it's basically unusable subagents with GPT 5.6 Sol there with Plus limits.
- _s_a_m_ 3mo agoi have globally disabled subagents for claude. otherwise one prompt ended my Pro account
- fearmerchant 3mo agoAgreed. The issue is that when working 1:1 you get a feel for how many tokens are being burned but the subagent spawn could be 3 or in one cases it spawned 171 to verify something. The latter was unexpected and burned through my token budget.
- rajeevbakshi 3mo ago[flagged]