7 ms·
Expensively Quadratic: The LLM Agent Cost Curve
- stuxf 8mo ago> Some coding agents (Shelley included!) refuse to return a large tool output back to the agent after some threshold. This is a mistake: it's going to read the whole file, and it may as well do it in one call rather than five. disagree with this: IMO the primary reason that these still need to exist is for when the agent messes up (e.g reads a file that is too large like a bundle file), or when you run a grep command in a large codebase and end up hitting way too many files, overloading context. Otherwise lots of interesting stuff in this article! Having a precise calculator was very useful for the idea of how many things we should be putting into an agent loop to get a cost optimum (and not just a performance optimum) for our tasks, which is something that's been pretty underserved.
- tekacs 8mo agoI think that's reasonable, but then they should have the ability for the agent to, on the next call, override it. Even if it requires the agent to have read the file once or something. In the absence of that you end up with what several of the harnesses ended up doing, where an agent will use a million tool calls to very slowly read a file in like 200 line chunks. I think they _might_ have fixed it now (or agent-fixes, my agent harness might be fixing it), but Codex used to do this and it made it unbelievably slow.
- reactordev 8mo agoYou’re describing peek. An agent needs to be able to peek before determining “Can I one shot this or does it need paging?”
- tekacs 8mo agoYep, I previously implemented it under that name in my own harness. That being said, there is value in actually performing a normal read, because you do often complete it on that first glance.
- reactordev 8mo agoConfession, I too implemented a “smart” read. A read unless it’s over a size, then it’s paged, or if it’s a specific format, a summary. However, I also supply `cat`
- inetknght 8mo ago> when you run a grep command in a large codebase and end up hitting way too many files, overloading context. On the other hand, I despise that it automatically pipes things through output-limiting things like `grep` with a filter, `head`, `tail`, etc. I would much rather it try to read a full grep and then decide to filter-down from there if the output is too large -- that's exactly what I do when I do the same workflow I told it to do. Why? Because piping through output liming things can hide the scope of the "problem" I'm looking at. I'd rather see the scope of that first so I can decide if I need to change from a tactical view/approach to a strategic view/approach. It would be handy if the agents could do the same thing -- and I suppose they could if I'm a little more explicit about it in my tool/prompt.
- kaibee 8mo agoIn my experience this is what Claude 4.5 (and 4.6) basically does, depending on why its grepping it in the first place. It'll sample the header, do a line count, etc. This is because the agent can't backtrack mid-'try to read full file'. If you put the 50,000 lines into the context, they are now in the context.
- jtbayly 8mo agoWhy can't the LLM/agent edit the context and dump that file if it decides it was dumb to have the whole thing in the context?
- cyanydeez 8mo agoBase model is content. If it reads to much it becomes the content. What you want is a harness that continually inserts file portions until a sufficiently bright light bulb goes off. When they say agentic AI, ITS BASICALLY: <command><content-chunk-1/></command> its the ugliest string mashing indeterministic garbage the bearded masters would face palm.
- inetknght 8mo ago> If you put the 50,000 lines into the context, they are now in the context. And you can't revert back to a previous context, and then add in new context summarizing to something like "the file is too large" with how to filter "there are too many unrelated lines matching '...', so use grep"? Using output-limiting stuff first won't tell you if you've limited too much. You should search again after changing something; and if you do search again then you need to remember which page you're on and how many there are. That's a bit more complex in my opinion, and agents don't handle that kind of complexity very well afaik.
- Areena_28 8mo ago[flagged]
- seanhunter 8mo agoTFA is talking about being quadratic in dollar cost as the conversation goes on, not quadratic in time complexity as n gets larger. Edit to add: I see you are a new account and all your comments thus far are of a similar format, which seems highly suspicious. In the unlikely event you are a human, please read the hacker news guidelines https://news.ycombinator.com/newsguidelines.html https://news.ycombinator.com/newsguidelines.html
- jauntywundrkind 8mo agoVery awesome to see these numbers, to see this explored so. Nice job exe.dev.
- TZubiri 8mo agoI'm not sure, but I think that cached read costs are not the most accurately priced, if you consider your costs to be costs when consuming an API endpoint, then the answer will be 50k tokens, sure. But if you consider how much it costs the provider, cached tokens probably have a way higher margin than (the probably negative margin of ) input and output inference tokens. Most caching is done without hints from the application at this point, but I think some APIs are starting to take hints or explicit controls for keeping state associated with specific input tokens in memory, so these costs will go down, in essence you really don't reprocess the input token at inference, if you own the hardware it's quite trivial to infer one output token at a time, there's no additional cost, if you have 50k input tokens, and you generate 1 output token, it's not like you have to "reinfer" the 50k input tokens before you output the second token. To put it in simple terms, the time it takes to generate the Millionth output token is the same as the first output token. This is relevant in an application I'm working on where I check the logprobs and not always choose the most likely token(for example by implementing a custom logit_bias mechanism client-side), so you can infer 1 output token at a time. This is not quite possible with most APIs, but if you control the hardware and use (virtually) 0 cost cached tokens, you can do it. So bottomline, cached input tokens are almost virtually free naturally (unless you hold them for a loong period of time), the price of cached input APIs is probably due to the lack of API negotiation as to what inputs you want to cache. As APIs and self-hosted solutions evolve, we will likely see the cost of cached inputs masssively drop down to almost 0. With efficient application programming the only accounting should be for output tokens and system prompts. Your output tokens shouldn't be charged again as inputs, at least not more than once.
- eshaham78 8mo ago[flagged]
- 2001zhaozhao 8mo agoAre you hosting your own infrastructure for coding agents? At least from first glance, sharing actual codebase context across compacts / multiple tasks seems pretty hard to pull off with good cost-benefit unless you have vertical integration from the inference all the way to the coding agent harness. I'm saying this because the current external LLM providers like OpenAI tend to charge quite a bit for longer-term caching, plus the 0.1x cache read cost multiplied by # LLM calls, so I doubt context sharing would actually be that beneficial considering you won't need all the repeated context every time, so caching context results in longer context for each agentic task which might increase API costs by more overall than you save by caching.
- devcraft_ai 8mo ago[flagged]
- Agent_Builder 8mo ago[dead]
- intellirim 8mo ago[dead]
- nivcmo 8mo ago[dead]
- alexhans 8mo agoNice article. I think a key part of the conversation is getting people to start thinking in terms of evals [1] and observability but it's been quite tough to combat the hype of "but X magic product just solves what you mentioned as a concern for you". You'd think cost is an easy talking point to help people care but the starting points for people are so heterogeneous that it's tough to show them they can take control of this measurement themselves. I say the latter because the article is a point in time and if they didn't have a recurrent observation around this, some aspects may radically change depending on the black box implementations of the integrations they depend on (or even the pricing strategies). [1] https://ai-evals.io/ https://ai-evals.io/
- intellirim 8mo agoIn my experience building agent pipelines, the real cost explosion happens at tool call chains - each hop multiplies tokens in ways that are hard to anticipate . Adding structured logging per step helped me identify and cut the worst offenders.
- the_harpia_io 8mo ago[flagged]
- aurareturn 8mo agoI disagree. I used to spend most of my time writing code, fixing syntax, thinking through how to structure the code, looking up documentation on how to use a library. Now I first discuss with an AI Agent or ChatGPT to write a thorough spec before handing it off to an agent to code it. I don’t read every line. Instead, I thoroughly test the outcome. Bugs that the AI agent would write, I would have also wrote. Example is unexpected data that doesn’t match expectations. Can’t fault the AI for those bugs. I also find that the AI writes more bug free code than I did. It handles cases that I wouldn’t have thought of. It used best practices more often than I did. Maybe I was a bad dev before LLMs but I find myself producing better quality applications much quicker.
- adrianN 8mo agoYou have way more trust in test suites than I do. How complex is the code you’re working with? In my line of work most serious bugs surface in complex interactions between different subsystems that are really hard to catch in a test suite. Additionally in my experience the bugs AI produces are completely alien. You can have perfect code for large functions and then somewhere in the middle absolutely nonsensical mistakes. Reviewing AI code is really hard because you can’t use your normal intuitions and really have to check everything meticulously.
- aurareturn 8mo agoIf it’s hard to catch with a comprehensive suit of test, what makes you think you can catch them by hand coding?
- skydhash 8mo agoA great lot of thinking about the code, which you can only do if you’re very familiar with it. Writing the code is trivial. I spend nearly all my work hours thinking about edge cases.
- anvevoice 8mo ago[flagged]
- formerly_proven 8mo ago> Too little and the agent loses coherence. Obviously you don't have to throw the data away, if the initial summary was missing some important detail, the agent can ask for additional information from a subthread/task/tool call.
- embedding-shape 8mo ago> Instead of feeding 500 lines of tool output back into the next prompt Applies for everything with LLMs. Somewhere along the idea, it seems like most people got the idea that "More text == better understanding" whereas reality seems to be the opposite, the less tokens you can give the LLM with only the absolute essentials, the better. The trick is to find the balance, but "more == better" which many users seem to operate under seems to be making things worse, not better.
- Tiberium 8mo agoAnother new LLM slop account on HN..
- seyz 8mo ago128k tokens sounds great until you see the bill
- 0-_-0 8mo agoThe cache gets read at every token generated, not at every turn on the conversation.
- mzl 8mo agoDepends on which cache you mean. The KV Cache gets read on every token generated, but the prompt cache (which is what incurs the cache read cost) is read on conversation starts.
- 0-_-0 8mo agoWhat's in the prompt cache?
- bsenftner 8mo agoWay too much. This has got to be the most expensive and most lacking in common sense way to make software ever devised.
- mzl 8mo agoThe prompt cache caches KV Cache states based on prefixes of previous prompts and conversations. Now, for a particular coding agent conversation, it might be more involved in how caching works (with cache handles and so on), I'm talking about the general case here. This is a way to avoid repeating the same quadratic cost computing over the prompt. Typically, LLM providers have much lower pricing for reading from this cache than computing again. Since the prompt cache is (by necessity, this is how LLMs work) prefix of a prompt, if you have repeated API calls in some service, there is a lot of savings possible by organizing queries to have less commonly varying things first, and more varying things later. For example, if you included the current date and time as the first data point in your call, then that would force a recomputation every time.
- lostmsu 8mo ago> The prompt cache caches KV Cache states Yes. The cache that caches KV cache states is called the KV cache. "Prompt cache" is just index from string prefixes into KV cache. It's tiny and has no computational impact. The parent was correct to question you. The cost of using it comes from the blend of the fact that you need more compute to calculate later tokens and the fact that you have to keep KV cache entries between requests of the same user somewhere while the system processes requests of other users.
- cs702 8mo ago> By 50,000 tokens, your conversation’s costs are probably being dominated by cache reads. Yeah, it's a well-known problem. Every AI company is working on ways to deal with it, one way or another, with clever data center design, and/or clever hardware and software engineering, and/or with clever algorithmic improvements, and/or with clever "agentic recursive LLM" workflows. Anything that actually works is treated like a priceless trade secret. Nothing that can put competitors at a disadvantage will get published any time soon. There are academics who have been working on it too, most notably Tri Dao and Albert Gu, the key people behind FlashAttention and SSMs like Mamba. There are also lots of ideas out there for compressing the KV cache. No idea if any of them work. I also saw this recently on HN: https://news.ycombinator.com/item?id=46886265 https://news.ycombinator.com/item?id=46886265 . No idea if it works but the authors are credible. Agentic recursive LLMs look most promising to me right now. See https://arxiv.org/abs/2512.24601 https://arxiv.org/abs/2512.24601 for an intro to them.
- yowlingcat 8mo agoWhat do you think about RLMs? At first blush it looks like sub agents with some sprinkles on top, but people who have become more adept with it seem to show its ability to handle sublinear context scaling behavior very effectively.
- cs702 8mo agoBy "agentic recursive LLMs," I mean all the approaches that involve agents recursively calling LLMs, including RLMs. My post in fact links to an RLM paper.
- vatsachak 8mo agoThe brain trims it's context through forgetting details that do not matter LLMs will have to eventually cross this hurdle before they become our replacements
- collinwilkins 8mo agowhat i've learned running multi-agent workflows... >use the expensive models for planning/design and the cheaper models for implementation >stick with small/tightly scoped requests >clear the context window often and let the AGENTS.md files control the basics
- rubicon33 8mo agothere’s something of a paradox there. Reduce the context window and work on smaller/tightly scoped requests? Isn’t the whole value proposition that I can work much faster? To do that, I naturally try to describe what I want at a higher, vaguer level.
- readyforbrunch 8mo agoThat's where something like openspec and beads come in. You work high level, create a spec and break it down into beads (small tasks). Your main agent then spawns workers that perform a task with limited scope.
- deleted 8mo ago[deleted]
- mergisi 8mo ago[dead]