3 ms·
The author only compared output token costs -- but for typical agentic workloads, input tokens dominate the costs by a large margin. Running inference locally,
by maho 5mo ago
The author only compared output token costs -- but for typical agentic workloads, input tokens dominate the costs by a large margin. Running inference locally, input tokens are, to first order, free. (They only generate implicit costs through higher time-to-first-token, higher power use, and lower token output speed).
- Wilya 5mo agoYeah, that completely invalidates his point. I looked at a couple random agentic sessions in my openrouter activity, and the input cost is 10x the output cost. Prompt caching on openrouter is complicated and unreliable. On local hardware with llama-cpp, it's mostly free.
- amluto 5mo agoEven ignoring superior caching on a local setup, Mac hardware can often process input token around 10x as quickly as they produce output tokens. Openrouter seems to have only a 2x difference on the same models.
- bigyabai 5mo agoFor larger contexts (eg. 20,000+ token agent workflows), being 10x faster still isn't enough. You have to be close to ~100x faster at crunching contexts for it to feel like realtime.