3 ms·
What worries me most are questions like this: these autonomous AI employees or automations - it doesn't change the essence - consume a lot of LLM tokens. Please
by ianberdin 18d ago
What worries me most are questions like this: these autonomous AI employees or automations - it doesn't change the essence - consume a lot of LLM tokens. Please tell me, how do you pay for this? Do you use, for example, the Anthropic or OpenAI API directly, or do you connect, for example, a Codex subscription?
- lnenad 18d agoIf you use a bajillion tokens your economical approach is to self host.
- noir_lord 18d agoThere is a point where those two lines do cross but with them effectively subsidising token cost by burning debt, it's further away than it will be at some point. That said I use local only models purely because I don't want to use remote models, never having to think about token costs is worth it and no one is training anything on my data either.
- mchusma 18d agoOur control panel is on a server but individual agents actually are run on anyone’s machine. This allows us to use the native harnesses including subscriptions. Yes it uses a lot more tokens, but i have have Claude $200, OpenAI $200, and SuperGrok Heavy $300 (or whatever it is called) that includes Cursor Ultra. My machine runs most of them, but some other team members have agents running on their machines using 1-2 $200/mo subs. I would say this setup probably costs us about $1,000/month total. (I’m excluding traditional engineering use of LLMs from this number. This is the cost of all the “AI employees”. One reason we did it this way was to use subs.
- grvdrm 18d agoDo I understand correctly that by using your machine and Claude Code or something like it you are avoiding API pricing?
- bleonard 17d agoI've been spending all my time recently thinking about LLM costs. From that, I am currently of the mind that the most interesting question is if they can get it to work in the first place. After that, like with "performance" in the past, I believe there are ways to optimize things. We see the code harnesses doing this out of necessity, for example. Some of the tools we use: https://www.induction.ai/docs/context-management https://www.induction.ai/docs/context-management