5 ms·
You made me realize something. I routinely spend upwards of 500$ per month on LLMs for coding (expensed towards clients). However I live in a place where 500$ i
by fbrncci 4mo ago
You made me realize something. I routinely spend upwards of 500$ per month on LLMs for coding (expensed towards clients). However I live in a place where 500$ is around the avg. salary. I’m lucky that I know my way around western clients. Clients who pay these expenses and are happy to work with me because I am still about 50% cheaper than local talent in EU/US, while my salary at home converts to an upper class income at the highest tax bracket.
Which of course causes some unfairness on both ends. Nobody here can compete with me. I often use left over tokens on local client projects; which despite lower pay, still pays off because they now take hours not days or weeks to complete. And nobody in the local clients talent pool can compete with me; unless they charge about half the market rate.
Take away my 500$ monthly grant; and I’d be more or less screwed. Better open models will more or less start to reduce this advantage. It’s not like I positioned myself here on purpose. But it’s definitely a „right place, right time“ situation.
- swader999 4mo agoIf you are running multiple agents your cost to them should be multiples less what their roi is.
- fbrncci 4mo agoMy costs are 0$ as any token or subscription spend on agents is invoiced as an expense to my clients.
- kreelman 4mo agoThanks so much for being bold enough to be fairly open about the costs, how you arrange billing and the advantages that's given you. I've been fooling around with DeepSeek 4 agentically. It's probably not as good as Anthropic offerings, but even those seem to be roiled in politics and strife and DeepSeek 4 is very good IMHO. I'll later try out GLM. I'm in Australia. The government has set up a "return and earn" scheme to keep aluminium cans, plastic bottles and paper drink cartons out of the waste stream. A laudable project. The money you make from return drink containers is pretty low, $AU 0.1 per container. I've participated to get the rubbish out of natural water streams and to make a nano amount of money on the side. When I looked at the costs of an app I was getting DeepSeek to help me with, I realised that the several hours I'd spent learning and building had cost something like 8 recycled containers. In my head after doing some DeepSeek stuff, I calculate a "cans per app" metric for myself for fun. I may even setup a simple graph to view my costs that way. I kind of hope the Anthropics of the world get enough price competition from sources like DeepSeek and GLM to drop their prices significantly. Time will tell. I'm using the Chinese DeepSeek provider, so everything done there could potentially be taken and used by the CCP... But this is hobbyist learning. There is probably a market for Deepseek/GLM served from non CCP available servers. I might even look into how hard that would be to setup here. I also hope that inference focused hardware will come to the fore, reducing energy use and cost. Realistically this will take time though, on the order of years. Here in Oz, we have community batteries that community members can charge and later draw from. Their electricity prices are competitive. I wonder if someone could setup something like a community battery to run data centres... That way reasonable environmental consideration could be given to inference power generation... This might not work in a market like the US or Europe, but small market size might be an advantage... Who knows.
- esperent 4mo ago> I'm using the Chinese DeepSeek provider, so everything done there could potentially be taken and used by the CCP As opposed to Anthropic or OpenAI where everything done could potentially be taken and used by the US government. Also, replace "could potentially" with "will definitely" in both cases, there's no conspiracy here. We're stuck between two bad positions, so just use the one that's best for you, and wait for a better solution to arrive.
- usef- 4mo agoIt's very easy to use other providers. See https://openrouter.ai/ https://openrouter.ai/ which also lets you filter by where the provider is hosted and their data retention policy. Jeremy Howard was recommending fireworks.ai as a host of you want to go direct. Or there's Cloudflare. For subscription alternatives people here on HN seem to mention Open Code Go a lot too https://opencode.ai/go https://opencode.ai/go
- SyneRyder 4mo ago> There is probably a market for Deepseek/GLM served from non CCP available servers. I might even look into how hard that would be to setup here. Please do. There is definitely a market for Deepseek / GLM hosted from non-China servers, there's over 20 providers for GLM 5.2 on OpenRouter alone... and they're all either Singapore (home of Z.AI / GLM), China, or US. There is nothing yet listed on OpenRouter from Europe (Inceptron still only has GLM 5.1). And of course, there is absolutely nothing hosted in Australia. We're in a particularly dire situation in Australia. We're about to be cut off from Claude Fable and premium American models. The European Mistral models are garbage, at least in comparison to US models. Our only hope is going to be Chinese models (GLM 5.2 is good), and we're not even hosting them in Australia. By the way, if you haven't tried an Anthropic model, it's worth spending at least $20 one month to give Opus 4.8 a try. I only got one night of access to Fable before I was cut off, but one single evening of Fable provided plans that I've been working through for about a week afterwards with Opus 4.8... and that was only Fable, not even Mythos. That's the kind of intelligence lead Australia is about to be cut off from. (And kudos on the Containers For Change, that's something I do as well - mostly as an exercise incentive to walk to the local recycling machine, because the money certainly doesn't compensate for the time spent on the recycling.)
- listic 4mo agoThanks for sharing your insight. Mind if I ask you for a few vibe coding tips? I failed to solve you gh puzzle in the profile though.
- lanthissa 4mo agoAI is the first technology that doesn't incentivize offshoring, and incentivizes co-location of talent. A NYC dev and a dev in india have the same ai costs, based the ratio tokens/salary it becomes less of comparative disadvantage to be in NYC. Now combine that with the fact that AI makes the act of generating code less a % time of the job, and the ability to get/refine requirements more of the job and you have a decent shift.
- Sammi 4mo agoErrr you just responded to someone that is offshore and is using AI to be much cheaper than local talent.
- fbrncci 3mo agoThe tokens/salary ratio is not relevant at all. Because while 200-500$ is a lot of money, it’s still a fraction of the salary you’d pay any dev in the world. It just comes out as a tooling expense. It also matters how those devs use the tools; you can’t assume everyone gets the same out of it. So that amount can last a day or it can last a month. I would say a dev in a developing nation would be more budget aware than someone being used to everything being priced in NYC rates. For example I build other AI products and I have been hyper aware of the token spend of our users. I was going crazy seeing that some users were having 5$ conversations. So that was optimized and I found ways to use sub agents to get it down to 1-2$. Just for management asking me why I was worrying to begin with? The users using these are consultants being paid 120$ per hour. They have a daily 10-20$ token expense, no problem. “But amazing job on the cost reduction.”.. well 5$ for me is what I spend on food daily. While the consultant is slamming: “yes” 10 times in a chat , for whatever reason for the same cost. Would the NYC dev care as much natively? No. You can still hire three devs in India for the price of a dev in NYC. Now you give them AI and you might only need 1-2. That makes offshoring even more appealing, not less. And the dev in India now having tooling to out compete local talent. Well that’s my reality (I am not in India though).
- whazor 4mo agoThe problem is that the differences between flagship and local models are compounding heavily. An 4% different could be massive when you keep iterating on the same code base.
- swiftcoder 4mo ago> The problem is that the differences between flagship and local models are compounding heavily This depends a lot on how you work, and how much of the architectural thinking you do yourself. People seem to lose sight of the fact that a flash model today is as powerful as a frontier model from a year ago. If you were happy with GPT 4.x, you should be ecstatic that equivalent power is now basically free...
- wolttam 4mo agoI am one of those ecstatic folk :)
- jmalicki 3mo agoI find that with a lot of the cheaper models, I end up spending a lot more time correcting the easy stuff. If I am 100% spot on on the architectural stuff, I have anecdata that some of the frontier models might actually be cheaper than "cheaper" alternatives once you look at what it takes to get to good output, since they require less correction. But that is on pure token costs. When you value the human overseer's time, there is just no competition. A model that is 10x more expensive that requires 10% less oversight is just a plain win.
- swiftcoder 3mo ago> A model that is 10x more expensive that requires 10% less oversight is just a plain win I think you are drastically underestimating the cost delta here. We're talking models we can run pretty much continuously for $10/month in tokens.
- jmalicki 3mo ago