3 ms·
Ask HN: What Makes Tokens Expensive?
Are there really such architectural differences between for example sol and astra that makes astra twice more in token price? Or maybe token cost covers training expenses? For me it's hard to believe that astra need twice the computing power that sol needs...
- verdverm 17d agocosts are in a large part tied to model parameter count, eg. a small model fits on one GPU, the biggest models require an entire rack the primary difference you see through those prices is model size that you don't seem much capability difference is why Big Ai is pushing for regulatory capture, because small models are nearly as good for many tasks at a fraction of the cost, undermining the business model they have used to justify the most expensive infra build out in human history
- PaulHoule 17d agoAlso I think there are diminishing returns to larger models and the harness can make up for some weaknesses. Like maybe the huge model can hypothetically answer a question off the cuff but the small model can get some search results and think about those and give you an answer.
- verdverm 17d ago100% I am big believer (based on personal experience) that there is a ton of ROI in harness engineering, to go along with the context engineering, invariably intertwined
- kbrannigan 17d agoIt's not always larger models. Often time it's a smaller model fit for a specific use case. In the future we will see 16B models that are for : Legal, Healthcare, Science, Software Engineering, Education, Forestry, Music production
- song_synth 16d agoI created a model for music production that doesn't require AI at all. It is algebraic by nature. This means determinism & creativity are first class citizens. https://monictheory.com/developer-api https://monictheory.com/developer-api
- yorwba 17d agoWillingness to pay. If people are willing to use Astra instead of Sol even if it costs twice as much, OpenAI has no need to make it cheaper. Prices only reflect costs in a competitive market where customers actively seek out cheaper alternatives.
- cyanydeez 15d agoThere still intrinsic loss to the model.
- harsh_yadav21 15d agoWillingness to pay is definitely the primary driver, but you aren't just paying for raw compute, you are paying for queue priority and uptime reliability under heavy concurrent load.
- dougame 14d ago[flagged]