3 ms·
I just want an API that takes these crazy small / cost effective open weight models and charges peanuts for access. Think, DeepSeek Flash (before the price hik
by apatheticonion 22d ago
I just want an API that takes these crazy small / cost effective open weight models and charges peanuts for access.
Think, DeepSeek Flash (before the price hikes) prices.
If I can run this on a 32gb card while they have a datacenter with wholesale electricity prices, why are we not seeing "cents per billion tokens" pricing?
- mrngld 22d agoBecause your 32gb card isn't running this at 1500 token/s. Serving these things at scale with the enormous context windows real use demands and doing some with usable performance takes a lot of expensive hardware. Yes, their margin on straight inference is allegedly really high, but that's severely offset by high capital costs. If you want to spend a new car worth of money and still not serve as fast as Cerebras because you can't simply buy their mammoth custom chips, then yes you too can self host a huge Deepseek or GLM model.
- apatheticonion 20d agoYes, that's what I'm saying. 200t/s is more than enough to 5 - 10x my productivity and the intelligence of current open weight models covers 90% of both my guided and software factory workflows