3 ms·
Because your 32gb card isn't running this at 1500 token/s. Serving these things at scale with the enormous context windows real use demands and doing some with
by mrngld 1mo ago
Because your 32gb card isn't running this at 1500 token/s. Serving these things at scale with the enormous context windows real use demands and doing some with usable performance takes a lot of expensive hardware. Yes, their margin on straight inference is allegedly really high, but that's severely offset by high capital costs.
If you want to spend a new car worth of money and still not serve as fast as Cerebras because you can't simply buy their mammoth custom chips, then yes you too can self host a huge Deepseek or GLM model.
- apatheticonion 29d agoYes, that's what I'm saying. 200t/s is more than enough to 5 - 10x my productivity and the intelligence of current open weight models covers 90% of both my guided and software factory workflows