2 ms·
i can provide some sources. 5.6 sol ultrafast on cerebras is 750tps, open models readily exceed this. just by using a smaller model cerebras serves qwen 3.8 27
by lukewarm707 20d ago
i can provide some sources.
5.6 sol ultrafast on cerebras is 750tps, open models readily exceed this. just by using a smaller model cerebras serves qwen 3.8 27b at 1850tps. or even larger models, mimo 2.5 pro was served for a while at 1000tps. and so on. [https://inference-docs.cerebras.ai/models/choose-a-model https://inference-docs.cerebras.ai/models/choose-a-model]
the chinese ai companies have 10% of the total compute resources of the US ones. since the USA tries to stop them from buying nvidia gpus. they maxed out the efficiency.
deepseek v4.1 has engram architecture. it has 550b params instead of 5T+ for astra/fable. it has 8b active instead of potentially hundreds active for astra/fable.
compare input/output/cache: $0.15/$0.60/$0.003 for v4.1 to $10.00/$50.00/$1.00 for astra and $10.00/$50.00/$0.25 for fable.
astra cache reads are over 330 times more expensive.
at the artificial analysis 7:2:1 ratio, deepseek is $0.18/m, fable is $7.18/m, astra is $7.7/m.
but what about intelligence? AA would rate deepseek v4.1 at AA 40, astra is AA 53.
so it cost 4,180% more for 32% more intelligence.
they are serving that at over 250tps at baseten. to get close to that on astra API you are paying double the cost for fast mode.
so it is now 8456% more expensive for a similar speed and 32% more intelligence. 84 times more expensive.
- vardalab 20d agoI don’t even think that regular Astra medium is anywhere close to 100 for tg
- user43928 20d agoThe proper comparison would be GPT 5.6 Luna at 38 on the intelligence score and $0.18 per task vs $0.27 for DeepSeek Flash 4.1.
- lukewarm707 20d agoi was trying to make a point about efficiency of serving the model. the cost per task itself would not be enough to show that. you could compare gpt 5.6 luna. if you did that the same way as before you would get a blended price of $0.17 for luna at AA 38. for baseten it would be $0.20 for v4.1 at AA 40. assume roughly the same intelligence. on AA openai gets 117tps. baseten gets 284tps. so 18% more expensive but 142% more tps. the fast mode is again double the cost, roughly same intelligence. so luna in that case would be 70% expensive. take the per task cost and it would still 9% more expensive. so i think there is something to be said about the efficiency of the model.