3 ms·
What quantization level is that? Because official endpoints are slow.
by ComputerGuru 2mo ago
What quantization level is that? Because official endpoints are slow.
- bel8 2mo agoFrom opencode go $10/mo plan I get between 60 t/s and 100 token/s even with large contexts of 150k+ tokens. I wouldn't call 80 t/s slow.
- ponyous 2mo agoYou are right, relatively to other llm providers this is not slow. But if you think what is possible when you have 1000t/s a sec you might find it slow.
- hatefulmoron 2mo agoThat's across 64 concurrent streams; you could make more concurrent requests to DeepSeek API no?
- ak_t 2mo agoIt doesn't need extra quantization. The official weights are natively mixed precision FP4/FP8, so it fits in ~160GB. The API slowness is probably from being batched with other concurrent user requests. The provider's aggregate throughput gets higher but per-stream speed slows down.
- zargon 2mo agoV4 Flash fits entirely in two RTX Pro 6000s without any quantization at all.
- olejorgenb 2mo agoI get around 60-110 TPS on tensorx.ai hosted models.