3 ms·
This is a 50B model. (Mixtral 8x7b)
by snowfield 3y ago
This is a 50B model. (Mixtral 8x7b)
- SeanAnderson 3y agoOh, sorry, I assumed the 8 was for quantization. 8x7b is a new syntax for me. Still, the NVIDIA chart shows Llama v2 70B at 750 tok/s, no?
- tome 3y agoI guess that's total throughput, rather than per user? You can increase total throughput by scaling horizontally. You can't increase throughput per user that way.