3 ms·
The 22 tokens per second claim for Mixtral conveniently fails to mention what type of quantization is going on with that benchmark.
by root_axis 3y ago
The 22 tokens per second claim for Mixtral conveniently fails to mention what type of quantization is going on with that benchmark.
- mmoskal 3y agoThey say 200GB/s of mem bandwidth; Mixtral uses 13B parameters for inference; they claim 22t/s, so 22 * 13B parameters per second, so (200 * 8) bits / (22 * 13) around 5.6 bits / parameter max. With overheads, it's probably 4 bit quant. edit: formatting