3 ms·
Smaller models are also generally faster, so thinking "more" may not matter and may even come out ahead.
by chmod775 25d ago
Smaller models are also generally faster, so thinking "more" may not matter and may even come out ahead.
- celrod 25d agoIf Q4 takes less than 1.3x as many tokens as bf16 or q8, it could still end up being faster, given how decode tends to be bandwidth bound. The kv cache was still bf16, so a few ops are the same between quants.