4 ms·
The thing they don't show is the one we really need, especially because model providers can skimp on quality (run lower quantization, lower kv cache precision,
by Gracana 23d ago
The thing they don't show is the one we really need, especially because model providers can skimp on quality (run lower quantization, lower kv cache precision, etc) to improve their pricing and performance. I agree that it's probably too expensive to keep running the benchmark, but we need some way to hold the providers to a certain standard, otherwise every user has to discover the problems on their own.