4 ms·
> The same model will benchmark very differently This is surprising to me. Does anyone have a definitive answer that accounts for this difference between provi
by plandis 21d ago
> The same model will benchmark very differently
This is surprising to me. Does anyone have a definitive answer that accounts for this difference between providers?
Do the benchmarks that Open Router runs not account for the stochastic nature of the models? Do the providers lie about quantization or context window sizing? Does Open Router not take into account variations for a given model?
I could understand latency/cost benchmarks varying but not the actual generated token responses of the models given the exact same model parameters.
- sheo 21d agoAs said in the article, all of the providers run different runtimes with proprietary, sometimes untested and weird optimisations There are multiple types of quantisation: weights and KV Cache. Quantizing kv cache can drastically hurt performance