2 ms·
> if you run one model, run glm-5.3 That is a horrible take-away from this, with only 28 tasks and a high pass rate for most models, it says almost nothing. T
by CMay 1mo ago
> if you run one model, run glm-5.3
That is a horrible take-away from this, with only 28 tasks and a high pass rate for most models, it says almost nothing.
Test a model for your use case and use the fastest, smallest, cheapest model that 100% satisfies your use case.
Or, if you truly do need a model with strong generalized performance, definitely do not take a benchmark like this serious with such a limited task set.
- ed-is-ai 1mo agoThe point I'm making is that most models are good enough for most tasks, so choose on speed/cost. Definitely benchmark on your own tasks. GLM-5.3 is the winner on mine, on yours maybe not. I am not trying to be a universal benchmark, as these serve no one but the person doing the benchmark. My previous post sank like a stone, but all the evidence + code to run this + what you have todo to adapt it for your own use cases is all here https://github.com/ed-is-ai/featherbench https://github.com/ed-is-ai/featherbench Encourage everyone to eval like the devil