2 ms·
Don't use any public benchmarks, every single one is worthless for your own use cases essentially. Spend a day or two going through your existing chat sessions
by embedding-shape 1mo ago
Don't use any public benchmarks, every single one is worthless for your own use cases essentially.
Spend a day or two going through your existing chat sessions, and create your own private benchmark with test cases based on real tasks, that you don't share with anyone nor publicly. Make it easy to add/remove new harnesses and model combinations, make it give you a final score, ideally avoid using other LLMs for scoring, then use this to figure out if the new model/harness actually improves things for you.
I've been doing this for some time, and while most new releases show big increases in the benchmarks/evaluations, my own benchmark usually barely moves.
- krashidov 1mo ago> make it give you a final score what does this mean exactly? A scored based on what?
- embedding-shape 1mo agoFor translations, the score is basically 1 or 0. For some tasks, the least amount of LOC gives the highest score, and so on. Basically, you need to figure out how to score it, so you can compare scores across agents/models.