4 ms·
I think this is a really interesting paper from Cohere, it really feels that at this point in time you can't trust any public benchmark, and you really need you
by pongogogo 1y ago
I think this is a really interesting paper from Cohere, it really feels that at this point in time you can't trust any public benchmark, and you really need your own private evals.
- ilrwbwrkhv 1y agoYup in my private evals I have repeatedly found that DeepSeek has the best models for everything and yet in a lot of these public ones it always seems like someone else is on the top. I don't know why.
- __alexs 1y agoPublishing them might help you find out.
- refulgentis 1y ago^ This. If I had to hazard a guess, as a poor soul doomed to maintain several closed and open source models acting agentically, I think you are hyper focused on chat trivia use cases (DeepSeek has a very, very, hard time tool calling and they say as much themselves in their API docs)
- deleted 1y ago[deleted]
- AstroBen 1y agoAny tips on coming up with good private evals?
- pongogogo 1y agoYes, I wrote something up here on how Andrei Kaparthy evaluated grok 3 -> https://tomhipwell.co/blog/karpathy_s_vibes_check/ https://tomhipwell.co/blog/karpathy_s_vibes_check/ I would pick one of two parts of that analysis that are most relevant to you and zoom in. I'd choose something difficult that the model fails at, then look carefully at how the model failures change as you test different model generations.