3 ms·
I stopped trusting benchmarks ever since LLMs started speaking alien like English, if I can't understand what they are saying how can I trust them.
by ouz-a 2mo ago
I stopped trusting benchmarks ever since LLMs started speaking alien like English, if I can't understand what they are saying how can I trust them.
- thomasnowhere 2mo agosame here, it reads exactly the same whether the number is real or completely made up, so the confidence stops meaning anything.
- puszczyk 2mo agoThis article is about using LLMs to overfit for a specific benchmark (or make a custom software for niche use cases) though. Not about LLMs benchmaxxxing
- dgellow 2mo agoIsn’t that the same? It’s a sort of recursive version of overfitting specific benchmarks
- r_lee 2mo agoThat's the load-bearing smoking gun—should I write a better benchmarks to catch the seams?
- lostmsu 2mo agoI wouldn't generalize this to all LLMs. So far I only saw Anthropic ones affected.