2 ms·
>I think it's a really clever idea that you could measure an LLM's prediction abilities and language understanding by some sort of held-out compression metric
by krackers 2mo ago
>I think it's a really clever idea that you could measure an LLM's prediction abilities and language understanding by some sort of held-out compression metric
This is the premise of https://huggingface.co/spaces/Jellyfish042/UncheatableEval https://huggingface.co/spaces/Jellyfish042/UncheatableEval
- adamgordonbell 2mo agovery cool
- andai 2mo agoNice. This ranking basically matches other benchmarks, from what I can tell. Which implies this would probably also hold for the larger models, which are sadly not included in the leaderboard.
- krackers 2mo agoI think it's limited to using base models (non post-trained), because the post-training would skew the logit distribution. There are ways to "coax" post-trained models back into behaving "like" a base model, I wonder if the benchmark could be unofficially updated with those somehow.