3 ms·
Looks very thorough, definitely a step up from that paper that got a lot of publicity, and was based on a single convoluted task. I think an approach like this
by epups 3y ago
Looks very thorough, definitely a step up from that paper that got a lot of publicity, and was based on a single convoluted task. I think an approach like this is the future for benchmarking LLM's, in general.
A minor nitpick here is that there is no human baseline to compare. How would an average human perform on these tests?