3 ms·
The most interesting thing about it is that it’s the type of task where you'd expect LLMs to do well, yet the best models only score around 30%, while top human
by zone411 2y ago
The most interesting thing about it is that it’s the type of task where you'd expect LLMs to do well, yet the best models only score around 30%, while top humans get 100%. Many other benchmarks are also getting close to saturation.