3 ms·
In research we are still seeing massive jumps. Subjects that LLMs were completely useless for half a year ago are now definitely in scope. And there are benchm
by Certhas 2mo ago
In research we are still seeing massive jumps. Subjects that LLMs were completely useless for half a year ago are now definitely in scope.
And there are benchmarks that cleanly separate the SOTA models:
https://epoch.ai/MirrorCode https://epoch.ai/MirrorCode
Saturation of benchmarks is a property of benchmarks just as much as of the models.