3 ms·
We're gonna need some new benchmarks... ARC-AGI-3 might be the only remaining benchmark below 50%
by pants2 6mo ago
We're gonna need some new benchmarks...
ARC-AGI-3 might be the only remaining benchmark below 50%
- randomtoast 6mo agoHumanity's Last Exam (HLE) is already insanely difficult. It introduces 2,500 questions spanning mathematics, humanities, natural sciences, ancient languages, ... Here is an example question: https://i.redd.it/5jl000p9csee1.jpeg https://i.redd.it/5jl000p9csee1.jpeg No human could even score 5% on HLE.
- saberience 6mo agoI've never understood the point of things like HLE, it doesn't really prove or show anything since 99.99% of humans can't do a single question on this exam. That is, it's easy to make benchmarks which humans are bad at, humans are really bad at many things. Divide 123094382345234523452345111 by 0.1234243131324, guess what, humans would find that hard, computers easy. But it doesn't mean much. Humanity's last exam (HLE) couldn't be completed by most of humanity, the vast majority, so it doesn't really capture anything about humanity or mean much if a computer can do it.
- DroneBetter 6mo agothe point is that each question is something that a specialist in a field would be able to do, but deems challenging enough that the ability to solve it would imply significant general usefulness in that domain
- Leynos 6mo agoOpus 4.6 currently leads the remote labor index at 4.17. GPT-5.4 isn't measured on that one though: https://www.remotelabor.ai/ https://www.remotelabor.ai/ GPT 5.4 Pro leads Frontier Maths Tier 4 at 35%: https://epoch.ai/benchmarks/frontiermath-tier-4/ https://epoch.ai/benchmarks/frontiermath-tier-4/
- mbesto 6mo ago> We're gonna need some new benchmarks... You can't consistently benchmark something that is qualitative by nature. I'm struggling to understand how people don't understand this.