3 ms·
Something always seemed incomplete about testing models against standardized tests; I would expect AI models to first do well on standardized tests, much better
by techfeathers 2y ago
Something always seemed incomplete about testing models against standardized tests; I would expect AI models to first do well on standardized tests, much better than humans, but it makes me wonder if there’s something else that humans possess that isn’t tested by these tests. We test humans in these tests to, and I would guess that loosely speaking there’s a correlation between a persons success on an advanced math test or the bar and success in their career, but we also sort of know examples where there’s an appearance of an inverse correlation, people who do great as a phd student or mathlete who can’t operate in a day to day job.
So when AI companies start saying these AI are as intelligent as a PHD student, it makes me wonder, most people aren’t as smart as a phd student and yet AI still seems to choke on some basic tasks.
- pants2 2y agoAI companies love standardized tests because that's what LLMs are really good at: short problems with objective solutions. LLMs do worse on problems that are more open-ended, like SWE-bench, which GPT-4 achieved a whopping 1.74% on.