4 ms·
> We have models that are doing better than humans at IMO. Not really. From my brief experience they can guess the final answer but the intermediate justificat
by otabdeveloper4 8mo ago
> We have models that are doing better than humans at IMO.
Not really. From my brief experience they can guess the final answer but the intermediate justifications and proofs are complete hallucinated bullshit.
(Possibly because the final answer is usually some sort of neat and beatiful answer and human evaluators don't care about the final answer anyways, in any olympiad you're graded on the soundness of your reasoning.)
- simianwords 8mo agowhat's the best way to falsify it?
- tveita 8mo agoYou could start by reading research on the topic instead of disregarding expert opinion based on your own gut feeling E.g. https://www.anthropic.com/research/tracing-thoughts-language-model#mental-math https://www.anthropic.com/research/tracing-thoughts-language...
- simianwords 8mo agoIt’s specific on Claude.
- otabdeveloper4 8mo agoFalsify what? The claim that LLM's are good for olympiad problems? I'm just an end user who tried to use these "frontier models" to actually solve real olympiad problems. They're useless.