10 ms·
The biggest story in AI was released a few weeks ago but was given little attention: on the recent USAMO, SOTA models scored on average 5% (IIRC, it was some ab
by InkCanon 2y ago
The biggest story in AI was released a few weeks ago but was given little attention: on the recent USAMO, SOTA models scored on average 5% (IIRC, it was some abysmal number). This is despite them supposedly having gotten 50%, 60% etc performance on IMO questions. This massively suggests AI models simply remember the past results, instead of actually solving these questions. I'm incredibly surprised no one mentions this, but it's ridiculous that these companies never tell us what (if any) efforts have been made to remove test data (IMO, ICPC, etc) from train data.
- AIPedant 2y agoYes, here's the link: https://arxiv.org/abs/2503.21934v1 https://arxiv.org/abs/2503.21934v1 Anecdotally, I've been playing around with o3-mini on undergraduate math questions: it is much better at "plug-and-chug" proofs than GPT-4, but those problems aren't independently interesting, they are explicitly pedagogical. For anything requiring insight, it's either: 1) A very good answer that reveals the LLM has seen the problem before (e.g. naming the theorem, presenting a "standard" proof, using a much more powerful result) 2) A bad answer that looks correct and takes an enormous amount of effort to falsify. (This is the secret sauce of LLM hype.) I dread undergraduate STEM majors using this thing - I asked it a problem about rotations and spherical geometry, but got back a pile of advanced geometric algebra, when I was looking for "draw a spherical triangle." If I didn't know the answer, I would have been badly confused. See also this real-world example of an LLM leading a recreational mathematician astray: https://xcancel.com/colin_fraser/status/1900655006996390172#m https://xcancel.com/colin_fraser/status/1900655006996390172#... I will add that in 10 years the field will be intensely criticized for its reliance on multiple-choice benchmarks; it is not surprising or interesting that next-token prediction can game multiple-choice questions!
- JohnKemeny 2y agoDiscussed here: https://news.ycombinator.com/item?id=43540985 https://news.ycombinator.com/item?id=43540985 (Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad, 4 points, 2 comments).
- otabdeveloper4 2y agoAnecdotally: schoolkids are at the leading edge of LLM innovation, and nowadays all homework assignments are explicitly made to be LLM-proof. (Well, at least in my son's school. Yours might be different.) This effectively makes LLMs useless for education. (Also sours the next generation on LLMs in general, these things are extremely lame to the proverbial "kids these days".)
- bambax 2y agoHow do you make homework assignments LLM-proof? There may be a huge business opportunity if that actually works, because LLMs are destroying education at a rapid pace.
- otabdeveloper4 2y agoYou just (lol) need to give non-standard problems and demand students to provide reasoning and explanations along with the answer. Yeah, LLMs can "reason" too, but it's obvious when the output comes from an LLM here. (Yes, that's a lot of work for a teacher. Gone are the days when you could just assign reports as homework.)
- itchyjunk 2y agoCan you provide sample questions that are "LLM proof" ?
- otabdeveloper4 2y agoIt's not about being "LLM-proff", it's about teacher involvement in making up novel questions and grading attentively. There's no magic trick.
- xeromal 2y agoPart of the proof is knowing your students and forcing an answer that will rat out whether they used an LLM. There is no universal question and it requires personal knowledge of each student. You're looking for something that doesn't exist.
- larodi 2y agoThis is a paper by INSAIT researchers - a very young institute which hired most of its PHD staff only in the last 2 years, basically onboarding anyone who wanted to be part of it. They were waiving their BG-GPT on national TV in the country as a major breakthrough, while it was basically was a Mistral fine-tuned model, that was eventually never released to the public, nor the training set. Not sure whether their (INSAIT's) agenda is purely scientific, as there's a lot of PR on linkedin by these guys, literally celebrating every PHD they get, which is at minimum very weird. I'd take anything they release with a grain of sand if not caution.
- apercu 2y agoIn my experience LLMs can't get basic western music theory right, there's no way I would use an LLM for something harder than that.
- waffletower 1y agoWhile I may be mistaken, but I don't believe that LLMs are trained on a large corpus of machine readable music representations, which would arguably be crucial to strong performance in common practice music theory. I would also surmise that most music theory related datasets largely arrive without musical representations altogether. A similar problem exists for many other fields, particularly mathematics, but it is much more profitable to invest the effort to span such representation gaps for them. I would not gauge LLM generality on music theory performance, when its niche representations are likely unavailable in training and it is widely perceived as having miniscule economic value.
- code_for_monkey 1y agomusic theory is a really good test because in my experience the AI is extremely bad at it
- motorest 1y ago> In my experience LLMs can't get basic western music theory right, there's no way I would use an LLM for something harder than that. This take is completely oblivious, and frankly sounds like a desperate jab. There are a myriad of activities whose core requirement is a) derive info from a complex context which happens to be supported by a deep and plentiful corpus, b) employ glorified template and rule engines. LLMs excel at what might be described as interpolating context following input and output in natural language. As in a chatbot that is extensivey trained in domain-specific tasks, which can also parse and generate content. There is absolutely zero lines of intellectual work that do not benefit extensively from this sort of tool. Zero.
- apercu 1y agoA desperate jab? But I _want_ LLM's to be able to do basic, deterministic things accurately. Seems like I touched a nerve? Lol.
- simonw 2y agoI had to look up these acronyms: - USAMO - United States of America Mathematical Olympiad - IMO - International Mathematical Olympiad - ICPC - International Collegiate Programming Contest Relevant paper: https://arxiv.org/abs/2503.21934 https://arxiv.org/abs/2503.21934 - "Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad" submitted 27th March 2025.
- usaar333 2y agoAnd then within a week, Gemini 2.5 was tested and got 25%. Point is AI is getting stronger. And this only suggested LLMs aren't trained well to write formal math proofs, which is true.
- selcuka 2y ago> within a week How do we know that Gemini 2.5 wasn't specifically trained or fine-tuned with the new questions? I don't buy that a new model could suddenly score 5 times better than the previous state-of-the-art models.
- levocardia 2y agoThey retrained their model less than a week before its release, just to juice one particular nonstandard eval? Seems implausible. Models get 5x better at things all the time. Challenges like the Winograd schema have gone from impossible to laughably easy practically overnight. Ditto for "Rs in strawberry," ferrying animals across a river, overflowing wine glass, ...
- AIPedant 2y agoThe "ferrying animals across a river" problem has definitely not been solved, they still don't understand the problem at all, overcomplicating it because they're using an off-the-shelf solution instead of actual reasoning: o1 screwing up a trivially easy variation: https://xcancel.com/colin_fraser/status/1864787124320387202 https://xcancel.com/colin_fraser/status/1864787124320387202 Claude 3.7, utterly incoherent: https://xcancel.com/colin_fraser/status/1898158943962271876 https://xcancel.com/colin_fraser/status/1898158943962271876 DeepSeek: https://xcancel.com/colin_fraser/status/1882510886163943443#m https://xcancel.com/colin_fraser/status/1882510886163943443#... Overflowing wine glass also isn't meaningfully solved! I understand it is sort of solved for wine glasses (even though it looks terrible and unphysical, always seems to have weird fizz). But asking GPT to "generate an image of a transparent vase with flowers which has been overfilled with water, so that water is spilling over" had the exact same problem as the old wine glasses: the vase was clearly half-full, yet water was mysteriously trickling over the sides. Presumably OpenAI RLHFed wine glasses since it was a well-known failure, but (as always) this is just whack-a-mole, it does not generalize into understanding the physical principle.
- AstroBen 2y agoThis seems fairly obvious at this point. If they were actually reasoning at all they'd be capable (even if not good) of complex games like chess Instead they're barely able to eek out wins against a bot that plays completely random moves: https://maxim-saplin.github.io/llm_chess/ https://maxim-saplin.github.io/llm_chess/
- kylebyte 2y agoEvery day I am more convinced that LLM hype is the equivalent of someone seeing a stage magician levitate a table across the stage and assuming this means hovercars must only be a few years away.
- Terr_ 2y agoI believe there's a widespread confusion between a fictional character that is described as a AI assistant, versus the actual algorithm building the play-story which humans imagine the character from. An illusion actively promoted by companies seeking investment and hype. AcmeAssistant is "helpful" and "clever" in the same way that Vampire Count Dracula is "brooding" and "immortal".
- famouswaffles 2y agoLLMs are capable of playing chess and 3.5 turbo instruct does so quite well (for a human) at 1800 ELO. Does this mean they can truly reason now ? https://github.com/adamkarvonen/chess_gpt_eval https://github.com/adamkarvonen/chess_gpt_eval
- hatefulmoron 2y ago3.5 turbo instruct is a huge outlier. https://dynomight.substack.com/p/chess https://dynomight.substack.com/p/chess Discussion here: https://news.ycombinator.com/item?id=42138289 https://news.ycombinator.com/item?id=42138289
- famouswaffles 2y ago
- bglazer 2y agoYeah I’m a computational biology researcher. I’m working on a novel machine learning approach to inferring cellular behavior. I’m currently stumped why my algorithm won’t converge. So, I describe the mathematics to ChatGPT-o3-mini-high to try to help reason about what’s going on. It was almost completely useless. Like blog-slop “intro to ML” solutions and ideas. It ignores all the mathematical context, and zeros in on “doesn’t converge” and suggests that I lower the learning rate. Like, no shit I tried that three weeks ago. No amount of cajoling can get it to meaningfully “reason” about the problem, because it hasn’t seen the problem before. The closest point in latent space is apparently a thousand identical Medium articles about Adam, so I get the statistical average of those. I can’t stress how frustrating this is, especially with people like Terence Tao saying that these models are like a mediocre grad student. I would really love to have a mediocre (in Terry’s eyes) grad student looking at this, but I can’t seem to elicit that. Instead I get low tier ML blogspam author. **PS** if anyone read this far (doubtful) and knows about density estimation and wants to help my email is bglazer1@gmail.com I promise its a fun mathematical puzzle and the biology is pretty wild too
- root_axis 2y agoIt's funny, I have the same problem all the time with typical day to day programming roadblocks that these models are supposed to excel at. I'm talking about any type of bug or unexpected behavior that requires even 5 minutes of deeper analysis. Sometimes when I'm anxious just to get on with my original task, I'll paste the code and output/errors into the LLM and iterate over its solutions, but the experience is like rolling dice, cycling through possible solutions without any kind of deductive analysis that might bring it gradually closer to a solution. If I keep asking, it eventually just starts cycling through variants of previous answers with solutions that contradict the established logic of the error/output feedback up to this point. Not to say that the LLMs aren't productive tools, but they're more like calculators of language than agents that reason.
- jwrallie 2y agoTrue. There’s a small bonus that trying to explain the issue to the llm may sometimes be essentially rubber ducking, and that can lead to insights. I feel most of the time the llm can give erroneous output that still might trigger some thinking on a different direction, and sometimes I’m inclined to think it’s helping me more than it actually is.
- sanxiyn 2y agoNope, no LLMs reported 50~60% performance on IMO, and SOTA LLMs scoring 5% on USAMO is expected. For 50~60% performance on IMO, you are thinking of AlphaProof, but AlphaProof is not a LLM. We don't have the full paper yet, but clearly AlphaProof is a system built on top of LLM with lots of bells and whistles, just like AlphaFold is.
- InkCanon 2y agoo1 reportedly got 83% on IMO, and 89th percentile on Codeforces. https://openai.com/index/learning-to-reason-with-llms/ https://openai.com/index/learning-to-reason-with-llms/ The paper tested it on o1-pro as well. Correct me if I'm getting some versioning mixed up here.
- alexlikeits1999 2y agoI've gone through the link you posted and the o1 system card and can't see any reference to IMO. Are you sure they were referring to IMO or were they referring to AIME?
- sanxiyn 2y agoAIME is so not IMO.
- billforsternz 2y agoI asked Google "how many golf balls can fit in a Boeing 737 cabin" last week. The "AI" answer helpfully broke the solution into 4 stages; 1) A Boeing 737 cabin is about 3000 cubic metres [wrong, about 4x2x40 ~ 300 cubic metres] 2) A golf ball is about 0.000004 cubic metres [wrong, it's about 40cc = 0.00004 cubic metres] 3) 3000 / 0.000004 = 750,000 [wrong, it's 750,000,000] 4) We have to make an adjustment because seats etc. take up room, and we can't pack perfectly. So perhaps 1,500,000 to 2,000,000 golf balls final answer [wrong, you should have been reducing the number!] So 1) 2) and 3) were out by 1,1 and 3 orders of magnitude respectively (the errors partially cancelled out) and 4) was nonsensical. This little experiment made my skeptical about the state of the art of AI. I have seen much AI output which is extraordinary it's funny how one serious fail can impact my point of view so dramatically.
- Sunspark 2y agoIt's fascinating to me when you tell one that you'd like to see translated passages of work from authors who never have written or translated the item in question, especially if they passed away before the piece was written. The AI will create something for you and tell you it was them.
- prawn 2y ago"That's impossible because..." "Good point! Blah blah blah..." Absolutely shameless!
- senordevnyc 2y agoJust tried with o3-mini-high and it came up with something pretty reasonable: https://chatgpt.com/share/67f35ae9-5ce4-800c-ba39-6288cb4685f0 https://chatgpt.com/share/67f35ae9-5ce4-800c-ba39-6288cb4685...
- CamperBob2 2y agoIt's just the usual HN sport: ask a low-end, obsolete or unspecified model, get a bad answer, brag about how you "proved" AI is pointless hype, collect karma. Edit: Then again, maybe they have a point, going by an answer I just got from Google's best current model ( https://g.co/gemini/share/374ac006497d https://g.co/gemini/share/374ac006497d ) I haven't seen anything that ridiculous from a leading-edge model for a year or more.
- cma 2y agoOpenAI told how they removed it for GPT-4 in its release paper: only exact string matches. So all discussion of bar exam questions from memory on test taking forums etc., that wouldnn't exactly match, made it in.
- geuis 2y agoQuery: Could you explain the terminology to people who don't follow this that closely?
- BlanketLogic 2y agoNot the OP but USAMO : USA Math Olympiad. Referred here https://arxiv.org/pdf/2503.21934v1 https://arxiv.org/pdf/2503.21934v1 IMO : International Math Olympiad SOTA : State of the Art OP is probably referring to this referred to this paper here https://arxiv.org/pdf/2503.21934v1 https://arxiv.org/pdf/2503.21934v1. The paper explains out how a rigorous testing revealed abysmal performance of LLMs (results that are at odds with how they are hyped about).
- KolibriFly 2y agoYeah, this is one of those red flags that keeps getting hand-waved away, but really shouldn't be.
- TrackerFF 2y agoWhat would the average human score be? I.e. if you randomly sampled N humans to take those tests.
- sanxiyn 2y agoThe average human score on USAMO (let alone IMO) is zero, of course. Source: I won medals at Korean Mathematical Olympiad.
- vintermann 2y agoAverage, hmmm?
- lordgrenville 2y agoI am hesitant to correct a math Olympian, but don't you mean the median?
- nhinck3 1y agoAverage is fine.
- hyperbovine 2y agoThis is a disappointing answer from an MO alum. Pick a quantile, any quantile...
- anonzzzies 2y agoThat type of news might make investors worry / scared.
- sigmoid10 2y ago>I'm incredibly surprised no one mentions this If you don't see anyone mentioning what you wrote that's not surprising at all, because you totally misunderstood the paper. The models didn't suddenly drop to 5% accuracy on math olympiad questions. Instead this paper came up with a human evaluation that looks at the whole reasoning process (instead of just the final answer) and their finding is that the "thoughts" of reasoning models are not sufficiently human understandable or rigorous (at least for expert mathematicians). This is something that was already well known, because "reasoning" is essentially CoT prompting baked into normal responses. But the empirics also tell us it greatly helps for final outputs nonetheless.
- Workaccount2 2y agoOn top of that, what the model prints out in the CoT window is not necessarily what the model is actually thinking. Anthropic just showed this in their paper from last week where they got models to cheat at a question by "accidentally" slipping them the answer, and the CoT had no mention of answer being slipped to them.
- yahoozoo 2y agoLLMs are “next token” predictors. Yes, I realize that there’s a bit more to it and it’s not always just the “next” token, but at a very high level that’s what they are. So why are we so surprised when it turns out they can’t actually “do” math? Clearly the high benchmark scores are a result of the training sets being polluted with the answers.
- colonial 2y agoLess than 5%. OpenAI's O1 burned through over $100 in tokens during the test as well!
- hyperbovine 2y agoIs that really so surprising given what we know about how these models actually work? I feel vindicated on behalf of myself and all the other commenters who have been mercilessly downvoted over the past three years for pointing out the obvious fact that next token prediction != reasoning.
- aoeusnth1 2y ago2.5 pro scores 25%. It’s just a much harder math benchmark which will fall by the end of next year just like all the others. You won’t be vindicated.
- hyperbovine 2y agoBold claim! Let's see what that 25% is. I guarantee it is the portion of the exam which is trivially answerable if you have a stored database of all previous math exams ever written to consult.
- aoeusnth1 2y agoThere is 0% of the exam which is trivially answerable. The entire point of USAMO problems is that they demand novel insight and rigorous, original proofs. They are intentionally designed not to be variations of things you can just look up. You have to reason your way through, step by logical step. Getting 25% (~11 points) is exceptionally difficult. That often means fully solving one problem and maybe getting solid partial credit on another. The median score is often in the single digits.
- hyperbovine 1y ago> There is 0% of the exam which is trivially answerable. That's true, but of course, not what I claimed. The claim is that, given the ability to memorize an every mathematical result that has ever been published (in print or online), it is not so difficult to get 25% correct on an exam by pattern matching. Note that this is skill is, by definition, completely out of the reach of any human being, but that possessing it does not imply creativity or the ability to "think".
- utopcell 1y agoThis is simply using LLMs directly. Google has demonstrated that this is not the way to go when it comes to solving math problems. AlphaProof, which used AlphaZero code, got a silver medal in last year's IMO. It also didn't use any human proofs(!), only theorem statements in lean, without their corresponding proofs [1]. [1] https://www.youtube.com/watch?v=zzXyPGEtseI https://www.youtube.com/watch?v=zzXyPGEtseI
- SergeAx 1y agoBecause of the vast number of problems reused, removing those data from training sets will just make models worse. Why would anyone do it?