5 ms·
How good are frontier models at physics?
- qt31415926 17d agoArticle: "How Good Are Frontier Models at Physics? Expert Re-Grading Reveals Broken Evaluations and Near-Saturation of Leading Benchmarks" John Sous from Yale posted a fairly solid study on how nearly all the physics benchmarks are broken, as they consistently mark correct answers as incorrect. When hand grading instead, they found out that the models have actually already saturated the benchmarks which is a little bit scary.
- fsh 17d agoI would be very surprised if any of the frontier models wasn't trained on all public physics benchmarks. Training data providers have been hiring people for exactly this task.
- letmevoteplease 17d agoThis study appears to be evidence against that: the model failed the benchmark but arrived at the correct answer.
- fsh 17d agoHalf of the benchmarks don't have published answers. That's why training data providers have been hiring physicists to solve them.
- deleted 17d ago[deleted]
- red75prime 17d agoThis is unconventional benchmaxxing then, when they decrease the benchmark scores to allow models to generalize on correct solutions.
- bobmarleybiceps 17d agoyeah, it would be almost shocking if an open source benchmark was NOT used ~somewhere in training. Perhaps just pre-training, but still. Neural networks can be fairly robust to some mistakes in their training data, so maybe it doesn't even matter if some of them are incorrect. Who knows.
- redwood 17d agoI'd have thought the same but this article from yesterday blew my mind https://www.amazon.science/blog/why-dont-machine-learning-research-agents-overfit https://www.amazon.science/blog/why-dont-machine-learning-re... As it essential implies that these models compress knowledge well in a way that what remains is what's generalizeable more so than remembering every specific detail... Anyway more understanding necessary but thought provoking
- respectattentio 17d agoThis is interesting and actually very important for robotics. I've been waiting for this, but all companies seem to not care much now. There is a way out of this by supplying right context (needs a bit of expertise in physics) 1 more year and frontier will become crazy good at this as well.
- amluto 17d ago(Trained physicist here) From personal experience, frontier models absolutely struggle with understanding a physical situation based on words. (Okay, I haven't played with Astra much. GPT-5.6 Sol makes outrageous errors that anyone understanding a real world object would not make. And I was just asking it about NPT threads, not advanced physics.) But seriously, what's up with these benchmarks? The example question in the paper is: > PHYBench, problem 140: equivalent expressions for the same rope tension > Problem statement. Three identical homogeneous balls are placed on a smooth horizontal surface, touching each other and are close enough to each other. A rope is wrapped around the spheres at the height of their centers, tying them together. A fourth identical sphere is placed on top of the three spheres. Find the tension T in the rope. It is given that the weight of each sphere is P. For some reason the paper was focused on the fact that the grader didn't notice that some models were producing answers that were trivially algebraically equivalent to the reference answer. But this is missing the elephants in the room: 1. "touching each other and are close enough to each other": the right response is "hey, Professor, what do you mean 'close enough to each other'? They're sitting on a table in an equilateral triangle, all touching (i.e. tangent at their equators), right? Did you have a different configuration in mind?" 2. The answer is 0. Go find four baseballs or foursquare balls or whatever, make a little triangle with three of them, and balance the fourth one on top. It's not especially hard on an appropriate surface. Now loosely wrap an imaginary rope around them (but see below) to keep them from moving - no tension is needed because they're not moving anyway. So the models and the reference answer are wrong, IMO. 3. How, exactly, do you plan to wrap a rope around the spheres, at equator height, with no built-in tension (not pre-stretched), such that the rope does not immediately fall off? Friction? But I suspect you need to pretend there is no friction to get the reference answer. (Or maybe that the marble-marble interface has friction but the marble-table interface doesn't? Again, I haven't tried to reverse engineer it.) So maybe the right answer is "infinity or impossible -- in the scenario where the rope is needed, the rope will promptly fall off because it cannot be stable in the described configuration and gravity pulls it down, and once the rope falls off the tension will be zero and the top marble will fall and the other three will roll over the rope." 4. The answer might be "any tension you like -- just wrap the rope with the desired amount of tension". Imagine three baseballs in a triangle with a rubber band around them and a fourth baseball on top for good measure. The tension is a function of what rubber band you choose. I'm sure there's an interpretation of the question that makes the reference answer correct, and I was not inspired to try to reverse engineer it. My tentative conclusion is that LLMs are almost unbelievably good at solving problems that are fully contained within the inputs and (training/verification) outputs, and that they and the people training them are not actually particularly good at the input and output parts. If you are training a model to benchmaxx this benchmark, you are training a bad model.
- ArashEdalat 17d ago[flagged]
- Founderarcstone 17d agoNice thanks for sharing this!
- RomanKornev 17d agoStarting to feel more and more like chinese room experiment The models are confidently answering physics questions, treating it as a math problem, but they don't fundamentally "get it" and even recently failed simple "should i drive to car wash" test The sample efficiency is just crazy low Still surprising that even with this they managed to saturate the benchmarks
- red75prime 17d ago> they don't fundamentally "get it" There's no clear decision criteria for this. Do trick questions demonstrate that most people don't "get it"? And, well, older model saying dumb things doesn't establish a general principle that LLMs don't "get it" in general. > The sample efficiency is just crazy low Autoregressive pretraining requires huge amount of data to go from a blank state to a somewhat functional model. Fine-tuning, LORA, reinforcement learning of foundation models and in-context learning are much more sample efficient. > Chinese room ...creates a wrong intuition that by cranking a Leibniz's mill you are somehow responsible for whether it understands something or not.
- rogerrogerr 17d ago> Do trick questions demonstrate that most people don't "get it"? "Should I walk to the car wash" is hardly a trick question. If a human told me to walk to the car wash because it's so close, I would say that demonstrates they don't "get it".
- red75prime 17d agoI don't see that much difference with "A plane crashes on the border of the United States and Canada. Where do they bury the survivors?" Fast thinking fails to activate slow thinking.
- webern777 17d agoSearle's thought experiment just kicks the can down the road. Since the room as a whole produces fluent Chinese it becomes a language philosophy debate about where meaning comes from. Brandom's inferentialism gets into this more than any normal person would ever care to read about. LLMs really show the absurdity of the Chinese Room rulebook. The whole thought experiment should really be left to history at this point. Searle’s social ontology is much more interesting also but this outdated thought experiment completely overshadows his more interesting stuff.
- mch82 17d agoAre models able to do math now, or do they still rely on “tools” to do the math?
- nomel 17d agoWhy is this a concern? Besides a few savants, humans also use tools to do non-trivial math. I'm in engineering, and it's extremely rare to do anything non trivial in your head, because getting a decimal place wrong has real world consequences. Tools are just another way to say "deterministic", which is always nice.
- totallymike 17d agoWhy don’t we throw a data center off a cliff and find out
- deleted 14d ago[deleted]
- lee_ward 17d ago[dead]
- paidx 17d ago[flagged]
- quantumtwist 17d agoThe last author also wrote up a quite-readable blogpost here that accompanies the article: https://jsous.github.io/blogs/is-physics-dead/ https://jsous.github.io/blogs/is-physics-dead/ One tidbit I found particularly interesting: "We tested a GPT-based agentic system, which previously succeeded in resolving several open mathematical conjectures, on open problems in theoretical physics. To our temporary relief and encouragement, it has not managed to fully resolve even one autonomously. We also noted that these agents made considerably less progress on the partially resolved physics problems than on open mathematics problems of comparable difficulty, both by our analysis and independent agent-based analysis of partial results in each domain."
- treebeard901 17d ago> To our temporary relief and encouragement, it has not managed to fully resolve even one ... Why would that be encouraging if the purpose of their chosen profession is to advance physics? This is sort of like mathematics being solved by AI with human researchers complaining that they are losing potential awards. It's an entirely self serving way to be.
- webern777 17d agoThis is the same social process as to why science tends to progress one funeral at a time. So many people have their financial and self worth tied to their expertise in phlogiston theory that the truth is second to advancing the field of phlogiston theory. From that perspective it is hard to think of something worse than a machine that disproves phlogiston theory.
- aadyachinubhai 17d ago[flagged]