11 ms·
The headline may make it seem like AI just discovered some new result in physics all on its own, but reading the post, humans started off trying to solve some p
by outlace 8mo ago
The headline may make it seem like AI just discovered some new result in physics all on its own, but reading the post, humans started off trying to solve some problem, it got complex, GPT simplified it and found a solution with the simpler representation. It took 12 hours for GPT pro to do this. In my experience LLM’s can make new things when they are some linear combination of existing things but I haven’t been to get them to do something totally out of distribution yet from first principles.
- getnormality 8mo ago[flagged]
- emil-lp 8mo ago"GPT did this". Authored by Guevara (Institute for Advanced Study), Lupsasca (Vanderbilt University), Skinner (University of Cambridge), and Strominger (Harvard University). Probably not something that the average GI Joe would be able to prompt their way to... I am skeptical until they show the chat log leading up to the conjecture and proof.
- famouswaffles 8mo agoThe paper has all those prominent institutions who acknowledge the contribution so realistically, why would you be skeptical ?
- kristopolous 8mo agothey probably also acknowledge pytorch, numpy, R ... but we don't attribute those tools as the agent who did the work. I know we've been primed by sci-fi movies and comic books, but like pytorch, gpt-5.2 is just a piece of software running on a computer instrumented by humans.
- famouswaffles 8mo agoI don't see the authors of those libraries getting a credit on the paper, do you ? >I know we've been primed by sci-fi movies and comic books, but like pytorch, gpt-5.2 is just a piece of software running on a computer instrumented by humans. Sure
- what 8mo agoAt least one of the authors works for OpenAI. It’s a puff piece.
- name_taken_duh 8mo agoAnd we are just a system running on carbon-based biology in our physics computer run by whomever. What makes us special, to say that we are different than GPT-5.2?
- palmotea 8mo ago> And we are just a system running on carbon-based biology in our physics computer run by whomever. What makes us special, to say that we are different than GPT-5.2? Do you really want to be treated like an old PC (dismembered, stripped for parts, and discarded) when your boss is done with you (i.e. not treated specially compared to a computer system)? But I think if you want a fuller answer, you've got a lot of reading to do. It's not like you're the first person in the world to ask that question.
- name_taken_duh 8mo agoYou misunderstood, I am prohumanism. My comment was about challenging the believe that models cant be as intelligent as we are, which cant be answered definitely, though a lot of empirical evidence seems to point to the fact, that we are not fundamentally different intelligence wise. Just closing our eyes will not help in preserving humanism, so we have to shape the world with models in a human friendly way, aka alignment.
- kristopolous 8mo agoIt's always a value decision. You can say shiny rocks are more important than people and worth murdering over. Not an uncommon belief. Here you are saying you personally value a computer program more than people It exposes a value that you personally hold and that's it That is separate from the material reality that all this AI stuff is ultimately just computer software... It's an epistemological tautology in the same way that say, a plane, car and refrigerator are all just machines - they can break, need maintenance, take expertise, can be dangerous... LLMs haven't broken the categorical constraints - you've just been primed to think such a thing is supposed to be different through movies and entertainment. I hate to tell you but most movie AIs are just allegories for institutional power. They're narrative devices about how callous and indifferent power structures are to our underlying shared humanity
- Refreeze5224 8mo agoTheir point is, would you be able to prompt your way to this result? No. Already trained physicists working at world-leading institutions could. So what progress have we really made here?
- famouswaffles 8mo agoIt's a stupid point then. Are you able to work with a world leading physicist to any significant degree? No
- emil-lp 8mo agoIt's like saying: calculator drives new result in theoretical physics (In the hands of leading experts.)
- famouswaffles 8mo agoNo it's not like saying that at all, which is why Open AI have a credit on the paper.
- not_kurt_godel 8mo agoAnd even if it were, calculators (computers) were world-changing technology when they were new.
- camdenreslink 8mo agoOpen AI have a credit on the paper because it is marketing.
- famouswaffles 8mo agoLol Okay
- Nevermark 8mo agoNo it’s like saying: New expert drives new results with existing experts. The humans put in significant effort and couldn’t do it. They didn’t then crank it out with some search/match algorithm. They tried a new technology, modeled (literally) on us as reasoners, that is only just being able to reason at their level and it did what they couldn’t. The fact that the experts were a critical context for the model, doesn’t make the models performance any less significant. Collaborators always provide important context for each other.
- Sharlin 8mo agoI'm a big LLM sceptic but that's… moving the goalposts a little too far. How could an average Joe even understand the conjecture enough to write the initial prompt? Or do you mean that experts would give him the prompt to copy-paste, and hope that the proverbial monkey can come up with a Henry V? At the very least posit someone like a grad student in particle physics as the human user.
- slopusila 8mo agohey, GPT, solve this tough conjecture I've read about on Quanta. make no mistakes
- co_king_3 8mo ago[dead]
- terminalbraid 8mo ago"Hey GPT thanks for the result. But is it actually true?"
- buttered_toast 8mo agoI would interpret it as implying that the result was due to a lot more hand-holding that what is let on. Was the initial conjecture based on leading info from the other authors or was it simply the authors presenting all information and asking for a conjecture? Did the authors know that there was a simpler means of expressing the conjecture and lead GPT to its conclusion, or did it spontaneously do so on its own after seeing the hand-written expressions. These aren't my personal views, but there is some handwaving about the process in such a way that reads as if this was all spontaneous involvement on GPTs end. But regardless, a result is a result so I'm content with it.
- lupsasca 8mo agoHi I am an author of the paper. We believed that a simple formula should exist but had not been able to find it despite significant effort. It was a collaborative effort but GPT definitely solved the problem for us.
- hgfda 8mo ago[dead]
- jmalicki 8mo ago"Grad Student did this". Co-authored by <Famous advisor 1>, <Famous advisor 2>, <Famous advisor 3>. Is this so different?
- sejje 8mo agoThe Average Joe reads at an 8th grade level. 21% are illiterate in the US. LLMs surpassed the average human a long time ago IMO. When LLMs fail to measure up to humans, it's that they fail to measure up against human experts in a given field, not the Average Joe. We are surrounded by NPCs.
- bpodgursky 8mo agoI don't want to be rude but like, maybe you should pre-register some statement like "LLMs will not be able to do X" in some concrete domain, because I suspect your goalposts are shifting without you noticing. We're talking about significant contributions to theoretical physics. You can nitpick but honestly go back to your expectations 4 years ago and think — would I be pretty surprised and impressed if an AI could do this? The answer is obviously yes, I don't really care whether you have a selective memory of that time.
- deleted 8mo ago[deleted]
- nozzlegear 8mo ago> We're talking about significant contributions to theoretical physics. Whoever wrote the prompts and guided ChatGPT made significant contributions to theoretical physics. ChatGPT is just a tool they used to get there. I'm sure AI-bloviators and pelican bike-enjoyers are all quite impressed, but the humans should be getting the research credit for using their tools correctly. Let's not pretend the calculator doing its job as a calculator at the behest of the researcher is actually a researcher as well.
- famouswaffles 8mo agoIf this worked for 12 hours to derive the simplified formula along with its proof then it guided itself and made significant contributions by any useful definition of the word, hence Open AI having an author credit.
- nozzlegear 8mo ago> hence Open AI having an author credit. How much precedence is there for machines or tools getting an author credit in research? Genuine question, I don't actually know. Would we give an author credit to e.g. a chimpanzee if it happened to circle the right page of a text book while working with researchers, leading them to a eureka moment?
- CGMthrowaway 8mo agoThis is the critical bit (paraphrasing): Humans have worked out the amplitudes for integer n up to n = 6 by hand, obtaining very complicated expressions, which correspond to a “Feynman diagram expansion” whose complexity grows superexponentially in n. But no one has been able to greatly reduce the complexity of these expressions, providing much simpler forms. And from these base cases, no one was then able to spot a pattern and posit a formula valid for all n. GPT did that. Basically, they used GPT to refactor a formula and then generalize it for all n. Then verified it themselves. I think this was all already figured out in 1986 though: https://journals.aps.org/prl/abstract/10.1103/PhysRevLett.56.2459 https://journals.aps.org/prl/abstract/10.1103/PhysRevLett.56... see also https://en.wikipedia.org/wiki/MHV_amplitudes https://en.wikipedia.org/wiki/MHV_amplitudes
- ericmay 8mo agoStill pretty awesome though, if you ask me.
- woeirua 8mo agoYou should probably email the authors if you think that's true. I highly doubt they didn't do a literature search first though...
- emp17344 8mo agoYou should be more skeptical of marketing releases like this. This is an advertisement.
- 8mo ago
- bottlepalm 8mo agoIs every new thing not just combinations of existing things? What does out of distribution even mean? What advancement has ever made that there wasn’t a lead up of prior work to it? Is there some fundamental thing that prevents AI from recombining ideas and testing theories?
- outlace 8mo agoFor example, ever since the first GPT 4 I’ve tried to get LLM’s to build me a specific type of heart simulation that to my knowledge does not exist anywhere on the public internet (otherwise I wouldn’t try to build it myself) and even up to GPT 5.3 it still cannot do it. But I’ve successfully made it build me a great Poker training app, a specific form that also didn’t exist, but the ingredients are well represented on the internet. And I’m not trying to imply AI is inherently incapable, it’s just an empirical (and anecdotal) observation for me. Maybe tomorrow it’ll figure it out. I have no dogmatic ideology on the matter.
- fpgaminer 8mo ago> Is every new thing not just combinations of existing things? If all ideas are recombinations of old ideas, where did the first ideas come from? And wouldn't the complexity of ideas be thus limited to the combined complexity of the "seed" ideas? I think it's more fair to say that recombining ideas is an efficient way to quickly explore a very complex, hyperdimensional space. In some cases that's enough to land on new, useful ideas, but not always. A) the new, useful idea might be _near_ the area you land on, but not exactly at. B) there are whole classes of new, useful ideas that cannot be reached by any combination of existing "idea vectors". Therefore there is still the necessity to explore the space manually, even if you're using these idea vectors to give you starting points to explore from. All this to say: Every new thing is a combination of existing things + sweat and tears. The question everyone has is, are current LLMs capable of the latter component. Historically the answer is _no_, because they had no real capacity to iterate. Without iteration you cannot explore. But now that they can reliably iterate, and to some extent plan their iterations, we are starting to see their first meaningful, fledgling attempts at the "sweat and tears" part of building new ideas.
- verdverm 8mo ago[flagged]
- buttered_toast 8mo agoAbsolutely no way this is true right? Ilya left around the time 4o was released. I can't imagine they haven't had a single successful run since then.
- verdverm 8mo agoWhen's the last time they talked about it? I heard this from people who know more than me
- buttered_toast 8mo agoCan't say, just seems implausible, but I am a nobody anyways ¯\_(ツ)_/¯
- verdverm 8mo agoI'm pretty sure it is widely known that the early 5.x series were built from 4.5 (unreleased). It seems more plausible the 5.x series is still in that continuation. For some extra context, pre-training is ~1/3 of the training, where it gains the basic concepts of how tokens go together. Mid & late training are where you instill the kinds of anthropic behaviors we see today. I expect pre-training to increasingly become a lower percentage of overall training, putting aside any shifts of what happens in each phase. So to me, it is plausible they can take the 4.x pre-training and keep pushing in the later phases. There is a lot of results out there to show scaling laws (limits) have not peaked yet. I would not be surprised to learn that Gemini 3 Deep Research had 50% late-training / RL
- buttered_toast 8mo agoOkay I see what you mean, and yeah that sounds reasonable too. Do you have any context on that first part? I would like to know more about how/why they might not have been able to pursue more training runs.
- ctoth 8mo agoIn my experience humans can make new things when they are some linear combination of existing things but I haven’t been able to get them to do something totally out of distribution yet from first principles[0]. [0]: https://slatestarcodex.com/2019/02/19/gpt-2-as-step-toward-general-intelligence/ https://slatestarcodex.com/2019/02/19/gpt-2-as-step-toward-g...
- randomtoast 8mo ago> but I haven’t been to get them to do something totally out of distribution yet from first principles Can humans actually do that? Sometimes it appears as if we have made a completely new discovery. However, if you look more closely, you will find that many events and developments led up to this breakthrough, and that it is actually an improvement on something that already existed. We are always building on the shoulders of giants.
- dotancohen 8mo agoRelativity comes to mind. You could nitpick a rebuttal, but no matter how many people you give credit, general relativity was a completely novel idea when it was proposed. I'd argue for special relatively as well.
- johnfn 8mo agoEven if I grant you that, surely we’ve moved the goal posts a bit if we’re saying the only thing we can think of that AI can’t do is the life’s work of a man who’s last name is literally synonymous with genius.
- poplarsol 8mo agoThat's not exactly true. Lorentz contraction is a clear antecedent to special relativity.
- deleted 8mo ago[deleted]
- DonaldFisk 8mo agoIt isn't an anteceent, it's part of special relativity, discovered by Lorentz. It's well known that special relativity is the work of several people as well as Einstein.
- lamontcg 8mo agoNot really. Pretty sure I read recently that Newton appreciated that his theory was non-local and didn't like what Einstein later called "spooky action at a distance". The Lorentz transform was also known from 1887. Time dilation was understood from 1900. Poincaré figured out in 1905 that it was a mathematical group. Einstein put a bow on it all by figuring out that you could derive it from the principle of relativity and keeping the speed of light constant in all inertial reference frames. I'm not sure about GR, but I know that it is built on the foundations of differential geometry, which Einstein definitely didn't invent (I think that's the source of his "I assure you whatever your difficulties in mathematics are, that mine are much greater" quote because he was struggling to understand Hilbert's math). And really Cauchy, Hilbert, and those kinds of mathematicians I'd put above Einstein in building entirely new worlds of mathematics...
- epolanski 8mo agoSerious questions, I often hear about this "let the LLM cook for hours" but how do you do that in practice and how does it manages its own context? How doesn't it get lost at all after so many tokens?
- javier123454321 8mo agoFrom what I've seen is a process of compacting the session once it reaches some limit, which basically means summarizing all the previous work and feeding it as the initial prompt for the next session.
- lovecg 8mo agoI’m guessing, would love someone who has first hand knowledge to comment. But my guess is it’s some combination of trying many different approaches in parallel (each in a fresh context), then picking the one that works, and splitting up the task into sequential steps, where the output of one step is condensed and is used as an input to the next step (with possibly human steering between steps)
- 8note 8mo agothe annoying part is that with tool calls, a lot of those hours is time spent on netowrk round trips. over long periods of time, checklists are the biggest thing, so the LLM can track whats already done and whats left. after a compact, it can pull the relevant stuff back up and make progress. having some level or hierarchy is also useful - requirements, high level designs, low level designs, etc
- amelius 8mo agoJust wait until LLMs are fast and cheap enough to be run in a breadth first search kind of way, with "fuzzy" pruning.
- stouset 8mo agoWhen chess engines were first developed, they were strictly worse than the best humans. After many years of development, they became helpful to even the best humans even though they were still beatable (1985–1997). Eventually they caught up and surpassed humans but the combination of human and computer was better than either alone (~1997–2007). Since then, humans have been more or less obsoleted in the game of chess. Five years ago we were at Stage 1 with LLMs with regard to knowledge work. A few years later we hit Stage 2. We are currently somewhere between Stage 2 and Stage 3 for an extremely high percentage of knowledge work. Stage 4 will come, and I would wager it's sooner rather than later.
- bluecalm 8mo agoThe evolution was also interesting: first the engines were amazing tactically but pretty bad strategically so humans could guide them. With new NN based engines they were amazing strategically but they sucked tactically (first versions of Leela Chess Zero). Today they closed the gap and are amazing at both strategy and tactics and there is nothing humans can contribute anymore - all that is left is to just watch and learn.
- deleted 8mo ago[deleted]
- TGower 8mo agoWith a chess engine, you could ask any practitioner in the 90's what it would take to achieve "Stage 4" and they could estimate it quite accurately as a function of FLOPs and memory bandwidth. It's worth keeping in mind just how little we understand about LLM capability scaling. Ask 10 different AI researchers when we will get to Stage 4 for something like programming and you'll get wild guesses or an honest "we don't know".
- baq 8mo agoChess grandmasters are living proof that it’s possible to reach grandmaster level in chess on 20W of compute. We’ve got orders of magnitude of optimizations to discover in LLMs and/or future architectures, both software and hardware and with the amount of progress we’ve got basically every month those ten people will answer ‘we don’t know, but it won’t be too long’. Of course they may be wrong, but the trend line is clear; Moore’s law faced similar issues and they were successively overcome for half a century. IOW respect the trend line.
- hellisad 8mo agoHmm feels a bit trivializing, we don't know exactly how difficult it was to come up with the generic set of equations mentioned from the human starting point. I can claim some knowledge of physics from my degree, typically the easy part is coming up with complex dirty equations that work under special conditions, the hard part is the simplification into something elegant, 'natural' and general. Also "LLM’s can make new things when they are some linear combination of existing things" Doesn't really mean much, what is a linear combination of things you first have to define precisely what a thing is?
- slibhb 8mo ago> In my experience LLM’s can make new things when they are some linear combination of existing things but I haven’t been to get them to do something totally out of distribution yet from first principles. What's the distinction between "first principles" and "existing things"? I'm sympathetic to the idea that LLMs can't produce path-breaking results, but I think that's true only for a strict definition of path-breaking (that is quite rare for humnans too).
- tedd4u 8mo agoWhat does a 12-hour solution cost an OpenAI customer?
- int_19h 8mo ago$200/month would cover many such sessions every month. The real question is, what does it cost OpenAI? I'm pretty sure both their plans are well below cost, at least for users who max them out (and if you pay $200 for something then you'll probably do that!). How long before the money runs out? Can they get it cheap enough to be profitable at this price level, or is this going to be "get them addicted then jack it up" kind of strategy?
- revlsas 8mo agoNo because open source models are close behind Compute costs will fall drastically for existing models But it's likely that frontier models of the future won't be released to the public at all, because they'll be too good
- int_19h 7mo agoI have yet to see an open model that is close to the previous gen of frontier models. Benchmarks are one thing, real world performance is very different.
- malshe 8mo agoMy physics professor once claimed that imagination is just mental manipulation of past experiences. I never thought it was true for human beings but for LLMs it makes perfect sense.
- waynesonfire 8mo ago> It took 12 hours for GPT pro to do this Thanks for the summary; but this is a huge hand-wave. was GPT Pro just spinning for 12 hours and returend 42?!
- deleted 8mo ago[deleted]
- acchow 8mo ago> In my experience LLM’s can make new things when they are some linear combination of existing things It seems to me that all “new ideas” are basically linear combinations of existing things with exceeding rare exceptions… Maybe Godel’s Incompleteness? Darwinian evolution? General Relativity? Buddhist non-duality?
- sathish316 8mo ago> I haven’t been to get them to do something totally out of distribution yet from first principles. Agree with this. I’ve been trying to make LLMs come up with creative and unique word games like Wordle and Uncrossy (uncrossy.com), but so far GPT-5.2 has been disappointing. Comparatively, Opus 4.5 has been doing better on this. But it’s good to know that it’s breaking new ground in Theoretical Physics!
- zaphirplane 8mo agoI must be a Luddite, how do you have a model working for 12 hours on a problem. Mine is ready with an answer and always interrupts to ask confirmation or show answer
- arjie 8mo agoThat's on the harness - the device actually sending the prompt to the model. You can write a different harness that feeds the problem back in for however long you want. Ask Claude Code or Codex to build it for you in as minimal a fashion as possible and you'll see that a naïve version is not particularly more complex than `while true; do prompt $file >> file; done` (though it's not that precisely, obviously).
- anon291 8mo agoVery very few human individuals are capable of making new things that are not a linear combination of existing things. Even such things as special relativity were an application of two previous ideas. All of special relativity is deriveable from the principles of relative motion (known into antiquity) and the constant speed of light (which was known to Einstein). From there it is a straightforwards application of the Pythagorean theorem to realize there is a contradiction and the lorentz factor falls out naturally via basic algebra.
- FranklinJabar 8mo agoSurely higher level math is just linear combinations of the syntax and implications of lower level math. LLMs are taught syntax of basically all existing math notation, I assume. Much of math is, after all, just linguistic manipulation and detection of contradiction in said language with a more formal, a priori language.
- MITSardine 8mo agoLLMs can write theorems, but can they come up with meaningful definitions?
- FranklinJabar 8mo agoI intended to imply this with "detection of contradiction". Coherence seems to me to be the only a priori meaning. Most of the meaning of "meaning" seems to me to be a posteriori. After all, what is the point of an a priori floating signifier?
- MITSardine 8mo agoSetting the framework (what I short-handed by "definitions") precludes the exploration of results (or at least an efficient one) that would yield to some framework-defining analysis. The search space is much too rich to be explored anything but greedily in timid steps off the trodden path, and the frameworks (arbitrary) set both the highways and the vehicles by which we move along and out of them. Now, the argument can be made that the "meta-mathematical" (but outright mathematical, really) setting of frameworks follows the same structure, and LLMs could also explore that space. Even assuming that, a major roadblock remains: mathematics should remain understandable by humans, and yield fast progress in desirable (by whom? until now, by humans) directions, so the constraints on admissible frameworks are not as simple as "yields to coherent results". Also, to take a step back, I wonder what the pertinence of using numerical math is to derive analytical math when we can already solve a great deal of problems through numerical methods. For instance, is it worth spending however many MWh on LLMs to derive an analytical solution to an optimization problem, which might itself be very expensive to compute (human-derived expressions tend to be particularly cheap to evaluate precisely because we are so limited; machines (formal calculus for instance) will happily give you multi-page formulae with thousands of operations to evaluate), when there's a vast array of algorithms at the ready to provide arbitrarily precise solutions? What remains is the kind of math that, arguably, is much more precious to understand than to derive.
- DeathArrow 8mo ago>LLM’s can make new things when they are some linear combination of existing things Aren't most new things linear combinations of existing things (up to a point)?
- bamboozled 8mo agoAll you have to do is see "openai.com" in the submission URL to know it's bullshit.
- Sparkyte 8mo agoAI cough LLMs don't discover things they simply surface information that already existed.
- slibhb 8mo agoYou're assuming there aren't "new things" latent inside currently existing information. That's definitely false, particulary for math/physics. But it's worth thinking more about this. What gives humans the ability to discover "new things"? I would say it's due to our interaction with the universe via our senses, and not due to some special powers intrinsic to our brains that LLMs lack. And the thing is, we can feed novel measurements to LLMs (or, eventually, hook them up to camera feeds to "give them senses")
- Sparkyte 8mo agoNo it isn't false. If it is new it is novel, novel because it is known to some degree and two other abstracted known things prove the third. Just pattern matching connecting dots.
- slibhb 8mo agoThe vast majority of work by mathematicians uses n abstracted known things to prove something that is unproven. In fact, there is a view in philosophy that all math consists only of this.
- mirsadm 8mo agoMy issue with any of these claims is the lack of proof. Just share the chat and now it got to the discovery. I'll believe it when I can see it for myself at this point. It's too easy to make all sorts of claims without proof these days. Elon Musk makes them all the time.