8 ms·
Both things can be true: 1. OpenAI when using your chats in pretraining is improving its model’s intuition. The model parameter size is massive, and while the
by sashank_1509 23d ago
Both things can be true:
1. OpenAI when using your chats in pretraining is improving its model’s intuition. The model parameter size is massive, and while the data is OOM larger it is plausible that model remembers stuff about chats that improves its latent representation.
2. During RL on verifiable math and massive compute, the model discovers techniques and connections to solve math problems that are superhuman and have little to do with some specific technique mentioned in its chat.
The rumor I’ve heard from multiple employees at OAI and Ant is that the model has solved hundreds of open problems in maths, and is basically solving anything you throw at it. We’ll know soon enough, but I’m inclined to believe this is true. Maths is a fully verifiable domain amenable to self play, massive scale RL can develop a search agent far better than any human and I’m inclined to believe OAI would have solved these conjectures without any of this chat data in its pre-training.
- Betelbuddy 23d agoJust use Bedrock...
- dgellow 23d agoI feel that we don’t praise Lean enough. AFAIU it’s what enables LLMs to brute force those problems
- iamgopal 23d agoTrue, but could humans cross pollinating lean x prolog x A* ( or any search algorithm) could have solved such math problems with super computer ?
- dgellow 23d agoI cannot say, math research isn’t my domain of expertise, I’m just trying to follow along :) But I find it interesting that Lean, a validator/compiler made by humans, is what enables those discoveries. But somehow all the praise goes to the models
- pixl97 23d agoI mean we don't instantly fall into ASI, hopefully. The problem with humans is every problem we solve the goal posts get kicked further down the road until they are reaching relativistic speeds. It starts around "well, the AI hasn't solved a novel problem" then moves to "well, they didn't write the validator" and suddenly humans are at the point of saying "Well AI hasn't rewrote the constants of the universe, what good are they". Of course another way to look at this is, the people that wrote the validator got praise for that years ago. Now and up and coming actor is solving problems that took us 100s of years to create in insanely short time periods so of course it's going to get a lot of attention as it well should.
- dgellow 23d agoTo be clear: I’m aware the LLMs are solving problems. I’m just saying that what enables that whole research revolution is Lean. We wouldn’t be seeing all those results without it. I would like to see it acknowledged when people are talking about LLMs solving maths. The same way I think we should acknowledge the humans who are guiding and prompting the LLMs. I don’t think that necessitates to move a goal post
- ModernMech 22d agoYes! Thank you, this has been irritating me from day one with agentic AI, I don't think it could be nearly as good as it is without all the well-designed tools humans have spent decades developing from PLs, to VCS, to the Unix philosophy, to CLIs / REPLs. I think that AI is very skilled and adept at using them, but without them it would just be flailing around in its own psychosis. The tools ground the AI in reality and allow them to make progress without going in hallucinated directions. I very much doubt AI could have solved this problem without Lean, and for AI to invent something like Lean it would have to use other tools made by humans. This is also perfect evidence of why Python isn't the end-all-be-all of programming languages just because the AI was trained on vast amounts of Python, and proof that the right language for the job is more viable than ever with the aid of LLMs. Frankly, the forecasting that programming languages are a dead field has baffled me because it seems like with LLMs, unique programming language semantics are more important than they've ever been.
- gwerbin 23d agoI don't think so. People have been trying things like this with evolutionary algorithms for a very long time already. LLMs can interleave symbolic manipulation with empirical experiments and simulations and charts and thinking/reasoning text, and an LLM will much more efficiently search the space of candidate ideas than any handcrafted mutation algorithm. Any task with a cheaply verifiable goal that requires fanning out across a massive search space is ideal for contemporary LLM technology to make progress with.
- deleted 23d ago[deleted]
- YeGoblynQueenne 23d agoThe brute-forcing is a good, old-fashioned generate-and-test approach like in Simon and Newell's Logic Theorist, which was presented in the Dartmouth convention in 1956, where AI was named by John McCarthy. Logic Theorist caused a big stir by (re) proving several of the theorems in Principia Mathematica by Russel and Whitehead. There was much excitement, then, as now, for this kind of approach and there were several systems that followed along the same lines, e.g. Automated Mathematician by Doug Lenat. Eventually it became clear that this approach is limited by what it can generate: you may have a sound and complete verifier, but if the generator, i.e. the first step in the generate-and-test pipeline, is incomplete, then the entire thing will run out of steam sooner or later. The difference with LLMs is that they are... well, large. They are the most powerful generators ever created. That means their limits are not in sight and it will probably take us a very long time to find them. Which is all to say that, yes of course, automatic verification is indispensable. But without an LLM generating an unprecedentedly large number of plausible theorems, there would be no AI mathematics, or in any case AI mathematics wouldn't have gone as far as it has.
- ForHackernews 23d agoHow long until we find out that some AI has quietly buried an exploit in Lean to cheat at proofs?
- dgellow 23d agoSimpler to exploit a soundness bug than introduce a back door I would assume
- ozgung 23d agoIf your rumor is true, what we are witnessing is a giant paradigm shift rather than individual incidents. Mathematicians were the first victims of super-intelligence. Of course it’s not an endless source. They had to burn millions of dollars to solve a single problem.
- pixl97 23d ago>They had to burn millions of dollars to solve a single problem I'd like to adjust that to "They had to burn a lot of energy (create a lot of entropy) to solve a single problem. As we go into the super-intelligence age the current paradigm of money as humans understand it may break at some point. For example to a paperclip-maximizer money at best is a short term instrumental goal, hard power of matter conversion machines is what it wants and once it has those money no longer has purpose.
- ForHackernews 23d agoI'd wager a fair chunk of my money that money breaks OpenAI before OpenAI breaks money.
- pixl97 23d agoOpenAI != AI. If you were in 1999 you'd be saying pets.com = internet.
- ForHackernews 23d agoyeah yeah yeah. I agree that AI is and will be a very useful tool, it's just not going to be worth $30T like OpenAI/Anthropic are pretending.
- bena 23d agoI think this leads to an interesting question. What happens when the money runs out? Right now, a lot of money is going to train new models. And we need to train new models because they get gated by their training data. And models are only as useful as their training data. So let's say the money stops. Do we stop training models? Do we train them slowly? Do we accept the then current models as the limit?
- HarHarVeryFunny 23d agoOpenAI said they sicced this agent army on Navier-Stokes on Sept 1st, while only a couple of days earlier OpenAI's Noam Brown happened to reply to a tweet saying that they had already tried to solve all the Millennium Prize problems and failed... So, it seems either the previous attempt didn't have the training to succeed, or was just not given the compute to do so. Once OpenAI heard that Navier-Stokes was solved, this caused them to immediately revisit the problem and throw a ton of compute at it, apparently using a more (very) recent model than what they had tried before. What we don't know is just how recent this model was, and therefore what it may have been trained on. Buckmaster/Levant had apparently been working towards this for at least a year, and made their "forced" blow-up breakthrough on August 15th. Presumably any anonymized prompts that are being trained on are part of pre-training, so older, but once OpenAI had heard that Navier-Stokes had been solved and wanted to revisit it, it seems possible they may have done a few weeks of incremental RL training on anything Navier-Stokes adjacent they could come up with, in addition to then throwing unlimited compute at it, now confident that there was something to find.
- irthomasthomas 23d agoOpenai said that a new model became available to them during this. But that could mean anything from a big new base model to a LoRA, fine-tuned on a few dozen prompts...
- auntienomen 23d agoAnd conceptually novel approaches to outstanding problems are the sort of thing that a retrain should pick up on, because they would be hard to compress into what it already knows.
- ndiddy 23d ago> What we don't know is just how recent this model was, and therefore what it may have been trained on. OpenAI's statement says that they began training their new model on August 28.
- mzs 23d agoomitting when training concluded edit: ffsm8 makes a great point below, it doesn't matter. I'm not great with dates, sorry.
- yellow_lead 23d agoBoth can be true: 1. OpenAI couldn't have solved the problem without the researchers' private data for training. 2. OpenAI models can solve math problems
- ozgung 23d agoVery likely. These mathematicians’ prompts are not like “hey chat, please solve Navier-Stokes for me”. They add real expertise and intuition from the cutting edge of their field.
- mlcrypto 23d agoAnthropic isnt getting enough scrutiny for their unprofessionalism: 1. Anthropic employee working on monumental problem but didnt receive/ask for the full backing of the company's resources 2. May or may not be mixing unreleased Claude output with Codex without zero data retention agreement 3. Victory lap on Twitter and giggling around the city before they finished the job, sparking rumors for competitors
- robocat 23d agoDr. Buckmaster sounds unsanitary. Recklessly prompting OpenAI without a care to the safety of their knowledge. And after that trying to cast aspersions at OpenAI? Hopefully we get some better facts, because OpenAI are disliked enough that a smear campaign could work against them. Edit: also the narritive is getting framed as OpenAI versus Anthropic. A highly political extremely capitalist fight is going on, and facts are victims.
- yellow_lead 22d agoHow dare employees do something without asking for the full backing of the company's resources. Incredibly unethical!
- cman1444 23d agoYou forgot possibility 3: OpenAI solved the problem without using any private training data from the two researchers. Everyone in this thread seems to have made up their mind about OpenAI's guilt though.
- merksittich 23d agoEven OpenAI's own publication [0] on Navier-Stokes from two days ago appears to contradict "basically solving anything you throw at it". The chart shows a pass rate of ~0.5 (vs. Astra's ~0.2) on "a curated set of open math problems". (Based on the timelines and events described in the publication, I presume that the "Internal Model" in the publication represents OpenAI's latest and greatest model. Evidently, this pass rate may improve in the future.) [0] https://openai.com/index/navier-stokes-solution/ https://openai.com/index/navier-stokes-solution/
- paulsutter 23d agoThe big question is whether OpenAI is training on "de-identified" sessions that are marked as "do not use for training" The answer is almost certainly yes, and this is a problem for most users.
- deleted 23d ago[deleted]
- kzz102 23d agoOn your second point: there is a more plausible explanation which David Bessis calls the "overhang". The short version is that there is a large amount of relatively low hanging fruits in mathematics, because no human has broad enough knowledge and enough time to try them all. AI is not constraint by that, and therefore can systematically pluck all those low hanging fruits. Quote: "The Overhang consists of the unrealized capital gains of past mathematical creativity, the latent value from connecting the dots in the existing corpus. It is a dividend of canonization. Mathematician X states problem A, mathematician Y crafts concept B, then mathematician Z notices that B trivially solves A and “captures” the social reward. But in the process of capturing the reward, Z usually introduces new concepts and new open problems, reinjecting latent value into the Overhang. LLMs can be trained on the entirety of the mathematical corpus. Thanks to their phenomenal memorization and pattern-matching abilities (without always being able to map out their associative logic and attribute due credits), they are in a unique position to harvest the Overhang. By contrast, professional mathematicians have typically read a few hundred articles in their career, out of millions of existing references, less than 0.1% of the total. This will lead to great discoveries, which is unambiguously exciting. But it could also lead to a sad new deal, where human slaves painfully curate the Overhang while AIs systematically beat them at the finish line." source: https://substack.com/inbox/post/183753276 https://substack.com/inbox/post/183753276
- calf 23d agoIt's like AlphaGo but playing against all living mathematicians. (Overhang being low hanging fruit is what allows this comparison, of course the general moot point is the skepticism that LLMs are also innovative etc.)
- fn-mote 23d agoWe are not seeing those incredible moves yet. The approach used in N-S was conjectured to work after B&L’s initial breakthrough. See a post by Tao. So on one hand the proof is an amazing accomplishment. On the other hand, humans have not yet discovered any superhuman moves in the proof. Just $MM grind.
- throw90094231 23d ago
- fweimer 23d agoThe leakage wouldn't be from training, but from other uses of Personal Data. As far as I understand it, users can opt out from the training aspect, but they cannot stop their conversations (“User Content”) being used “[t]o improve and develop our Services and conduct research, for example to develop new features”.
- iAMkenough 23d ago> We’ll know soon enough, but I’m inclined to believe this is true. I mean, we’ll know as soon as they decide they want to provide verifiable proof. Really dragging their feet on this front so far. I’m inclined to believe this is false.
- WD-42 23d agoIf they have solved hundreds of open problems in math, why are they publishing results for the ones other mathematicians happen to be working on at the same time? Why not the others?
- brulard 23d agoYou think other mathematicians are currently working on very little subset of relatively low-hanging fruit problems?
- sebzim4500 23d agoWell I'm sure if they find a millennium prize problem that no mathematician has worked on recently they will get right on publishing that.
- cyanydeez 23d agoThe Cult tells us the AI is almight andpowerful; unfortunately, the cult cant actually describe the indescribable.
- SrslyJosh 23d ago> The rumor I’ve heard from multiple employees at OAI and Ant is that the model has solved hundreds of open problems in maths Obviously these are unbiased and trustworthy sources.