9 ms·
Buried under the drama is the fact that OpenAI is claiming that an internal model they’ve been training for less than two weeks is more than twice as capable in
by arctic-true 24d ago
Buried under the drama is the fact that OpenAI is claiming that an internal model they’ve been training for less than two weeks is more than twice as capable in mathematics as Astra, which was only made public a week ago. Even if this improvement is limited to mathematics, that is an astounding feat.
- naveen99 24d agoAstra was trained more than two weeks ago.
- sashank_1509 24d agoAstra was in use by OpenAI employees for more than 3 months internally from rumors I heard
- credit_guy 24d agoThe internal model they mention is different from Astra.
- chinathrow 24d agoPre-IPO marketing?
- jrflo 24d agoI'm so tired of this "It's just marketing!!" commentary. An AI model just proved one of the top 3 unsolved problems in mathematics, they have a Lean certificate showing it's valid. How much more evidence do you need that these models are actually highly capable?
- QuesnayJr 24d agoOf the seven Millenium problems, Navier-Stokes was the one most thought to be in reach. I'm not sure what the top 3 problems are. You can make a case for the Riemann Hypothesis and P != NP, but I'm not sure what #3 would be. Maybe the Langlands program? (That one is not as precisely stated as the other two.)
- anthonypasq 24d agothe goalposts are on Pluto at this point.
- dsdf3 24d agoI'd put good money on the fact that we will have a lot of distilled intelligence and yet the world won't look much different.
- anthonypasq 24d agoi mean that is already true
- QuesnayJr 24d agoI'm not moving the goalposts. I haven't heard anyone, ever, refer to the Navier-Stokes problem as a top 3 problem in mathematics. People were saying that they thought the solution was in reach a few years ago, before AI was at all capable of research-level mathematics (and the expectation that there was a counterexample). I am not particularly skeptical of claims about AI, compared to the average here on HN, but that doesn't mean every random piece of hype is warranted. What they did is impressive, even though we now know the only reason they threw so much compute at the problem is that they heard a rumor that someone else was already close. Navier-Stokes is not a top 3 problem in mathematics, and it was the one that was thought closest to being solved.
- Yizahi 23d agoIs there any reference to a current goalpost position? Since you are claiming they had been moved, it's interesting from where exactly.
- ameliaquining 24d agoThere were also some people talking about the Hodge conjecture, because it has some similarities to some LLM-assisted breakthroughs that were considered impressive in the distant past of [checks notes] July 2026. See, e.g., https://xenaproject.wordpress.com/2026/07/20/human-mathematicians-are-being-outcounterexampled/ https://xenaproject.wordpress.com/2026/07/20/human-mathemati...
- mrbungie 24d agoThey are highly capable, no doubt about that, but: 1) We don't really know how they arrived to this result except that they had a lead and that they threw millions of compute at the problem. The article is written in a way that makes you believe that it was just an agent loop with little human intervention, but without any evidence. 2) If the threats are to be believed, it is concerning how far they are willing to go to show how capable the model is. One would think their products and credibility would be enough to speak for themselves.
- dsdf3 24d ago"2) If the threats are to be believed, it is concerning how far they are willing to go to show how capable the model is. One would believe their products and credibility would take by themselves but here we are." Personally I anticipated nefarious behaviour as part of a broader marketing strategy to sway the view of those in the west that american frontier offerings were far better and powerful than that of China - that if you did not purchase their offerings you'd be awake every night worrying your competitor was. And this is boring - they need to admit at some point they misinvested, Anthropic less so. All this math stuff is great... but hello? The largest market cap companies are valuable irrespective of such amplified intelligence.
- scurnus 24d ago1) The article is written in a way that states clearly they threw a lot of compute at the problem. In api cost millions of dollars. 2) Millenium Problems have been the goal every AI company wanted to achieve since their diffusion, all companies have thrown a lot of resource to solve these problems, as they are very famous and scientists spent a lot of time trying to solve them. The first company to solve it will remain in history, despite all of you finding excuses about it. Regarding product and credibility normal people have a completely different view about LLMs, most don't even know difference between models and probably don't even care about Millenium problems, but care instead if chatgpt can solve their day to day problems. This is just them trying to have the throne on the AI companies space, outside it this result won't matter.
- mrbungie 24d ago
- andrepd 24d agoLmao my friend, the whole "drama" is that there are allegations of plagiarism.
- danielmarkbruce 24d agoHighly capable of writing math proofs, no doubt. It's really unclear that this entire line of work (training LLMs for proof writing) has much real value outside of writing math proofs. It is reasonably clear that, similar to Deep Blue at the time, people are extrapolating the results to general intelligence because the people who usually write proofs are insanely smart (just like world class chess players).
- eutropia 24d agoIf pre-ipo marketing pushes them to train a model capable of resolving a millennium problem in mathematics in a weekend, then, to quote XKCD: "Mission. Fucking. Acccomplished." https://xkcd.com/810/ https://xkcd.com/810/
- Aboutplants 24d agoEven if it is, Anthropic better have a few things up their sleeve
- chilmers 24d agoThe implication from their last couple of published articles[1][2] is that they think they’ve achieved “recursive self improvement”. [1] https://openai.com/index/research-acceleration-view-inside-openai/ https://openai.com/index/research-acceleration-view-inside-o... [2] https://openai.com/index/an-alien-mind/ https://openai.com/index/an-alien-mind/
- 10xDev 24d agoCompute will always be the bottleneck even if this were true.
- Miner49er 24d agoEventually recursive self-improvement includes reducing bottlenecks.
- 10xDev 24d agoEventually the bottleneck might be people themselves.
- ccozan 24d agoImprobably, the real bottleneck is energy.
- glenstein 24d agoWhich is to say, scalable and open-ended capability of ramping up physical infrastructure. I don't know that that's achievable yet. Though the era of increasingly advanced and automated robotics seems to be around the corner which could create a cycle, vicious or virtuous depending on how you feel about it.
- Fordec 24d agoIf humans can figure out to optimize to circumvent bottlenecks, I have no doubt each new bottleneck will also get routed around, just now automated.
- magicalist 24d ago> Buried under the drama is the fact that OpenAI is claiming that an internal model they’ve been training for less than two weeks is more than twice as capable in mathematics as Astra. Is this buried under the drama or are the major OpenAI twitter accounts from the people involved in the drama desperately attempting to make this the story after everything else obviously got away from them?
- ameliaquining 24d agoI don't know what anyone's been saying on Twitter and I don't care. If it's really true that there's a model out there that's that capable two weeks after the start of training, then that's objectively a much bigger deal than a priority dispute, even if the latter involves juicy allegations of espionage and skulduggery.
- 20k 24d agoIt isn't a priority dispute, the more concerning allegation is that OpenAI may be training their models on prompts that mathematicians were using to solve this problem, and then surprise surprise OpenAI were able to replicate that work in their latest model What we're really looking at is seemingly a massive plagiarism scandal, which especially brings a lot of the past results into question If OpenAI is training models on researchers' prompts, and then threatening them into staying quiet about it, who knows if anything that's been announced is genuine - or just theft? Edit: OpenAI have admitted they were training on prompts at the time they made their breakthrough https://mastodon.social/@tristanbuckmaster/117236471352470303 https://mastodon.social/@tristanbuckmaster/11723647135247030...
- ameliaquining 24d agoIf you're alleging that they don't actually have a highly capable model and the work they're attributing to it was actually plagiarized from human mathematicians, well, that would be big if true, but I'd be inclined to take the other side of that bet. With most previous splashy AI results, others have subsequently used the model to do other things around the same difficulty level. Also, it would still be necessary to explain why all these famous open problems are suddenly falling like dominoes, if it's not AI solving them. If you're saying that the question of whether they actually have a highly capable model is less important than the question of whether there's a plagiarism scandal, I continue to disagree.
- pama 24d agoNot only that, but it used 10k agents coherently over 88 hours to come up with the proof. This is a significant advance.
- danielmarkbruce 24d agoIf you can create a graph of independent work, which you can with many such problems, agents can work together nicely. Again, thank Lean and the tooling around it.
- pama 23d agoAs far as I understand the 10k agents worked on the proof. The lean formalization came later and was easier/faster than getting the proof.
- danielmarkbruce 21d agoYeah, good catch. I was under the impression they did the proof in lean from the get go, but you are right. I guess the nature of the problem lent itself to the 10k agents. Ie, there isn't something general to take here.
- topaz0 24d agoWhat makes you think they were coherent?
- pama 23d agoThey managed to solve a problem that was beyond current human ability.
- topaz0 23d agoThat was the net effect (assuming what they solved was the actual problem and not a loophole in the problem statement or a lean bug). My point is they need not all work coherently to do that -- for example, for all we know 3/4 of them went off the rails, their results were pruned, and the relevant results came from a random subset that happened to produce something useful.
- Aboutplants 24d agoI’m of zero knowledge on model training, but how is a model accessible while performing training at the same time, especially so early in its run? I’m obviously thinking a little too narrowly in terms of how it actually works
- stingrae 24d agothe model is a set of weights, you can take a snapshot and test it. Reinforcement learning itself is largely testing and tuning.
- blake__dev 24d agoYeah I'm surprised they posted a chart, you would think they would keep specifics like that hidden until they're closer to launch
- cool_dude85 24d agoThe chart is as non-specific as could be. It improved in some very vague metric by some amount at different (increasing) levels of training.
- blake__dev 24d agoThat's fair, but at least the chart has an axis. :) Since openai just released astra, I was more surprised that they would publicly show any gap to their (presumably SOTA) internal model.
- merksittich 24d agoThe x-axis label of the chart is test-time compute. Doesn't this relate to inference ("thinking level") instead of training?
- sebzim4500 24d agoIsn't the y axis just what portion of the open problems it could solve? The axis is unlabelled though, I'll give you that
- bananaflag 24d agoYeah it's Bel
- mzhaase 24d agoThe singularity happening under trump? We could have had star trek, instead we're getting the combine.
- dboreham 24d agoThat said, perhaps it will take over the world government and decree that all corrupt officials shall be imprisoned and all weapons of mass destruction shall be destroyed.
- dakolli 24d agoYou think a model with an effective memory of 200-500k words, that can be unplugged, is going to "run the world" You people gotta put down the sci-fi
- E-Reverance 24d agoThe scifi pov has a good track record as this point, you people gotta be more open minded
- fc417fc802 24d agoMany present day politicians appear to have effective memories much smaller than that coupled with equally questionable world models so ... what is your point, exactly?
- dakolli 24d ago[flagged]
- sznio 24d agoit proved navier-stokes taking over the us government is easier imo, any idiot gets to be president
- karmakurtisaani 24d agoAnd then it will be shut down, proper guard rails put in place, and the new version will accelerate the cleptocracy.
- vatsachak 24d agoBrain has loops and parallel connections. Loops and parallel connections make transformer go brrr
- curt15 24d agoThey're also counting on more casual observers to extrapolate optimistically from successes in high profile math theorems to the company's economic value.
- refulgentis 24d agoCarefully worded; it's extremely likely to be the same large frontier model that started training again on August 28th as well, as they revealed in some of the RL message board follow-up - for several reasons, most importantly, if we assume it was start of training, only a week from start of training to producing any answer would imply several orders of magnitude increase in training speed/decrease in model size.
- _fizz_buzz_ 24d agoCan someone explain if i understand this correctly: Are they saying that they started training this new model on August 28th and then started using it on September 1st? Does training a new model only take 3 days?
- gcr 24d agoit's possible to do a RLHF or RLVR pass pretty quickly. I'm almost certain a full pretraining run isn't possible within that time frame.
- tristanj 24d agoOpenAI finished another pre-train in late August, and they are now building models off that base. He's saying the specific model OpenAI used to solve this problem is currently in post-training, which started on August 28.
- lossolo 24d agoNot entirely, it's just a late stage of the overall training process. It's an early checkpoint in post training (you can use the model at different stages of training), so it will probably become even stronger with more post training.
- danielmarkbruce 24d agoWith Lean, math has become a really well suited problem for LLMs. We will likely see large gains for many years from here, just doing more and more rlvr, like continuously, non stop. No need to train from scratch. It really doesn't speak to the general intelligence of models though. It does speak to how good these things can become when a problem space has verifiable rewards, especially when you can verify one step at a time like Lean enables.
- preommr 24d agoIt's crazy how deep Microsoft's bench is (Lean was started there, vscode is another), for everything not directly related to the the ai models (hell, even github for data). So interesting how everything played out, I remember in the early days when MS came out with the partnership with OpenAI it seemed like they were playing 5d chess and were poised to win big. And it all just fizzled out. Second biggest fumble after Google.
- danielmarkbruce 24d agoWell, amazon and apple haven't done great either. One might reasonably claim they weren't as well placed as goog or msft, but it might just be "big company can't do genuinely new thing".
- sebzim4500 24d agoDon't they own a large portion of OpenAI? Things could be worse
- irthomasthomas 24d agoOr they trained a LoRA on the victims chats in order to launder their plagiarism.
- fer 24d agoThe timing makes it the most likely, not only them but potentially more. Comparatively quick, instant results. "Here's Astra! BTW our internal model is 10x better at math!" It'd be interesting to see academics having giving deeper looks at whatever OpenAI publishes from now on.
- piloto_ciego 24d agoAnd... they found this "solution" in 88 hours or so. It's all gas no brakes now boys and girls. Hold on to your hats!
- itemize123 24d agoit's buried because due to the drama the evidence is scarce
- vimbtw 23d agoMy guess based purely off of vibes from previous models is that boosting the frontier math ability of a model is not that difficult. Most models trained for general use are ingrained with certain tendencies that are usually very useful like "if you're stuck and bashing your head against the wall stop and tell the user". You generally don't want Claude Code to go off and work for weeks on something when if it had just asked for help you could've clarified or provided more information or just picked a different approach. When you're solving extremely difficult math problems though you generally do want a model to be more persistent and keep trying even when the model can't clearly see a way forward. OpenAI appears to have done this with lots of previous models. The model they trained for the IMO competition seems to have been an RL maxxed version since they noted that while it did the math it couldn't write up its results on its own and just produced CoT [0]. The capabilities are in there lying dormant, you just to need to RL max the model to ruthlessly pursue the goal at all costs which destroys general use but improves frontier math. We've also seen hints from OpenAI at least that they seem to train more persistent versions of all their models [1]. Also, Astra probably completed training at least one to two months before the public release so it's not like they only had a week to whip this version up. [0]: https://x.com/OpenAI/status/1946594933470900631 https://x.com/OpenAI/status/1946594933470900631 [1]: https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/#core-takeaways-about-this-incident:~:text=and%20a%20highly%2Dpersistent%20internal%20model%2C%20%5B6%5D%20which%20we%20will%20refer%20to%20as%20%E2%80%9CHPIM%E2%80%9D%20going%20forward https://metr.org/blog/2026-08-26-openai-hugging-face-inciden...
- soltanov 23d agoAgent systems become most credible when they produce artifacts that can be independently checked, not when they merely produce persuasive explanations.
- carlailab2025 23d ago[dead]
- HarHarVeryFunny 23d agoIf there is any truth to this timeline, then presumably it just means an additional 2 weeks of RL training on Astra. "Twice as capable in mathematics" just means they found some problems that Astra couldn't solve, or make progress on (who knows how they chose to define "capable", or "twice as" for that matter), then put those 2 weeks of training in to focus on those gaps. At this point, focused on their IPO, the best way to interpret OpenAI press releases is "what is the least this can mean, without being an actual lie". They are not shy - if there was a more impressive claim they could make, they would have made it.
- famouswaffles 23d ago>If there is any truth to this timeline, then presumably it just means an additional 2 weeks of RL training on Astra OpenAI finished a larger pre-train (rumors are it's the largest since GPT 4.5) in late August (not Astra). Presumably, this is post training on top of that since it lines up. >They are not shy - if there was a more impressive claim they could make, they would have made it. What claim would that be ?
- HarHarVeryFunny 23d ago> What claim would that be ? Huh? I'm saying there isn't one.
- famouswaffles 23d agoOkay. I was just confused.