4 ms·
My main gripe here is the lack of transparency around the total experiment and construction. I doubt that they simply pointed their model at these ten specific
by aabhay 2mo ago
My main gripe here is the lack of transparency around the total experiment and construction. I doubt that they simply pointed their model at these ten specific problems alone and gave the model one shot; therefore the $2000 number could be completely misleading, similar to P-value hacking by not disclosing the total experimental setup.
I want to know:
1. How many total problems were given to the model, and what percent were left unsolved at what cost before giving up?
2. How many attempts did you give the model at solving these problems?
3. How expensive was the harness, e.g. did the model have access to a job cluster?
- einpoklum 2mo agoAlso, have there been examples of researchers not affiliated with OpenAI (or another LLM creator), who have done something similar? Another question I have is whether or not OpenAI 'simply' hired capable combinatorics researchers to work on problems, and they have, and the use of the model is incidental / secondary to their work.
- traes 2mo ago> Also, have there been examples of researchers not affiliated with OpenAI (or another LLM creator), who have done something similar? A couple small ones that I've seen (example here [0]), but not anything of the magnitude that OpenAI and Anthropic have put out. Likely just related to token limits. > Another question I have is whether or not OpenAI 'simply' hired capable combinatorics researchers to work on problems, and they have, and the use of the model is incidental / secondary to their work. I think their output has reached a level that precludes this possibility, but I of course don't have any hard proof. [0]: https://www.reddit.com/r/math/comments/1uxj3cy/after_openais_cdc_proof_announcement_gpt56_used_a/ https://www.reddit.com/r/math/comments/1uxj3cy/after_openais...
- irthomasthomas 2mo agoWhy you think that?
- azan_ 2mo agoI guess that's because there are serious problems on which many professional mathematicians worked on years. If it was just a matter of hiring an expert, they would've been solved long time ago.
- irthomasthomas 2mo agoI guess expert+chatgpt beats chatgpt alone, so why not hire top experts to drive the search?
- energy123 2mo agoMany less important Erdos problems have been solved by amateurs prompting ChatGPT 5.{3,4,5,6} Pro using their $200 subscription.
- kittoes 2mo agohttps://blob.byteterrace.com/public/bds-theorem.html https://blob.byteterrace.com/public/bds-theorem.html I have no affiliation whatsoever with any AI company, nor any formal education outside high school, for what it's worth. Simply being curious and persistent can get you quite far in my anecdotal experience.
- brighteyes 2mo agoYes, here is another example of major work in this area: https://arxiv.org/html/2605.22763v1 https://arxiv.org/html/2605.22763v1 > Our most capable agent autonomously resolved 9 of 353 open Erdős problems at the per-problem cost of a few hundred dollars, proved 44/492 OEIS conjectures
- einpoklum 2mo agoThe actual quote: > Our full-featured agent autonomously solved 9 Erdős problems out of 353 attempted, including two questions that had been open for 56 years Note _had_ been open, not _have_ been open. Can you clarify?
- jsnell 2mo agoThe original was an actual quote? But "had" still doesn't mean what you are implying: once the model solved the problems and the solutions were verified, the problems weren't open any more, so a later description using the past tense is totally consistent.
- simianwords 2mo ago[flagged]
- traes 2mo agoIt's a very important clarification if it took $2000/problem on 20 problem attempts or on 1,000 problem attempts for each successful one. That may be the deciding factor on whether or not it's economically viable to replace a mathematician with a ChatGPT subscription.
- simianwords 2mo agoYeah fair I concede that this is somewhat crucial information. The parent seems to write it in a tone that suggests deliberate misleading “lack of transparency” etc.
- esperent 2mo ago> deliberate misleading “lack of transparency” etc It's a marketing post from a huge company. Only the naive would view it uncritically without assuming it's been written carefully to present the results in the best possible light while skirting the boundaries of outright lying.
- dist-epoch 2mo agoThe results speak for themselves. Imagine 2 years from now: "yes, GPT solved the Riemann Hypothesis, but cmon, it's just a marketing stunt to hype their stuff, it was probably Terence Tao doing the work but he's so obsessed with hyping AI that he doesn't want to take credit"
- esperent 2mo agoNobody is claiming the results are false. We're saying look critically at the claims for how it was done, that it only cost $2000, etc. it would be extremely easy to run 100 sessions that failed, each costing ~$2000, and then just publishing an article about the one that succeeded, for example. This goes double since it's an internal secret model (Astra) so nobody else can verify the results.
- dist-epoch 2mo agoI don't think you want to bring cost into this argument. Even if the cost was $1 mil for these 10 problems, that's maybe 10-20 math researchers for a year. Do you really think that if you paid that to humans, they will deliver the same results?
- mungaihaha 2mo agoGrad students on zero pay solve problems like this everyday. What exactly is your point here?
- mirzap 2mo agoEven if they can solve problems like this every day, you still have a very limited number of grad students who can solve them. With model capabilities like this, you can have the equivalent of millions of grad students who can solve problems like this.
- whattheheckheck 2mo agoGive the grad students these resources and they can do even more!!!
- gbnwl 2mo agoEveryday? Which 10 problems were solved by mathematics grad students in the past 10 days? OK I’ll grant that it’s not your obligation to be my search function (despite you making the wild assertion in the first place), so instead can you just point us to the latest grad student solved problem of this level that you know of?
- gowld 2mo ago[dead]
- r0uv3n 2mo agoGrad students do not solve problems such as the existence of non-sofic groups every day.
- irthomasthomas 2mo ago[flagged]
- azan_ 2mo ago> therefore the $2000 number could be completely misleading, similar to P-value hacking by not disclosing the total experimental setup. I don't think that comparison to p-hacking is fair. I mean not reporting price of all run is nothing like committing scientific fraud and fake results.
- whattheheckheck 2mo agoYeah I remember reading about something along the lines of Mathematics is now about the scaffolding around you find the problems/solutions not just the problems and solutions. For teaching purposes. This was before this ai craze
- c7b 2mo agoI believe we're seeing a new kind of mathematics that will require completely new formats for publication, a bit similar to those used in experimental sciences. AI-powered mathematics should be fully reproducible, so it's the authors' responsibility to disclose the exact model type, inference settings/seeds and the full prompt history leading to the result. Of course that would ideally require open weights models. It's not just about requiring to disclose AI use. AI-powered mathematics is a completely valid discipline that doesn't need to be shy, but it should develop its own publication culture.
- lkirk 2mo agoI think this is a bit optimistic compared to my view (wrt portability). There's a large stack of software that is involved in training and probably less so in inference. I'm not saying it's impossible but there are definitely different levels of reproducibility and the academic incentive structure doesn't really prioritize reproducibility in my experience. I'm sure it varies quite a bit, I'd be curious to know how those in this problem space are thinking about reproducibility and at what level.
- c7b 2mo agoI know it sounds unrealistic and not aligned with academic incentive structures. But those are the exact structures that gave us a lot of headaches in the experimental sciences. I think it would be a good north star to aim for something that resembles how those are trying to address the reproducibility crisis. Better than to embrace the most black-box version of math that AI systems can produce (million-line proofs without context). Even if a reproducibility crisis is seemingly impossible (although agents so far have also been pretty good at finding compiler bugs).
- jsenn 2mo agoI can see this being important if you only care about the results as evaluations of AI progress, but if what you care about is the math itself why should you care about the prompt or anything other than the proof?
- 2mo ago
- wrsh07 2mo agoIt seems like they threw it a decently large battery of open math problems and probably limited it to something like $200-500 per problem: https://x.com/polynoamial/status/2083478171975082334 https://x.com/polynoamial/status/2083478171975082334 As a complete guess, it seems like they tested hundreds to thousands of problems with a relatively low per-problem budget -- The linked tweet from Noam Brown at OpenAI reads: > And yes we did try other major problems without success. Sadly no Millennium Prize problems (yet). > But also, we didn’t spend a lot on each problem. It’s possible to push test-time compute much further.
- moscoe 2mo agoI guess people will always find something to gripe about.