5 ms·
It's really neat that the prompt was released! I'm curious how many unsolved problems are tried against frontier models when they come out. Are we trying every
by bgirard 3mo ago
It's really neat that the prompt was released!
I'm curious how many unsolved problems are tried against frontier models when they come out. Are we trying every problems against every release? What is the solve success rate? Is there a sub-community within Mathematics that is coordinating this effort? How much untapped opportunity is there here?
- emil-lp 3mo agoThe prompt was released, but not the cost of the result.
- riknos314 3mo agoAssuming all 64 subagents were running for a full hour (the tweet states just under an hour): Throughput Output tokens Output cost ---------------------------- ------------- ----------- 40 tok/s (5.5 low) ~9.2M ~$275 55 tok/s (5.5 base) ~12.7M ~$380 70 tok/s (5.5 high) ~16.1M ~$485 750 tok/s (Sol Fast, $75/M) ~172.8M ~$13,000 Claude estimates that tool use / input tokens might add 10-15% on top of that depending on exactly how the model went about the task. Edit: better tok/s estimate buckets based on GPT 5.5 actual speeds since I couldn't find real benchmarks on 5.6 published anywhere. Also account for Sol Fast pricing.
- conradkay 3mo agoSol fast isn't the Cerebras 750 tok/s version, it's just 1.5x speed at 2.5x price I assume they didn't use the Cerebras version for this since it's probably very supply-constrained right now
- throw1234567891 3mo agoBut Sol is running on Cerebras. That’s the whole point of this. That’s how they get 750 tokens per second. There is no other way.
- Anuiran 3mo agoRegular Sol does not run on Cerebra’s. I don’t think anyone public has access to that. https://x.com/thsottiaux/status/2075596669958472146?s=46&t=Z63pc8n9_5mIhTJJzzSOTA https://x.com/thsottiaux/status/2075596669958472146?s=46&t=Z...
- judge2020 3mo agoYeah they posted an update below it > apparently that is just Sol being Sol on fast mode, not 750 tps. :x O.o > This is real and it is not 750 TPS. Anyone with 5.6 Sol Ultra on fast mode can reproduce this! It’s all GUI interactions with CUA. No MCP or bpy needed. https://x.com/kimmonismus/status/2075493505011482922?s=20 https://x.com/kimmonismus/status/2075493505011482922?s=20
- throw1234567891 3mo agoWho’s they, who’s chubby, why would I care.
- deleted 3mo ago[deleted]
- therobots927 3mo agoAnd not how many times it was prompted before it returned a working solution. Or how many prior variants of this prompt were tried. Or if proof checking software was used to hone in on the final winning prompt / LLM output.
- not-a-llm 3mo agopretty sure already millions of dollars (in inference costs) were already thrown at the Riehmann hypothesis as the models get stronger, larger amounts will be thrown at it imagine paying "just $1 bil" to go down in history as the company who's model solved the hardest/most famous open problem in mathematics. imagine the worldwide press headlines. as they say, the Riehmann Hypothesis is the hardest way to earn a million dollar
- Frost1x 3mo agoI’m all for it since it’s value directly returned to humanity.
- deleted 3mo ago[deleted]
- hnisfulomrons 3mo ago[dead]
- CSMastermind 3mo agoI mean if there's something I'd bet against being solved by LLMs in my lifetime it's that one. We truly do not have line of sight into what a proof would even look like.
- blovescoffee 3mo agoWhy would you bet against it being solved by LLMs? Isn't this very post proof that LLMs in an agentic harness are capable of doing real math? If you just keep cranking away at the tokens I don't see a principled argument against that leading to more solutions to unsolved math, even the hardest problems.
- CSMastermind 3mo agoThere are different classes of mathematical problems. The one in the post definitely shows the advantages that LLMs have compared to humans for some problems but it's in an entirely different class than the Riemann Hypothesis. Riemann is one of the most studied math problems of all times and all of humanity has basically collectively failed to make progress. The idea that there's some technique that just hasn't been tried yet (like in the post) is very very unlikely. The general consensus is that we'll need an entirely new branch of mathematics to solve Riemann - our current tools aren't just inadequate; they're of the wrong class entirely. I suspect inventing new branches of math will remain beyond LLMs for the remainder of my life.
- edflsafoiewq 3mo agoI find it kind of interesting the whole output wasn't released. A common criticism of mathematical writing is results are "pulled out of a hat"; you only write up a polished, final proof, but hide everything that went into developing it. It's kind of ironic the practice is even carried on when an LLM writes the proof.
- singularity2001 3mo agoVery good question I can only answer for one subset tracked by Terence Tao https://github.com/teorth/erdosproblems https://github.com/teorth/erdosproblems