3 ms·
But it still remains far away from mathematics research. Solving any of the problems would not result in a new research paper.
by christianstump 4mo ago
But it still remains far away from mathematics research. Solving any of the problems would not result in a new research paper.
- jona-f 4mo agoWas this event sponsored by Surge AI? Why didn't you run the prompts yourself?
- christianstump 4mo agoNo, they only provided large-scale model runs for us (this is explained in the ackonowledgements). These runs would have been too expensive to perform myself, so I am happy they offered to provide them.
- jona-f 4mo agoThanks for answering this random internet guy's question. It's a bit sad that a german math prof doesn't have sufficient funds to run a few prompts. I would have paid for them for this amount of advertising. I don't like that you gave them to a silicon valley company. On that note, the tests are very US-centric. Only one chinese model and you unfairly nerfed it by limiting it's context window, when the compressed context is deepseek v4's main innovation and even with full context it is much cheaper to run than all the others.
- christianstump 4mo agoPlease indicate which other models you would like to see included. (And I agree that the context window limitations were not reasonable to have.) Finally: running this few prompts would have been $10-20k if I would have run them myself via the API. (And the company didn't asked to contribute, but I asked whether they would be willing to do so, just saying.)
- jona-f 4mo agoKimi K2.6 and mimo 2.5 pro are ahead of deepseek v4 in other benchmarks. Anyhow, great work, the benchmark seems to show great separation, so should be very useful to improve the math capabilities of the next generation of ai. I'm more interested in the prompt engineering/orchestration and technical details (what I can do without millions), but I get that you are mathematicians, so your focus is obviously on the math. Sorry for the nagging.
- jll29 4mo agoCan anyone comment on the "distance to publication-worthiness" of the typical question from this set?
- christianstump 4mo agoInfinitly far away. These questions are about "have you understood and can you apply existing research" not about "create new research". For humans, these two correlate quite strongly. So we ask PhD students to work on the former to prepare and to become better at the latter. For LLMs, it remains unclear if there is any correlation.