4 ms·
Another interesting question is why the frontier labs are piling on pure maths, which has little direct economic value compared to something like law or improvi
by curt15 2mo ago
Another interesting question is why the frontier labs are piling on pure maths, which has little direct economic value compared to something like law or improving the efficiency of their own models? How much OpenAI and Anthropic are paying to serve these models for ordinary users is the elephant in the room. A cynical take is that the frontier labs are trying their best to pump up their pre-IPO valuation through flashy headlines.
- gpm 2mo agoBecause it's a tool in search of a use case (or many use cases) and mathematics is the most natural use case for it. Mathematics is by definition the art of putting words on a page in a rigorously defined "correct manner" (i.e. in the form of a valid logical argument, a proof) and all LLMs do is put words on pages and evaluating if they're good words is by far easiest when there is a strict definition of right and wrong.
- danielmarkbruce 2mo agoIt's one of the few areas where you can verify results. That fits nicely into training models. They aren't just making judgement calls on what would be nice, it's "what can we do?".
- beering 2mo agoIf everyone publicly said that the models can only do things that humans have already done, but you know they can do more, wouldn’t you want to show them otherwise? Math ability also helps with other things like making models more efficient.
- anon373839 2mo ago> Another interesting question is why the frontier labs are piling on pure maths The reason is that the original scaling axes (parameters, training tokens, test-time compute) have saturated already, but RLVR (reinforcement learning from verifiable rewards) is still scaling well. And math has this nice property where you can synthetically generate arbitrary volumes of rewards to train the model, because math is self-contained and completely objective. Open-ended reasoning and analysis don't have that convenient property, and that is why progress is much slower outside of math and coding.
- robotpepi 2mo agodo we have any concrete idea of how well the models are scaling now? I agree with these 10 results being impressive, and it is easy to think "wow, and last year the models were barely able to solve IMOs problems". But for me it is perfectly possible that a non sofic group could be found by 100 good IMO students working on all the different strategies that have been proposed (OpenAIs solution was based on "expander graphs", which were introduced to solve the problem some years ago), so it could be that current AI is simply many (say 1000) old models working in parallel. This is linear scaling, not exponential. It could be I'm completely wrong also, the problem is that we have little information.
- qingcharles 2mo agoWhat's the best way to apply it to legal problems? Finding bugs in statutes? (there are often statutes with wording errors, missing negatives, things like that which don't get picked up for ages)