3 ms·
As an ex-mathematician I was really interested to see how well o3 etc handles difficult, unseen math questions, so I tried giving it some hard-ish questions fro
by sweezyjeezy 2y ago
As an ex-mathematician I was really interested to see how well o3 etc handles difficult, unseen math questions, so I tried giving it some hard-ish questions from mathoverflow [1] (mainly non-trivial questions on graduate+ level topics). It definitely isn't great and may even be more harmful than useful currently. The main issue it will never say "I'm not sure how to do this", it will almost always give a complete answer from start to finish, with 'bugs' along the way that can be very subtle.
But I found it genuinely shocking some of the steps it manages to take successfully, and it definitely doesn't feel like we're a million years from something could replace big parts of researchers' work. I honestly found some things it could do extremely unsettling as a thought-worker.
[1] https://mathoverflow.net/ https://mathoverflow.net/
- 2-3-7-43-1807 2y agonot sure _how_ relevant it is at the bottom line but you are aware that o3 was trained very likely on those same problems form mathoverflow? cause you know - the underlying llm was trained on pretty much everything available oline - especially high quality sources like ... mathoverflow.
- sweezyjeezy 2y agoI was posting questions from the last week, so not in o3 training data.