3 ms·
I haven't read TFA as I'm at work, but I would be very interested to know what the system was doing in those three days. Were there failed branches it explored?
by ZenMikey 2y ago
I haven't read TFA as I'm at work, but I would be very interested to know what the system was doing in those three days. Were there failed branches it explored? Was it just fumbling its way around until it guessed correctly? What did the feedback loop look like?
- qsort 2y agoI can't find a link to an actual paper, that just seems to be a blog post. But from what I gather the problems were manually translated to Lean 4, and then the program is doing some kind of tree search. I'm assuming they are leveraging the proof checker to provide feedback to the model.
- tsoj 2y agoThis is NOT the paper, but probably a very similar solution: https://arxiv.org/abs/2009.03393 https://arxiv.org/abs/2009.03393
- visarga 2y ago> just fumbling its way around until it guessed correctly As opposed to 0.999999% of the human population who can't do it even if their life depends on it?
- dsign 2y agoI was going to come here to say that. I remember being a teenager and giving up in frustration at IMO problems. And I was competing at IPhO.
- kevinventullo 2y agoYeah, as a former research mathematician, I think “fumbling around blindly” is not an entirely unfair description of the research process. I believe even Wiles in a documentary described his search for the proof of Fermat’s last theorem as groping around in a pitch black room, but once the proof was discovered it was like someone turned the lights on.
- logicchains 2y agoI guess you mean 99.9999%?
- lacker 2y agoThe training loop was also applied during the contest, reinforcing proofs of self-generated variations of the contest problems until a full solution could be found. So they had three days to keep training the model, on synthetic variations of each IMO problem.
- thomasahle 2y agoThey just write "it's like alpha zero". So presumably they used a version of MCTS where each terminal node is scored by LEAN as either correct or incorrect. Then they can train a network to evaluate intermediate positions (score network) and one to suggest things to try next (policy network).
- utopcell 2y agoI'm at work and reading this article is the first thing I did this morning. What's your point ?