2 ms·
yes, huge for pure math and activities that look like it.
by jaykru 11d ago
yes, huge for pure math and activities that look like it.
- danielmarkbruce 11d agoDoesn't really even need to look like it. If you can verify rewards, RLVR will optimize really really well. If you can't... it's a struggle. There are probably fewer fields where you can verify rewards than one might hope.
- skydhash 10d ago> There are probably fewer fields where you can verify rewards than one might hope. 2 tasks I've done today that I believe robots are nowhere near being able to do: Cleaning my wardrobe and draining bad fuel out of my generator. As in generic use cases.
- danielmarkbruce 10d agoHard to verify that your wardrobe is clean. Also hard to verify that the bad fuel is out without physical sensors. Many, many tasks are quite difficult to verify beyond "you know it when you see it". That doesn't work so well for training a model.