4 ms·
The biggest problem is the word "predictor". Once you get into post training with RLHF and RLVR, it simply isn't doing that. It is not predicting anything. It's
by danielmarkbruce 1mo ago
The biggest problem is the word "predictor". Once you get into post training with RLHF and RLVR, it simply isn't doing that. It is not predicting anything. It's producing tokens, but it isn't predicting them. The chess analogy in the post is a good one - it's closer to searching for a set of moves that give a result than predict. It's search for a set of ideas, represented as locations in very high dimensional space, that when put together in the right order lead to a result.
- hippietrail 27d ago"Guessing" is more accurate than "predicting". It takes educated guesses.
- danielmarkbruce 27d agoIt's closer to strategizing once you've run a model through RL. It's optimized to take steps which will lead it to a good outcome down the road.