4 ms·
>the ONLY thing these LLMs know how to do is predict the probability that their next word This is super incorrect. The base model is trained to predict the dis
by smus 2y ago
>the ONLY thing these LLMs know how to do is predict the probability that their next word
This is super incorrect. The base model is trained to predict the distribution of next words (which obviously necessitates a ton of understanding about the language)
Then there's the RLHF step, which teaches the model about what humans want to see
But o1 (which is one of these LLMs) is trained entirely differently to do reinforcement learning on problem solving (we think), so it's a pretty different paradigm. I could see o1 planning very well