4 ms·
Actually I didn't. Correct me if I am wrong, but my understanding is that RL is still an LLM tuning approach, i.e. an optimization of its parameter set, no matt
by physix 1y ago
Actually I didn't. Correct me if I am wrong, but my understanding is that RL is still an LLM tuning approach, i.e. an optimization of its parameter set, no matter if it's done at scale or via HF.
- sva_ 1y agoRL is a lot more general than that, it is basically a way in which an agent learns to make optimal decisions by learning from experience to maximize rewards. So you can do all kinds of stuff other than finetuning LLMs with it, like controlling a robotic arm, playing/mastering videogames, etc. For example, AlphaGo was also RL.