3 ms·
The issue is that it is a stretch to call it reinforcement learning when all we do currently (in the context of LLMs) is to multiply the reward with the learnin
by randomNumber7 1y ago
The issue is that it is a stretch to call it reinforcement learning when all we do currently (in the context of LLMs) is to multiply the reward with the learning rate.
It sounds cool as marketing. It helps improve LLMs a bit. And it will never yield s.th. like an AGI or anything that is "reasoning". Unless you also redefine the word reasoning of course.