3 ms·
> The most promising idea is to use reasoning models to generate data, and then train our non-reasoning models with the reasoning-embedded data. Why is it prom
by ClumsyPilot 2y ago
> The most promising idea is to use reasoning models to generate data, and then train our non-reasoning models with the reasoning-embedded data.
Why is it promising, aren’t you potentially amplifying AI biases and errors?
- pas 2y agoit seems to work and seems very scalable, "reasoning" helps to counter biases (answers become longer, ie. the system uses more tokens which means more time to answer a question -- likely longer answers allow better differentiation of answers from each other in the "answer space") https://newsletter.languagemodels.co/i/155812052/large-scale-reasoning-oriented-reinforcement-learning-r-zero https://newsletter.languagemodels.co/i/155812052/large-scale... also from the posted article """ The R1-Zero training process is capable of creating its own internal domain specific language (“DSL”) in token space via RL optimization. This makes intuitive sense, as language itself is effectively a reasoning DSL. """