3 ms·
> This is the first open research to validate that reasoning capabilities of LLMs can be incentivized purely through RL, without the need for SFT. This is a no
by justinl33 2y ago
> This is the first open research to validate that reasoning capabilities of LLMs can be incentivized purely through RL, without the need for SFT.
This is a noteworthy achievement.
- throwaway314155 2y agoExcuse my ignorance. What does SFT refer to here?
- josephcsible 2y agoSupervised fine-tuning