3 ms·
But that's still false. RLHF is not instruction fine-tuning. It is alignment. GPT 3.5 was first fine-tuned (supervised, not RL) on an instruction dataset, and t
by elcomet 3y ago
But that's still false. RLHF is not instruction fine-tuning. It is alignment.
GPT 3.5 was first fine-tuned (supervised, not RL) on an instruction dataset, and then aligned to human expectations using RLHF.
- stevenhuang 3y agoYou're right, thanks for the correction