4 ms·
>Generate chains-of-thought (CoT) for a problem domain. >Label the intermediary CoT steps using a combination of human experts (“supervised fine tuning” or SFT)
by dakshgupta 2y ago
>Generate chains-of-thought (CoT) for a problem domain.
>Label the intermediary CoT steps using a combination of human experts (“supervised fine tuning” or SFT) and automated machines (“reinforcement learning” or RL).
>Train base model using (2).
This is remarkably intuitive and elegant. Seems analogous to the idea that humans can come up with new knowledge by synthesizing from their current knowledge. Theoretical sciences or creative arts for example.