4 ms·
Much of the ‘secret sauce’ at the AI labs is not from the corpus on which they are trained but from the reinforcement training done afterwards. This step can in
by dan-robertson 2mo ago
Much of the ‘secret sauce’ at the AI labs is not from the corpus on which they are trained but from the reinforcement training done afterwards. This step can influence the ‘personalities’ of the models and it is what makes them better at being ‘agents’, able to string together various individual steps to achieve your goal.
You could imagine that telling the LLM you have a holdout test makes it ‘feel’ more like an environment in which it was being RLed and therefore makes it better at seeking the reward by doing a good job.