3 ms·
Is it a good insight though? The training data for many of these premier models is already available, and that's even better than some input-output pairs.
by substation13 4y ago
Is it a good insight though? The training data for many of these premier models is already available, and that's even better than some input-output pairs.
- zarzavat 4y agoHe’s talking about RLHF - reinforcement learning with human feedback (the process that trained ChatGPT), the training data for which is not publicly available. And the point is that you don’t need to RLHF as long as you have access to another model that has been trained with RLHF that you can blackbox.
- substation13 4y agoI see, I didn't realize that training data for tuning was not available. Thanks!