3 ms·
I don’t have any insider insight on this but the GPT3 paper discusses some of their data sets and curation techniques (https://arxiv.org/pdf/2005.14165.pdf http
by jeeeb 3y ago
I don’t have any insider insight on this but the GPT3 paper discusses some of their data sets and curation techniques (https://arxiv.org/pdf/2005.14165.pdf https://arxiv.org/pdf/2005.14165.pdf).
The recent DinoV2 paper is also interesting reading (https://arxiv.org/pdf/2304.07193.pdf https://arxiv.org/pdf/2304.07193.pdf), as they particularly focus on techniques for improving the training set.
OpenAI also have been open about making heavy use of RL (via PPO) to fine tune the models.
For RL it seems they’ve basically developed a second model that can be used to score the quality of responses based on the encoded preferences of human evaluators. I.e you build a ranking of different responses based on desired characteristics (e.g. polite, helpful etc) and use those to train a second model which models the RL reward function. This can then be used to fine tune the main model.