3 ms·
Hard to measure these days. The training sets are so large they might contain leaks of test sets. Take these numbers with a grain of salt.
by vletal 4y ago
Hard to measure these days. The training sets are so large they might contain leaks of test sets. Take these numbers with a grain of salt.
- code51 4y agoOr... it could be that Chinchilla study has deficiencies in measuring capabilities of models maybe? Either that or your explanation. Frankly I don't think 13B is better than GPT-3 (text-davinci-001 which I think is not RLHF - but maybe better than base)
- simonw 4y agotext-davinci-001 is currently classed as "GPT 3.5" by OpenAI, and it did indeed have RLHF in the form of instruction tuning: https://openai.com/research/instruction-following https://openai.com/research/instruction-following MY MISTAKE: 002 and 003 are 3.5, but 001 looks to have pre-dated the InstructGPT work.