3 ms·
I don't know the scope of OpenAIs data cleaning or RLHF efforts. But I'd guess this is a big part of what's made them the leader. It's one thing to train an LLM
by version_five 3y ago
I don't know the scope of OpenAIs data cleaning or RLHF efforts. But I'd guess this is a big part of what's made them the leader. It's one thing to train an LLM on a big text corpus, as researchers do. It's another to actually wade into it and do the curation to turn it into a real product. I know there are public instruction datasets out there, but my guess is OpenAI has put a lot of money into getting cleaner data and more supervised / feedback data. Doing that extra work is the kind if thing that wins when building a product.
- famouswaffles 3y agoIt's part of it but it's not necessary to get humans to do it. If you're aware of the Claude Models, they use reinforcement learning from AI feedback and they're great. The best Claude model is only behind GPT-4