5 ms·
To use RLHF you need a dataset that includes instructions with good & bad answers - do many of those exist? I know there are a few datasets of just plain instru
by brofallon 3y ago
To use RLHF you need a dataset that includes instructions with good & bad answers - do many of those exist? I know there are a few datasets of just plain instructions-with-responses, but I'm not aware of any that have both good and bad (or ranked) responses. Is that trivial, or an important missing element here?
- sdenton4 3y agoAll of the UX interface have little up/down thumb icons... that's where the boolean feedback comes from. If people stop using that, sentiment analysis on the human responses will likely go a long way.
- valine 3y agoOpenAssistant has been collecting instruction/response data. They’ve already used that data to refine several llama models with good success. You can also bootstrap RLHF training data from the gpt4 api. Vicuna is probably the best public model created with gpt4 data available as of today.
- ttul 3y agoYou’re not supposed to do this, but GPT-4 can generate RLHF data.
- tzekid 3y agoIf I understood correctly, the OpenAssistant team wants to open-source their community built RLHF dataset. On the other hand, if you're being cheeky, I bet there's a way to datamine from websites like ShareGPT and profit off shared ChatGPT <> User interactions.