3 ms·
Training data can be anything. They scrape the entire internet, plenty of which is inaccurate or poorly written. That doesn't prevent it from being useful train
by RevEng 2y ago
Training data can be anything. They scrape the entire internet, plenty of which is inaccurate or poorly written. That doesn't prevent it from being useful training data because the point of an LLM trained for text generation is to predict what someone would write. Every question you ask and your responses to their responses is valuable data that isn't already available on the public internet. Even if some of it is unreliable or even intentionally adversarial, on average the responses will be useful for training. This is training data that they have exclusive access to it.
- fsndz 2y agoTypically, when people aren't satisfied with an answer, they continue asking questions or prompting further, which makes the whole exchange valuable even without the use of a thumbs-up or thumbs-down button. That's the secret weapon of OpenAI, and a part of their 'moat'
- AmericanChopper 2y agoThis is what I’m skeptical of, because my experience with LLMs doesn’t align with this description at all. Even if a response is good, I will typically give it follow up prompts to get more details, or answer other questions that the (potentially high quality) response raised. If the response is bad I might try some follow ups to see if it improves. In either case, I’m submitting a prompt, I might accept the answer immediately, give up immediately, accept the answer after further prompting, or give up after further prompting. With my own use there is no correlation between the number of prompts I submit, and the quality of the responses given. If this is the metric OpenAI is using to perform crowdsourced RLHF, then the reinforcement is going to be garbage.
- fsndz 2y agoThe chain of thought that is apparent in the conversation is what's really interesting and what OpenAI exploits. OpenAI does not have to evaluate the quality because the conversation itself is already valuable. Whether the conversation was good or bad can easily be inferred from the back-and-forth. That's the key: they just have to take all those conversations and fine-tune their models with them.
- ben_w 2y agoWhat about the nature of your follow up responses? Do you say "No, not x, y"? Or perhaps "Now Baz the Foos"?