3 ms·
All that I have seen OpenAI employees "admit" is that if you press the Thumbs up button on a response, this can be used as a signal for training. That's it. Th
by user43928 29d ago
All that I have seen OpenAI employees "admit" is that if you press the Thumbs up button on a response, this can be used as a signal for training.
That's it. The rest appears to be wild speculation.
- jsw97 29d agoYeah this is a land mine. Even if you opt out of them training on your conversations, giving “feedback on a new version” or answering “how are we doing” or whatever can slurp up all relevant context, which if you think about it can probably be construed to include memories, into the belly of their flying saucer. At least the last time I checked their terms. Never ever touch those requests. If you get a side by side comparison just resend the prompt.
- lackoftactics 29d agodo you even needs thumbs up? I've been long suspecting that code models get better because they use our data and our results from feedback, status codes, green tests for reinforcement learning
- user43928 29d agoI was thinking that the training is more curated, so that the methods are learned from experts, and that measurably successful behavior is reinforced. Throwing in random chats with some sentiment analysis doesn't seem like the most promising method to me, but I can only speculate.
- lackoftactics 29d agohmm, maybe not sentiments, running commands can produce binary results to reinforce, but that's also speculation