3 ms·
Nobody except OpenAI knows whether or not OpenAI trained on their data. So the burden remains on OpenAI here.
by nulld3v 25d ago
Nobody except OpenAI knows whether or not OpenAI trained on their data. So the burden remains on OpenAI here.
- fc417fc802 25d agoThat is an absurd and entirely untenable position that breaks with approximately all western conventions. Only the CIA knows whether or not they're actively covering up reptilian space aliens exerting control over the US government. Therefore the burden of proof remains on the CIA to prove that they are not actively participating in such a scheme.
- nulld3v 25d agoI don't understand, OpenAI can just say: "yes/no we did/did not train on your data". It's not a hard question to answer, and it is a question that OpenAI should be able to answer for all data we feed into ChatGPT.
- fc417fc802 25d ago> It's not a hard question to answer I didn't realize you had insider knowledge about their systems. Do please explain for the class. As I understand it they will only have trained on his data if he consented to it. Do you have evidence that they do otherwise?
- Timon3 25d agoThis whole discussion is about evidence. That's not proof and it is not certain, but it is evidence pointing into the direction that OpenAI might be doing something that they're strongly incentivized to do. What kind of "evidence" do you see as necessary?
- hellohello2 25d agoWhen someone authors a paper, is it on others to proove the author did not use their work as inspiration? No, it is on the author to give credit where it is due. You guys are acting as if it its legal issue, when it is not.
- machomaster 25d agoYou can never prove the negative.
- tristanj 25d agoIncorrect, Buckmaster and Alpöge can comment if they had the ChatGPT "Improve the model for everyone" setting enabled or disabled. If it was enabled, then their work was included in the training dataset.
- opello 25d agoIn order for this to be the strong evidence everyone also has to believe that the setting is absolutely true. That some logging from some piece of the system could not also leak the prompt information in such a way that it could have been included as training data. Perhaps the design of how data is collected for the training dataset is so rigorous as to make this a practical impossibility. But, it's asking a lot without sufficient detail to completely exclude from possibility that one setting is all that could possibly have been absolutely load bearing in deciding if the other researcher's active efforts meaningfully contaminated the internal model. At least, as an ignorant outsider, that's how it seems to me.
- pesacharia 24d agoAs I understand it, that setting does not prevent them training on user data, just which derivatives are used (i.e. just PII scrubbed vs certain types of synthetic summarization)