3 ms·
>we did not read any private chats The question I am interested in is not "did we read private chats", but "was this new model trained using any of Tristan and
by Imnimo 18d ago
>we did not read any private chats
The question I am interested in is not "did we read private chats", but "was this new model trained using any of Tristan and Levent's chats, regardless of whether they were marked private". Can you comment on that?
- moralestapia 18d ago>Can you comment on that? No answer is also an answer. He's a human, like everybody else. Mostly a bunch of hungry animals looking to put bread in our mouths. It's rarely ever something a bit more sophisticated than that.
- lukewarm707 17d agohis choice to defend the indefensible.
- tedsanders 18d agoIf they opted out of training, then we definitely did not train on them. If they did not opt out, then I don't personally know if training signals came from their chats, and I don't think we'd be able to tell without their cooperation in identifying them. And even if signals were trained on in some manner, I highly doubt it made a difference to a problem as challenging as the NS proof. Reasons for my doubt: - I know most of our training recipes - Our model's proof is very different from theirs - The proof took a tremendous amount of tokens to derive (it wasn't a recall/lookup type question) - This unreleased model has beastly performance on many unsolved math problems, not just the Euler solution I acknowledge that this requires trust, and if you think we'd lie shamelessly about this stuff, then nothing we say can really help our case here. Reminds me a bit of the Frontier Math fiasco, where people accused us of training on the eval set (we didn't), but it's hard to convince someone if they think you're lying. If you're convinced we lie and cheat, then nothing I say may help. But if you're not sure, then hopefully providing my perspective is helpful.
- lossolo 18d ago> If they opted out of training, then we definitely did not train on them. Can't you guys just check their account settings so the public knows what was set? EDIT: Why was this downvoted? I'm genuinely asking because I have no idea. Opting out is just a normal setting in the profile, It's not like I'm asking for their private conversations or PII. If I were the person claiming that they trained on my conversations, I'd make sure to disclose that I had opted out and hadn't given them permission to do so. And if I were the accused party, I'd disclose whether that setting was turned on or off to provide evidence against the accusation.
- derangedHorse 17d agoI don’t think your question is unfair*. They can check and so can Buckmaster. If he didn’t opt out, there’s a good chance his data was used for training. I believe this to be the case myself. What I’m more skeptical about is the purported impact of this data on the model’s behavior.
- lossolo 17d agoYeah, I'm just curious about the setting. It's just weird to me that this wasn't disclosed by either party while the accusations were being made, that's all. Even if it was used in the training data, I don't believe it had that much of an impact myself, since the solutions are quite different.
- mucha 18d agoThat's not what your Chief Research Officer, Mark Chen, says on X: "Do we use user feedback and de-identified data to improve ChatGPT and Codex in a holistic way? Yes. And so does every LLM company." https://x.com/markchen90/status/2097400166554993041 https://x.com/markchen90/status/2097400166554993041
- vemacs 18d agoDoes OpenAI think de-identified data is no longer user data? Wild take for OpenAI and certainly not industry standard.