3 ms·
Do you not log the training data? Seems like you should be able to just check what was in the training data. To not keep track is just sloppy work and certainly
by Pulcinella 24d ago
Do you not log the training data? Seems like you should be able to just check what was in the training data. To not keep track is just sloppy work and certainly unprofessional science.
- tristanj 24d agoThe burden of proof is on Buckmaster and Alpöge to reveal if they had the "Improve the model for everyone" setting enabled or disabled. OpenAI shouldn't be expected to reveal private user configuration data. You're asking them to perform a user privacy violation.
- whimsicalism 24d agoIf the reason that OpenAI is unable to state whether they trained on this data is because they (as policy) do not reveal whether a given member has turned on/off the "Improve the model for everyone" setting, they can at least say so. FWIW, publicly facing OAI docs are very unclear about whether this setting even applies to Codex conversations.
- tristanj 24d agoAn OpenAI employee did say so: https://x.com/tszzl/status/2097393423808377173 https://x.com/tszzl/status/2097393423808377173 it is exceptionally unlikely that anything they ever did made it into any part of training, and the chances are zero if they have opted out (likely). it would be a terrible precedent to break the the PII-scrubbing boundary to go and round it down to 0, and we won’t do it
- amazingman 24d agoAnd we all know how good OpenAI is at containing models during training...
- unsupp0rted 24d agoThat's an entirely different question
- ted_dunning 24d agoNot really. We have lots of examples now of their model doing what they say is impossible. Now we have another example of something that they say is impossible or very unlikely. Do we take their word for it this time? Really?
- illiac786 24d agoDoes Anthropic, Grok, etc. log their training data? I had the impression it was rather a mess.
- tristanj 23d agoYes. Pretty much all the models that don't suck are trained on user data, either directly or via derived synthetic data. Many upstart Chinese labs got around the user data issue by just buying copious amounts of Claude and ChatGPT session logs from model routers.
- illiac786 21d agoBut I was asking about _logging_ training data, I don’t understand how your answer relates to my question. Maybe I’m being obtuse.
- RunSet 24d agoThat training data does not preserve provenance seems a "smoking gun" in terms of intent to plagiarize.