3 ms·
OpenAI's release explicitly says No. But then also caveats that with "we cannot rule out that de-identified data derived from their usage of our products" impac
by naniel 24d ago
OpenAI's release explicitly says No. But then also caveats that with "we cannot rule out that de-identified data derived from their usage of our products" impacted things.
What's most striking to me, and what may or may not be true, is the "we cannot rule out" bit.
"We (the researchers and the agents) did not see any of their work through any means until they released it publicly — in particular, no specific user data was accessed in order to solve this problem. While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models . However, our proofs differ significantly and even the precise results proved are different in the Euler case (forced vs unforced)." https://openai.com/index/navier-stokes-solution/ https://openai.com/index/navier-stokes-solution/
NOTE: there are a couple duped threads around this. i replied on a different one first before seeing this one
- matt3210 24d agoLook at openAIs history, they clearly have no idea what the agents are doing.
- golly_ned 24d agoThat was also striking to me too, the 'we cannot rule out'. I also do not believe it. They can, trivially, rule out their model being trained on if More to the point -- of _course_ they know what data their model is trained on and the lineage of that data, even if it ended up being anonymized and they cannot identify which precise user.
- burkaman 24d agoI don't understand how they can say it's unlikely. It's objectively true that they train on de-identified user data (https://openai.com/policies/how-your-data-is-used-to-improve-model-performance/ https://openai.com/policies/how-your-data-is-used-to-improve...), and objectively true that they encourage users to submit such data (the setting is on by default). Since it's de-identified I can believe that they can't give a straight yes or no answer here, but it seems more likely than not that at least some amount of his usage became training data. It takes an unusual level of awareness and effort for a user to ensure that all usage is opted out.