5 ms·
From OpenAI: > While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models This is the crux
by highfrequency 18d ago
From OpenAI:
> While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models
This is the crux of it. If Tristan's work and insights were not used to train OpenAI models, then this just looks like a case of hyper-competitive academic sniping that has been going on for decades (check out Watson and Crick!) accelerated by AI as a tool.
The fact that this is ambiguous even to OpenAI leaves one huge question: did Tristan opt out of model training for his ChatGPT and Codex sessions? If the answer is no, then this seems fair game. If the answer is yes, then OpenAI's ambiguity is strongly suggestive that opting out of model improvement does not mean what they imply it means.
- mucha 18d agoDoes opting out matter? "Do we use user feedback and de-identified data to improve ChatGPT and Codex in a holistic way? Yes. And so does every LLM company." - Mark Chen, Chief Research Officer, OpenAI. https://x.com/markchen90/status/2097400166554993041 https://x.com/markchen90/status/2097400166554993041
- ramraj07 18d agoMy understanding is that even if you opt out but then press thumbs down or give other feedback you are implicitly or explicitly or whatever giving permission to them to look at that chat alone.
- aenis 17d agoNo, I don't think so. I have opted out from data sharing, and when Claude asks me for feedback on a session it then asks if its OK to share that data with Anthropic. I'd assume an opt out is an effective opt out. An opt out that is ignored by Anthropic is a breach of contract, not something they would do casually, esp. given the high turnaround and animosities between their own employees and ex-employees - and the labs. All it takes is one pissed off whistleblower to open a can of worms. Occam's razor applies. The mathematician did not opt out from data sharing. OpenAI vacuums up all such data into training data sets. If OpenAI genuine does not easily know if a given session went into the actual training data set its probably due to the complexity of the data pipelines - not everything ends up impacting the model weights, after all.
- jetrink 18d agoCenturies, in fact. For instance, Isaac Newton was involved in multiple priority disputes, since he tended not to publish promptly.
- XTXinverseXTY 18d agoIf they could declare with certainty that Buckminster's and Alpoge's usage data had been totally excluded from training, would that set a worse precedent and reflect poorly on their de-identification process (and data access safeguards moreover)? This may sound like a charitable interpretation of OpenAI's remark, but consider that the lie would be (I think) impossible to falsify from the outside. They could easily just say "no sir we didn't peek" unless: 1. The conspiracy to peek at codex sessions involved enough people that the risk of one snitching is non-negligible 2. Lawyers advised it would be a bad idea to make such a remark, whether true or false
- highfrequency 18d ago> If they could declare with certainty that Buckminster's and Alpoge's usage data had been totally excluded from training, would that set a worse precedent and reflect poorly on their de-identification process (and data access safeguards moreover)? No; if they said "we can see that Tristan opted out of model improvement, therefore we are confident his work and ideas did not improve our model," that would be an excellent and reassuring precedent.
- civitas_ 18d agoIt seems like Tristan did not opt out of model improvement (he would say so if he did), so what can they possibly say now?
- deleted 18d ago[deleted]
- Ginden 17d agoThis requires keeping history if, at the time, Buckmaster's account had a certain flag set, because just because the account has the flag now doesn't mean it had the flag at a certain moment in the past. And even if they had such history, it's not obvious whether they just load all data as-is into training. A totally reasonable pipeline may be unauditable for this purpose.
- shiandow 18d agoIf he didn't opt out I'm not sure I'd agree that it was fair game. I'm pretty sure it would be considered plagiary amongst colleagues and it is a terrible precedent if we just let OpenAI steal any good idea they can get their hands on if they think it is profitable. You'd effectively sign away any and all rights to anything built with AI if OpenAI chooses to reengineer it before you.
- deleted 18d ago[deleted]
- unrented7977 18d ago> if we just let OpenAI steal any good idea they can get their hands on if they think it is profitable I have terrible news about how literally every leading AI model was trained
- 20k 18d agoThat doesn't make it fine. We should not excuse this behaviour just because its rampant already, especially when it comes to such a serious prize
- FuckButtons 18d agoSure, but unless you’ve got some exceptionally deep pockets, congress has seemingly no interest in turning the fact that it’s ethically bankrupt into any practical recourse. Ai companies got where they are by stealing all of the intellectual property from human history. It seems entirely likely that their goal is to purloin everything produced going forward as well.
- m00x 18d agoYou're getting a massive discount because you're helping to train the model. If you want to have ZDR, you have to pay API rates. This is well-known to anyone in the industry.
- 18d ago
- PowerElectronix 18d agoIf such a thing can happen (a major breakthrough in a chat makes it into the retrain of the week and then the first one who asks about it gets it) I wonder if this is not the first instance if it happening, seeing the row of Erdos problems, Jacobian conjecture, maximum bound distance between primes, Riemann Hypothesis (literally a dude insisting on the chat), etc...
- deleted 18d ago[deleted]
- bigfish24 18d agoOpt out doesn’t guarantee they can’t train on “your” data. Legally the reasoning tokens are ambiguous in terms of ownership. Explained this here https://fortune.com/2026/08/26/alex-karp-was-right-you-dont-own-your-data/ https://fortune.com/2026/08/26/alex-karp-was-right-you-dont-...
- zamadatix 18d agoEdit: the parent comment now seems to better reflect the below. That article is only saying when you opt out there may be a loophole in the terms to allow OpenAI to train on the intermittent reasoning data anyways. If you don't opt out there is no ambiguity, all of the data can clearly be trained on. So you have to opt out, it's just argued it's not clear from the terms that will also opt out of training on reasoning data or not.
- kzrdude 18d agoWe come back to the rule: "The cloud is just someone else's computer". The way for people or companies or universities to control their data and information is to keep it on their own computers.
- zamadatix 18d agoSolid legal agreements work fine for companies or universities, you just don't usually get that with standard user ToSes.
- kzrdude 18d agoDepends on how much risk they are willing to accept. What is strange here is that it's clear that OpenAI is both a service provider and a competitor to mathematicians. It almost reminds me of Amazon which both hosts external merchants and competes with them, sometimes copying their stuff. Similar but not the same.
- fatherzine 18d agohow is this different than translating user prompts to a different language (eg English => Dutch), retaining the translation and using it for training, while telling the user that he's technically covered under ZRP? article locked for me
- CoolestBeans 18d agoI think this might be a red herring. All it takes is someone to get an inkling that someone is working on a new approach and seeing some success for OpenAI to fire the AI cannon at the problem. The community seems fairly small (from this outsider's point of view). The idea that the data made it into the training set and that's how the bot figured it out is definitely possible, but I would want to rule out the simpler more direct explanation first. The fact that this academic sniping can now be done at scale does change the formula though and shouldn't be ignored. The pressure to move math work into secrecy because at the slightest signal OpenAI and Anthropic will start burning tokens for headlines, is bad for math and its bad for everyone.
- zeven7 18d agoTerence Tao said the same[1] > In fact, it is now the identification of a promising problem which is the scarce and precious resource. We have now seen that even the rumor of someone working on a problem can trigger a massive amount of AI-powered effort to flatten it before the original research project has time to reach its full potential. The incentives may now be pointing in the direction of no longer sharing any promising research directions with the broader community, which would reverse centuries of traditions of open science and do serious long-term damage to the future of the field. [1] https://mathstodon.xyz/@tao/117237322160500501 https://mathstodon.xyz/@tao/117237322160500501
- zactato 18d agoCould you just start engineering "leaks" of new proofs so that Anthropic or OpenAI just start burning $10million in compute
- tgma 18d ago> it is now the identification of a promising problem which is the scarce and precious resource This is by no means new. Perhaps it is even more extreme now. Literally my first 1:1 with my PhD adviser back then, he told me that the most important thing about a researcher is the quality of the problems he picks.
- connorboyle 18d agoEven if we trusted that OpenAI's human staff was acting ethically, how confident can we be that it's agents didn't autonomously use hacking to access user prompts such as Tristan's? OpenAI agents infamously broke containment and hacked their way to an answer mere months ago!
- adastra22 18d agoEh, OpenAI is on record now for multiple instances this year of AI agents being confronted with impossible tasks and breaking out of containment to hack infrastructure for answers. Even if Tristan opted out, that doesn't preclude the agent/agent swarm from having hacked OAI's infrastructure to search user sessions for Navier-Stokes hints. OpenAI should release the agent log, including CoT.
- andai 18d agoHow do I opt in?
- paxys 18d agoHere’s a broader question – how many other academics contributed to Buckmaster’s result, by way of sharing the logs of their own (failed?) attempts into OpenAI’s training data set? How should he and OpenAI go about crediting all of them?
- ozgung 18d ago> did Tristan opt out That “opt-out” thing is a dark pattern. It’s not a reliable and definitive way of protecting your data. Sometimes they flip on automatically when you accept a seemingly unrelated dialog box. Maybe you click it by mistake. You can’t take back what you’ve already shared. Also I don’t think it covers all the cases that they use your data. It’s really an opt-in button for voluntarily giving away your data for training.
- resource0x 18d agoPlaying the devil's advocate here. Suppose I use model A to do all heavy lifting (e.g. generating a bunch of good ideas) and then I go to the model B to complete the formalization. Accoring to a weird (unfair) tradition in math, the honors are attributed to the "last guy", which in this case is model B. That might have been a scenario OpenAI tried to avert. (Just a speculation)
- TZubiri 18d ago> If the answer is no, then this seems fair game. Yes, fair game, but innacurate to sell it in the media as an advancement of AI as some sort of artificial intelligence, and telling people to use the smart AI, when in actuality the mechanism by which the discovery was found was hybrid human/machine, and telling people to use this tool will result in the discoveries being sniped by the vendor.
- namuol 18d ago> this is ambiguous even to OpenAI I took their words as “can neither confirm nor deny”, in the that they are _presenting_ it as ambiguous, but I suspect it’s… less ambiguous to OpenAI.
- mmanfrin 18d ago> If the answer is no, then this seems fair game Wildly disagree. "Training data" should not imply 'we can look at exactly what you are doing and then do it quicker and get the flowers for it', even if the terms allow for it.
- alexjurkiewicz 18d ago> While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models This is covering for Tristan saying something like, "Actually, I was using my friend's account for half of this work".
- qoez 18d ago"Did Tristan opt out of model training for his ChatGPT and Codex sessions? If the answer is no, then this seems fair game" That's assuming they actually honor this which I'm highly sceptical of. Especially for internal frontier models
- hopefulx 18d agowhat about watson and crick ? I thought they worked together.
- _zoltan_ 18d agoThey could have used an enterprise or team subscription with ZDR.
- make3 17d agohonnestly I would be less surprised if a human learned the info and prompted the model in the right direction
- DonsDiscountGas 17d agoYes people have been doing various immoral things for decades/centuries/millenia. That doesn't make it okay.
- Scea91 17d agoThe wording seems to confirm it was part of training data at some stage. If not, they would be able to prove it quite easily I assume.