3 ms·
I'm stunned that people are taking this accusation as a fact. OpenAI is no stranger to rivalry with Anthropic but 1. it's not like user data is sitting around
by postalcoder 19d ago
I'm stunned that people are taking this accusation as a fact.
OpenAI is no stranger to rivalry with Anthropic but 1. it's not like user data is sitting around on some kitchen table somewhere and 2. I consider OpenAI to be as economically motivated as any other actor in this space and playing around with user data like that would destroy their business.
There are things that Buckmaster alleged and things that he speculated. The entire training data thing is speculation. If this is pissing you off, then you ought to evaluate how you ingest information.
- thereitgoes456 19d agoHe asked whether they used their chats as training data and received no response. Any speculation here seems quite appropriate?
- deleted 19d ago[deleted]
- dash2 19d agoHe didn't even make that accusation! > I asked whether the model had been trained on, or had access to, our sessions in Codex, into which we had been putting all our drafts for the whole of this project. I was told the model did not look up user data. I asked again, about training, and I did not get an answer. The shocking/interesting thing would be if it was trained on the sessions. I think it's very implausible that they gave the model access to someone else's sessions as input. That would be a huge privacy violation and would probably blow up a large proportion of their enterprise business. Does openAI train on user conversations in general? I assume so. But so fast as that? That seems unlikely in general. I expect OpenAI will come out denying this.
- actionfromafar 19d agoCouldn't the Enterprise have a different fine print?
- johnnienaked 19d agoIt wouldn't be shocking at all. They stole human data to train the first models and they've been stealing it ever since to train new models. Stealing mathematicians private chats and private research and taking credit for it would absolutely be par for the course. Enterprises are well aware of it and are fully on board. You didn't think every corporation in America has an OpenAI subscription because the models were good, did you? The whole reason they have subs is to train them on YOUR WORKFLOWS lol
- howdareme 19d agoThey are not training a whole model in a matter of days
- xdavidliu 18d agopretraining is months but they can totally fine tune in a few days
- irthomasthomas 18d agoThey where working on the problem for a year using codex.
- johnnienaked 18d agoThey don't need to train a whole model. They can feed it new information and fine tune it.
- impossiblefork 18d agoModels are very obviously continuously updated. Model editing to remove PII that slipped through, all sorts of things of that sort.
- revolvingthrow 19d agoIt would be shocking if it wasn’t trained on sessions. Have you read the ToS parts for both openai and anthropic that talk about it? It’s so obviously a weaselly way to say "no we do not train on your exact chats but we talked with legal and we think a cleanroom reimagining of your convo is probably fine and frankly where else are we going to get such a treasure trove of training data?" There’s potentially trillions on the line, do you seriously expect those companies to adhere to laws and regulations any more than, say, uber? The only unlikely part is the timeline - your sessions from a week ago probably haven’t made their way into the model. It’ll just take a while longer, and will be massaged just enough so that it isn’t really your exact session word for word so you can’t sure as easily.
- blini-kot 18d agoat first I thought your post was a bit revolting with "have you read ToS?" bit, but in the end I completely agree and understand I also don't get why it was downvoted, other than due to people not reading past the first sentence - although in the modern world's attention deficit that is understandable too
- Traster 19d agoParse that statement more carefully. > I was told the model did not look up user data. The naive way to read this is "Nothing you guys did influenced the way our model got to the solution". The less naive way to read this is "Of course the model isn't looking up your user data. I (the guy trying to blackmail you to remove the Anthropic employee from credit on your paper) looked up your sessions, and tipped our model off on how to solve this problem".
- YeGoblynQueenne 18d agoDuh. There are supposed to be limits to what OpenAI is allowed to access with respect to logs and user interactions but there is no technical limitation. It's a bit like sending unencrypted messages through a messaging app and the developer having a TOS that says they don't look at your messages. They might not, but they are fully capable of doing so. If they have a reason to do it, they will. Nobody's stopping them.
- ncruces 18d agoI read that as “the model didn't look up user data” as part of a “tool call,” i.e. they don't have an internal tool that loads user data (chats, sessions, attachments) for their internal models to read online while working. Or (likely) they do have it, but the model didn't use it (unless it's so powerful it escaped that guardrail, wouldn't that be ironic?) They declined to answer about anonymized aggregated user data being used for training. And even then, they may weasel out that they don't train on your “input” words, but that it's fair game go train on their “output” to your words.
- YeGoblynQueenne 18d ago>> Does openAI train on user conversations in general? I assume so. But so fast as that? That seems unlikely in general. I expect OpenAI will come out denying this. How "fast" does it have to be? Buckmaster and Alpoge have been working on this for just a day short of a year. See Alpoge's tweet announcing his collaboration with Bukmaster dated 9/19/25: https://x.com/__alpoge__/status/2097206973418611054 https://x.com/__alpoge__/status/2097206973418611054 It takes a few months to train a model these days but not a whole year. OpenAI had all the time to train on Buckmaster and Alpoge's results of just a few months earlier at which point they must have been well on the path to their result.
- mentalpiracy 18d agoWhy would this be implausible? ChatGPT user sessions were found publicly exposed to the internet not too long ago. Moreover, OpenAI has continued to play a hype-marketing game by revealing how their models keep breaking out of the sandbox. Conspiracy minded thinking is not helpful, but why should OpenAI be granted the benefit of the doubt here after being caught doing underhanded/negligent shit on several previous occasion?
- charcircuit 18d ago>ChatGPT user sessions were found publicly exposed to the internet not too long ago Do you mean publicly shared chats were able to be accessed by the public? That's the point of the feature.
- mswphd 18d agoopenAI's claimed solution uses a model trained in the last 2 weeks. The prior work would definitely be included in the training set.
- s1artibartfast 18d agoAnd the labs are all building panel of domin expert models, while simultaneously chasing open math problems. I would be shocked if they weren't tuning those models with the most relevant math texts and user material
- deleted 17d ago[deleted]
- johnnienaked 19d agoThey stole it.
- FiberBundle 19d agoWhenever I see comments defending AI companies, I look at the account's creation date, and interestingly almost all of them were created post 2024.
- az226 19d agoI don’t think you understand how brazen big tech companies are in practice.
- paxys 18d agoThis is how internet discourse works on Reddit/Twitter/HN and the rest. Someone said something which confirms your biases so it’ll now be treated as a fact and repeated endlessly in the echo chamber.
- defmacr0 18d agoI am stunned anyone is giving OpenAI the benefit of the doubt
- black_rabbit_ 18d agoI'd be surprised if all of it is organic discussion, shall we say. I reckon The Bot Factory just possibly might dogfood the astroturf machine.
- sensanaty 18d agoI guarantee you most of the comments regarding this aren't real humans. The homepage is full of crap meant to distract from what OAI did here, the comments are full of OAI employees. Dead internet theory pushed to the max
- hgoel 18d agoEspecially after the blatant cover up of their uncontrolled bot swarm infesting the internet, and the feckless "hopefully we do better" response upon being caught, I don't think OpenAI deserves much grace until they properly explain themselves. We had all assumed that surely the supposed smartest engineers in the world, with access to the most computing and a direct view of model capabilities, would take sandboxing and cybersecurity much more seriously than they have turned out to do. It follows that while we might assume they take user data privacy seriously and have tight controls on who can access it, it's possible they do not actually do that. At this point any initial trust is dead and has to be re-earned.
- haxiomic 18d ago> playing around with user data like that would destroy their business. Their entire business is based on stealing data. They can make a calculation that the cost stealing data is less than the cost of the positive publicity they can shape for solving Millennium NS
- sherburt3 18d agoGiven the history of OpenAI and current litigations, I would say they've developed a bit of a reputation for not respecting intellectual property. I'm dubious they have some unbreakable moral code that would prevent them from viewing and using user data.
- ozgung 18d ago> The entire training data thing is speculation. I think it's safe to assume AI labs DO train on your data and it's very hard to prevent that. I've just checked my inaptly named "Help improve our AI models" toggles. The toggle on the Claude settings had magically turned on. I asked about how this can happen. Claude says they show re-consent modals when terms change, and it is a "real and fairly common pattern" to re-opt in without noticing. All my work and conversations since I don't know are now part of their training corpus. No way to take it back. Google's Gemini/Antigravity didn't have opt-out toggles at all last time I checked. Codex also has a separate "include environments" setting which is hard to find (found it in Codex Cloud) and I don't know what it does. Lots of Dark UI Patterns here even if we assume they keep their promise. For this incident, Occam's Razor says their internal models somehow saw a version of the mathematicians' logs, during or after training. Maybe indirectly. These systems are literally designed to collect data. Privacy and safety is not trivial to achieve on the users' side. Simply because it's against the labs' best interest.
- cyclopeanutopia 18d agoBut the whoreshippers of The Holy Dollar will tell you it's all good and justified.
- cma 17d agoYou can opt-out of training in Gemini on personal plans, but it disables your chat history, just to be vindictive; there is no technical reason and the other companies don't do this.
- landdate 13d agoI prefer this anyway. And its not vindictive.
- iaw 18d agoAbsence of evidence is not evidence of absence. With the behaviors we know OpenAI engages in the accusations are wholly believable.
- ajkjk 18d agoEverybody knows it's not a sure thing, it's a question of trustworthiness. OpenAI is not trustworthy at all; this random researcher is and seems honest so far. iThe fact that people are corroborating Bubeck being a piece of shit in other settings add to credence. But nobody is over here saying it's an indisputable certainty. And your (2) is probably false, their history of deception suggests they would do just about anything as long as they didn't think it would backfire on them publicly.
- Tanjreeve 18d ago>The entire training data thing is speculation Quite literally in the terms of use.
- timdiggerm 17d ago> I consider OpenAI to be as economically motivated as any other actor in this space and playing around with user data like that would destroy their business. It's a gamble that this would be overlooked compared to the reputation they build for solving the thing
- sp527 17d ago> I consider OpenAI to be as economically motivated as any other actor in this space and playing around with user data like that would destroy their business. And that's exactly why you can only see evidence of this when the stakes are high enough, such as solving a Millenium problem. The legalese they outputted in response to the incident left them an escape hatch that permits the possibility of theft. People familiar with corporate damage control should recognize the verbal maneuvering, often used to paper over actual guilt. There was also already a high prior of shadiness. The company is run by someone who is close enough in reputation to "known sociopath". It's actually far more reasonable to assume that OpenAI stole user data to generate a breakthrough. They have an unbelievably high economic motive to do so. And even a low chance that they might be doing this implies catastrophic risk to anyone with valuable knowledge.