5 ms·
Both Sam Altman and Sebastien Bubeck admitted they only want Buckmaster to be the lead author on a rewrite of the OpenAI proof. https://x.com/sama/status/20973
by hkmaxpro 24d ago
Both Sam Altman and Sebastien Bubeck admitted they only want Buckmaster to be the lead author on a rewrite of the OpenAI proof.
https://x.com/sama/status/2097385167002415140 https://x.com/sama/status/2097385167002415140
https://x.com/SebastienBubeck/status/2097379411691516310 https://x.com/SebastienBubeck/status/2097379411691516310
A wake up call for using OpenAI models. If you discover something with their model and you work for a competitor, they “felt it would be inappropriate” for you “to author OpenAI’s work”.
- derangedHorse 24d agoIt sounds like OpenAI is trying to appease the author when they don’t have to by allowing him to rewrite their proof. They probably don’t believe he deserves to, so him asking for a coauthor from Anthropic might overextend their grace in their eyes.
- nezi 24d agoGiven that Tristan has said that the proofs that LLMs come up with are mostly "slop" and not up to the standard that human written papers achieve, maybe OpenAI needs an expert like him more than you think to get the result published?
- fn-mote 24d agoAlmost certainly the Lean proof needs to be decoded for humans and probably also made “human intelligible”. Now maybe LLMs can also simplify arguments and make sense of them for humans, but we haven’t seen that yet (unaided). (I haven’t looked at it, personally.)
- nl 23d agoIt's a Navier Stokes solution. They don't need any help publishing.
- reverius42 24d agoHe's not "asking for a coauthor from Anthropic"; he already has a coauthor, who he's already been collaborating with, who happens to also be employed by Anthropic (but whose research in this area is not done as part of their employment at Anthropic).
- Catloafdev 24d agoThese responses seem to me to make it abundantly clear who's telling the truth here. I wonder who this fools. It would be extraordinarily easy to simply say, this model was not trained on your work, if that were the case. It's telling that they refuse to acknowledge the root issue here, and are attempting to shift the conversation elsewhere.
- vlovich123 24d agoI'm not sure it's so easy to tell whether a given piece of data was in a training run at their scale. It's entirely possible they think the answer is no, but on the off-chance that it could be, they'd rather not say no and then later it turns out they did and then they're claimed to be lying. If you were them, unless you could 100% rule it out, you'd hedge and say you can't.
- Catloafdev 24d agoIt may not be easy, quick, or simple to figure that out - absolutely fair. But it is knowable. Their entire business is built around training models - they have the ability to know exactly what was in any given training run. I guess time will tell.
- m00x 24d agoIt would be very difficult to say. It confirms that Tristan's data is likely part of the data the models use, but a lot of filtering, pruning, and transform goes into training. Data has to be determined to be signal and not just noice, then it could go through processes of generating questions/answers from that data, then it RLHF's over this. OpenAI have petabytes of data, all anonymized. It could take months to say for sure it was part of the training, and even more time to determine if it made any difference.
- Catloafdev 24d agoFrankly, I don't buy this difficulty argument. They know which model was used to come up with that particular idea. A text search over the corpus of user data used in the training set can only take so long.
- davesque 24d agoHonestly this whole thing is so fucking weird. I feel like there's an argument that absolutely no one involved in the final crossing of the finish line to the proof actually did any work (other than just intelligently directing an LLM) and deserves any credit. As the author of this doc mentions, the mathematicians who did the actual work that led to the formulation of this approach (without the use of LLMs; just good ole' fashioned human intellect) are the ones who deserve the credit. Imagine that a no name janitor used their time in the evenings to go spelunking through the literature to push an LLM to this result. No one would care because that person isn't an anointed expert. So why would the expert deserve any more credit? Because they sort of understand the result, even if they couldn't have achieved it on their own? The whole issue of credit for AI-assisted discoveries seems like it's going to run into a brick wall pretty soon.
- GPerson 24d agoI agree for most of the people in the story except Buckmaster himself seems to have been supplying real ideas.
- dudeinjapan 24d agoSounds like the plot for Good Will Hunting 2.
- dd8601fn 24d agoGood Will Hunting 2: Hunting Season is taken. It’ll have to be Good Will Hunting 3.
- bitwize 23d agoThe Hunting for More Money
- johnnienaked 23d agoToken Bill Hunting
- WD-42 24d ago
- aaron695 24d ago[dead]
- lalalanananana 24d agoHere's a wake up call for everyone sending all of their ip to openai and anthropic. Especially in verticals they intend to dominate. Lol at all the biotech companies all in on Claude and paying millions in fdes creating huge lapses in security as they go.
- johnnienaked 23d agoIt's too late. Sub models are deployed at every major organization in the United States and all it will take is turning off the option to improve the model for them to train directly on your own personal workflow, which CEOs will greedily eat up instantly if they can reduce labor costs. If they can brute force N-S, automating your finance or SWE job will be trivial. GG to most jobs connected to a computer in the next 5 years. BTW, this was always the plan from day 1. You will pour all your training and experience into training the model and receive a pink slip as compensation.
- segmondy 24d agoFor those of us who are into local models and preach it, we are called paranoid. I have often said this, if you are doing any real novel work, or putting your profitable business data/workflow into these models, you're a fool.
- JbMaj9 24d agoA mathematician working for the competitor, solving a math problem using our model? Isn’t that the best marketing possible?
- deleted 24d ago[deleted]
- oliculipolicula 24d ago>only want Buckmaster to be the lead author Sam+Seb are struggling with their ideological allegiance. This amounts to a confession that there are no reseaechers, only research managers, left at OpenAI. Maybe they even know that they are losing credibility from their main investor(s). They desperately need a domain expert to salvage credibility. They have no credibility with academia left, obviously, but their main competitor still does. No Millennium prize incoming, I'd wager. For openAI. Let's see mAth get political for once!! One might be more certain that levent is now going to corner all the institutional support. Go go go!
- Tanjreeve 23d agoIf the work done is just "we made other people's work searchable without their consent" it's not quite the same as what they're implying in the marketing of "our model solved this problem".
- deleted 24d ago[deleted]
- esalman 23d agoI work in catastrophe risk modeling and it's a multi billion dollar industry. We often chat where the business might be heading in future. An uncomfortable scenario is what if a frontier tech company decides to offer our customers the same products that we do. There's a lot of pressure on AI adoption so the company has partnered with various tech companies to build intelligent systems on top of proprietary data and mathematical models. If OpenAI is indeed using customer data to train their models to win a $1m prize, then it throws a giant IP question at the partnerships that affects multi billion dollar businesses.
- DeepSeaTortoise 23d ago> If OpenAI is indeed using customer data to train their models to win a $1m prize Is that even a question? Of course everything not kept on premise at gunpoint is going to be trained on. The chances of getting caught are 0 and the consequences of getting caught are 0 (as we've seen with copyright laws going from sending people to jail for years to unenforced within months). Yet the benefits are through the roof. Your customers aren't going to pay for having the very same data vibe enriched twice, it's exclusive, extremely high value data your competitors will never have access to.
- rickdeckard 23d agoAgree, I think the practice is also very clear from the overall strategy of AI-companies and their ToS: Scale with subsidized pricing as fast as possible to gain more user-data for training --> Own the better model --> scale pricing. Scanning social media (e.g. Twitter, Reddit) posts only give a glimpse into the thought-process, chat logs on-scale give you the actual process in machine-readable format. There's a reason why Google considers the Emails of Spirit Airlines to be worth millions of dollars [0], they give insights into a process, not just into the results... [0] https://www.axios.com/2026/08/17/google-spirit-airlines-bankruptcy https://www.axios.com/2026/08/17/google-spirit-airlines-bank...
- leonidasrup 23d ago> - Tristan is suspicious of the timing, as only few others were trying this approach. OpenAI says the model didn't access his user data directly, but leaves unanswered whether Tristan's chat conversations were part of the training. The question, for AI customers, is when they build products using services of AI-companies, would AI-companies engage in theft of customer data for use in training?
- ghshephard 23d agoIf this is a "wake up call" - then your legal team needs immediate education. First - there is this - https://openai.com/policies/how-your-data-is-used-to-improve-model-performance/ https://openai.com/policies/how-your-data-is-used-to-improve... (linked from the Navier Stokes writeup) I don't know how much more clearly they can write: > When you use our services for individuals such as ChatGPT, Sora, or Operator, we may use your content to train our models. One of the key selling tactics that companies like Data Bricks or Palantir provides their customers is "Data Governance" - that is, some control over where the data is being used. It's also a reason why enterprises don't use the OpenAI or Anthropic APIs directly - but through secondary sources that have Enterprise Agreements that do their best to make sure that no Company IP is ever retained by a third party, or even exists on a multi-tenant GPU. AWS Bedrock, and companies like together.ai, fireworks.ai have tons of deals that focus very much on data confidentiality. The reality is - if you want any type of control - you run your own inference, on your own hardware. Anything else and you are at the mercy of third-parties, despite what their contracts might promise you.
- pms 23d agoChatGPT has this option "Improve the model for everyone" in user preferences, which comes with the attached description, meaning that training on user data can be deactivated: > Allow your content to be used to train our models, which makes ChatGPT better for you and everyone who uses it. We take steps to protect your privacy. Learn more The "Learn more" link takes you to the link you've shared.
- pesacharia 23d agoAs I understand it, there is substantial question as to whether that actually stops them training on your data, it just perhaps changes what derivative processes are applied and used.
- ghshephard 23d agoAll I can say with certainty that the legal team at our typical Multi-Billion dollar Silicon Valley company had zero faith in any licensing arrangements with Anthropic or OpenAI, regardless of they $$$ involved, and that even getting to the point where Amazon Bedrock on Dedicated GPUs (we're already a big AWS customer - so definite cost advantages to dealing with them) - took 4-6 months of legal review before we could allow our engineers to start using Claude and OpenAI coding agents. Still can't use Fable because of their Data Retention requirements.
- whateverboat 23d ago> Now that we can see their work, the approaches appear to be different. It is also worth noting that our latest model can solve many, many other math problems. Is that a threat?