6 ms·
Is this the one that was allegedly based on someone else's actual work & prompts? https://news.ycombinator.com/item?id=49605915 https://news.ycombinator.com/it
by pavel_lishin 25d ago
Is this the one that was allegedly based on someone else's actual work & prompts?
https://news.ycombinator.com/item?id=49605915 https://news.ycombinator.com/item?id=49605915
https://bsky.app/profile/quantian.bsky.social/post/3muyhwbcdys2x https://bsky.app/profile/quantian.bsky.social/post/3muyhwbcd...
https://cims.nyu.edu/~tristanb/statement.pdf https://cims.nyu.edu/~tristanb/statement.pdf
- beering 25d agoThat is addressed in the article.
- floatrock 25d agoOpenAI's position: > We (the researchers and the agents) did not see any of their work through any means until they released it publicly — in particular, no specific user data was accessed in order to solve this problem. While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models . However, our proofs differ significantly and even the precise results proved are different in the Euler case (forced vs unforced).
- cute_boi 25d agoI thought openai don't use any user data if we opt out of training and via api?
- andrewguenther 25d agoThat is correct. It is possible they didn't opt out and given the timeline and anonymization of data unclear whether a particular conversation would have made it into the training set if they hadn't.
- biophysboy 25d agoWhy is it unlikely?
- enraged_camel 25d agoBecause OpenAI says so, obviously!
- tristanj 24d agoBecause the models are trained on hundreds of billions of user conversations, across more than a billion different humans. The conversations are anonymized and not easily traceable back to a specific user. It's unknowable and not possible to prove if any one specific conversation contained the insights for solving Navier–Stokes. We also don't know if the authors unintentionally provided data to OpenAI through alternate means, such as via alternate accounts or model feedback queries.
- biophysboy 24d agoI understand that AI is not just cut and paste, but some documents will have more influence than others w/ power law scaling. I would be very surprised if this distribution were not extremely steep for arcane math
- JumpCrisscross 24d ago> It's unknowable Knowability and likelihood are almost orthogonal here. If I commit a crime and perfectly destroy the evidence, my deed may be unknowable. That doesn’t make it more or less likely. > not possible to prove if any one specific conversation contained the insights for solving Navier–Stokes It may be. We haven’t seen the researchers’ transcripts. We don’t know what Buckmaster or his co-author uploaded to OpenAI or with what permissions (or if OpenAI actually respects those toggles).
- tristanj 24d agoI think OpenAI desperately wants to make a blanket denial that they didn't look at or train on Buckmaster and Alpöge's chat transcripts, but know they cannot, because the data is anonymized. The fact they can't make a blanket denial triggers everyone's bullshit detectors, and they're getting eviscerated over it.
- tedsanders 25d agoYes, that was the allegation last night. I work at OpenAI, though not on the team that did this, and my understanding is: - we decided to ask our model for Millenium problem solutions because of two reasons: (a) our new model was looking incredibly good and (b) we heard rumors that some Millenium problems had been solved and were curious if our models could solve them (the goal here was not to scoop any particular individuals and we were looking at many problems beyond these) - we did not read any private chats (but of course the model was aware of prior research literature published to the internet) - the proof generated by our model was very different from theirs and also goes far beyond the published literature - we made an effort to jointly announce rather than immediately scoop (I understand Tristan was unhappy with the conversations; I know zero details here and I hope more is shared today) Edit: Here's is Sebastian's take: https://x.com/SebastienBubeck/status/2097379411691516310?s=20 https://x.com/SebastienBubeck/status/2097379411691516310?s=2...
- suddenlybananas 25d agoHow are people talking about this there? Why are so many employees posting nasty things about Tristan on twitter?
- tedsanders 25d agoCan you point me to any nasty things being posted? I'll ask them to delete.
- sk4rekr0w 25d agoHaven't seen a single post doing this on X or anywhere really from OAI employees. Only seen knives pointed at Sebastian on social media so this is extreme and shameful gaslighting.
- suddenlybananas 25d agoI do see people claiming he's abusive/unscrupulous which are pretty extreme allegations. https://news.ycombinator.com/item?id=49605915#49610498 https://news.ycombinator.com/item?id=49605915#49610498 https://x.com/dheeraj_nagaraj/status/2097266146445774924?s=61 https://x.com/dheeraj_nagaraj/status/2097266146445774924?s=6... (I can't reply to the below comment, but I was aware this was about Sebastien, I was trying to be charitable by including stuff said about both people)
- heaney-555 25d agoDid you actually read the article and the substance of the solution? >our proofs differ significantly and even the precise results proved are different in the Euler case (forced vs unforced)
- SpicyLemonZest 25d agoIt's not a meaningful response to the accusations. Any productive new research direction would be expected to lead to a number of different possible proofs of a number of similar problems. (Given their bizarrely compressed timescale here, it's possible that the proofs really are so different it's clear they came independently, and they just didn't have time to come up with that information before hitting publish.)
- octoberfranklin 24d ago> https://cims.nyu.edu/~tristanb/statement.pdf https://cims.nyu.edu/~tristanb/statement.pdf This really need to be a top-level story on HN.. I asked whether the model had been trained on, or had access to, our sessions in Codex, into which we had been putting all our drafts for the whole of this project. I was told the model did not look up user data. I asked again, about training, and I did not get an answer. This whole episode is more horrific than "AI is eating math". We now have a clear and economically damaging (or at least career damaging) example of the "training on customer tokens" problem. We can't ignore this problem any longer.
- dekhn 24d agoWe don't have any proof of that at all. Please stop rushing to judge without data.
- JumpCrisscross 24d ago> We don't have any proof of that It’s fair to give benefit of doubt to Buckmaster given OpenAI is currently being very credibly sued by Apple for openly stealing others’ original work in another context.
- dekhn 24d agoA prior of "one large company is current involved in an unrelated lawsuit with another large company" is pretty weak; the case is undecided and about an entirely different kind of IP theft. In short I think a lot of people are jumping to conclusions without supporting evidence and that's really not helping the situation.
- JumpCrisscross 23d ago> the case is undecided and about an entirely different kind of IP theft The guy mocked accessing his prior employer's circuit diagrams and was protected by OpenAI until Apple filed suit. Tabula rasa, sure, we need more evidence. But ignoring the priors should at least be explicitly acknowledged.
- scam-altman 24d ago[flagged]