6 ms·
I'm not one to comment often but this really pisses me off. OpenAI looked at user data, stole world class researchers' work, and then tried to threaten those r
by mayakacz 26d ago
I'm not one to comment often but this really pisses me off.
OpenAI looked at user data, stole world class researchers' work, and then tried to threaten those researchers to do what would make their corporation profit (which they would anyways!).
Imagine you have been working on a terribly difficult math problem for a decade. This is a result you have spent years on, and what you will likely be remembered for. And to have some punk from OpenAI lie to you, threaten you, and tell you that they are willing to go on the record that you "deserved" it? What is this, the Godfather?
If OpenAI solved Navier-Stokes, that is an astounding result! - yet they'll still be remembered as those who thought credit was more important than results. That winning was more important than collaboration. If this is true, they're burning any trust left with academia.
- sk4rekr0w 26d agoYes, but only if you take this one sided statement at face value.
- thereitgoes456 26d agoI’m open to evidence, but just using Bayesian reasoning, OpenAI is one of the most dishonest companies in history. They’re currently being sued for a dozen employees stealing Apple hardware! I don’t understand why I should give them any grace.
- tigershark 26d agoWhy would have they rushed the publication if this was not true? Are you also suggesting that he fully invented the call with Open AI?
- famouswaffles 26d agoThe results being true, the 'deal' that was made being true doesn't mean some of the implied accusations here are true, for example - that Open AI used their Codex logs to drive their breakthrough.
- tigershark 26d agoSo are you baselessly assuming that he is lying? He explicitly reported that he was threatened and your answer here is to defend OpenAI no matter what.
- famouswaffles 26d agoDo you not have reading comprehension? Did you even read the statement? He himself asserts at the end he doesn't know if the above example is true or not. What on earth are you going on about? Where in my comment am I assuming he's lying ?
- tigershark 26d agoFrom my post above: > He explicitly reported that *he was threatened* and your answer here is to defend OpenAI no matter what. Yes, I read fully the statement, what about you? Do you know what is a threat? What is this in your super-humble opinion if not a threat: > The reply was, “Why would you ruin your career?” I replied that I am an academic, and asked why he thought going public would ruin my career. The reply was, “If you don’t want me to be nice, then I don’t have to be nice.” And you are saying that I don't have reading comprehension...
- sk4rekr0w 26d agoYou really lack reading comprehension
- tigershark 26d agoMy reading comprehension is pretty good, I'm not the one that doesn't recognize a threat even when it's perfectly clear. Verbatim from the statement: > The reply was, “Why would you ruin your career?” I replied that I am an academic, and asked why he thought going public would ruin my career. The reply was, “If you don’t want me to be nice, then I don’t have to be nice.”
- PowerElectronix 26d ago
- scurnus 26d agoThings are more entangled than that. The contribute made from both OpenAI and Anthropic models to solve these problems are clear, now it really hard to quantify which one contributed more, if the role played by the human is major or minor. OpenAI tried to collaborate and share the results together with a fixed timeline, to avoid this mess but it was inevitable. There is a conflict of interest, where the other researcher works at Anthropic, who will also try to take credit. Where they may be in the wrong is if they took user data regarding the problem, how will we know if they did or not?
- Ar-Curunir 26d agoThey offered to collaborate by asking to drop a coauthor. That is not collaboration, and is not an academic norm.
- postalcoder 26d agoI'm stunned that people are taking this accusation as a fact. OpenAI is no stranger to rivalry with Anthropic but 1. it's not like user data is sitting around on some kitchen table somewhere and 2. I consider OpenAI to be as economically motivated as any other actor in this space and playing around with user data like that would destroy their business. There are things that Buckmaster alleged and things that he speculated. The entire training data thing is speculation. If this is pissing you off, then you ought to evaluate how you ingest information.
- thereitgoes456 26d agoHe asked whether they used their chats as training data and received no response. Any speculation here seems quite appropriate?
- deleted 26d ago[deleted]
- dash2 26d agoHe didn't even make that accusation! > I asked whether the model had been trained on, or had access to, our sessions in Codex, into which we had been putting all our drafts for the whole of this project. I was told the model did not look up user data. I asked again, about training, and I did not get an answer. The shocking/interesting thing would be if it was trained on the sessions. I think it's very implausible that they gave the model access to someone else's sessions as input. That would be a huge privacy violation and would probably blow up a large proportion of their enterprise business. Does openAI train on user conversations in general? I assume so. But so fast as that? That seems unlikely in general. I expect OpenAI will come out denying this.
- actionfromafar 26d agoCouldn't the Enterprise have a different fine print?
- johnnienaked 26d ago
- tristanj 26d agoThis conclusion is flawed. It's unclear at this point if OpenAI's model or employees actually looked at or stole the author's data. Having worked at large companies before, I'm leaning towards no, since very few employees have access to that data. And simply knowing a problem can be solved is half the battle.
- dbdr 26d agoFrom Buckmaster's text: The route to the Clay problem through a smooth force, options c and d in Fefferman’s statement of the problem, is the route Luis and Diego opened and the one Levent and I had quietly chosen to attack. Almost nobody else I know of was working on it. It is not the direction one arrives at in a few days by giving a model the problem statement. When I heard “forced,” it was a bright red flag. This is much more than the knowledge than the problem can be solved, it's also the specific, non-obvious approach to solving it. That's much more damning for OpenAI, if confirmed.
- Keyframe 26d ago[flagged]
- tristanj 26d agoThat's a stretch. The Luis and Diego paper was published in 2023 and is included in every frontier model's training dataset. An AI model could independently choose the same path route as Luis and Diego, without access to Buckmaster and Alpöge’s work. And the article states "an insane amount of compute had been used," which implies OpenAI brute-forced their way to a solution. I.e. they searched for every paper published on Navier-Stokes and exhaustively attempted every approach. Such an approach would lead them to a solution. There is not enough information at this time to reach a conclusion. The best option is to wait for statements from both sides, then reevaluate.
- mishellaneous 26d ago> An AI model could independently choose the same path route as Luis and Diego, without access to Buckmaster and Alpöge’s work. the post you were replying to quotes Buckmaster specifically denying this: "It is not the direction one arrives at in a few days by giving a model the problem statement." > And the article states "an insane amount of compute had been used," which implies OpenAI brute-forced their way to a solution. I.e. they searched for every paper published on Navier-Stokes and exhaustively attempted every approach. Such an approach would lead them to a solution. "implies" is a surprising choice of word here. that's certainly one interpretation of "an insane amount of compute had been used". what came to my mind, considering Buckmaster's statement that the AI would not head down this specific path on its own, is, though, that they prompted it in this specific direction and then used an insane amount of compute. this seems consistent as well with these other statements: > Over the course of the call, as members of their team sent Sebastien corrections and details over their internal chat, it emerged that an entire team had been working on the problem, that this was one of a number of things that was tried, that work had started on the unforced problem, that the team first set the model on easier problems, including Euler (...)
- johnnienaked 26d agoLLMs do nothing but steal, and the companies that own them are fully aware and eager to do it.
- Sol- 26d ago> OpenAI looked at user data, stole world class researchers' work This doesn't seem to be clear and is very implausible for a large company. Be as cynical as you want, but a normal researcher will simply not have access rights to this data, which will be siloed away somewhere else. It might very well be somewhat unfair to catch wind of a promising approach and then try to frontrun them by throwing compute at the problem, but this isn't really the same.
- alas44 26d agoNo, plausible given AI companies want/need session data to train their next models. Probably not someone peeking an eye to sessions directly, but probably not so hard to find the useful sessions in anonymized training data to post train a model on. As stated in the paper, OpenAI did not explicitely denied the researcher sessions were not used for training the model. So either they don't know, or don't want to tell "I asked whether the model had been trained on, or had access to, our sessions in Codex, into which we had been putting all our drafts for the whole of this project. I was told the model did not look up user data. I asked again, about training, and I did not get an answer." Let's see what statement OpenAI will come up with for their side of the story EDIT: precised my thought on user data vs session data
- alas44 26d agoI was downvoted initially, look what OpenAI shared https://openai.com/index/navier-stokes-solution/ https://openai.com/index/navier-stokes-solution/ ... "Since August 28 we have been training a new internal model that has exhibited unprecedented performance in our benchmarks, including mathematics. This model’s training is ongoing and its performance continues to improve." "When a further trained version of our internal model became available over the course of the effort" "While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models "
- cududa 26d agoThey likely train on logs.
- 26d ago
- paxys 26d ago[dead]
- ozgung 26d agoThis Godfather-like threat in particular pissed me off as well: > I said that if OpenAI released its result in the way proposed I would go public with what happened. The reply was, “Why would you ruin your career?” I replied that I am an academic, and asked why he thought going public would ruin my career. The reply was, “If you don’t want me to be nice, then I don’t have to be nice.”
- goolz 25d agoThis is what is truly problematic. People at OAI probably think they can do and say whatever they want.
- rsrsrs86 26d agoNot only did they look at user data The whole business is based on reselling user data scraped from the whole internet It’s plagiarism at scale
- deleted 26d ago[deleted]
- davidguetta 25d agohow much of the "codex discussion" was actually ideas also generated by open ai ? this is a bit the elephant in the room