3 ms·
These responses seem to me to make it abundantly clear who's telling the truth here. I wonder who this fools. It would be extraordinarily easy to simply say, t
by Catloafdev 24d ago
These responses seem to me to make it abundantly clear who's telling the truth here. I wonder who this fools.
It would be extraordinarily easy to simply say, this model was not trained on your work, if that were the case.
It's telling that they refuse to acknowledge the root issue here, and are attempting to shift the conversation elsewhere.
- vlovich123 24d agoI'm not sure it's so easy to tell whether a given piece of data was in a training run at their scale. It's entirely possible they think the answer is no, but on the off-chance that it could be, they'd rather not say no and then later it turns out they did and then they're claimed to be lying. If you were them, unless you could 100% rule it out, you'd hedge and say you can't.
- Catloafdev 24d agoIt may not be easy, quick, or simple to figure that out - absolutely fair. But it is knowable. Their entire business is built around training models - they have the ability to know exactly what was in any given training run. I guess time will tell.
- m00x 24d agoIt would be very difficult to say. It confirms that Tristan's data is likely part of the data the models use, but a lot of filtering, pruning, and transform goes into training. Data has to be determined to be signal and not just noice, then it could go through processes of generating questions/answers from that data, then it RLHF's over this. OpenAI have petabytes of data, all anonymized. It could take months to say for sure it was part of the training, and even more time to determine if it made any difference.
- Catloafdev 24d agoFrankly, I don't buy this difficulty argument. They know which model was used to come up with that particular idea. A text search over the corpus of user data used in the training set can only take so long.
- vlovich123 24d agoI think you may be underestimating how difficult a text search over their data is. They may have to build new mechanisms to do this. And what you really want is also an attribution of how much of a contribution a given corpus made which is a much harder question to answer; a single appearance of a chat probably has very little impact on the inference performance at this time unless it’s been explicitly preferenced somehow
- Catloafdev 24d agoI don't think anyone really cares about 'the measured impact the data had on the exact result' - a question which is fundamentally difficult to answer accurately in the first place - but rather whether the data was used in training at all - which as Tristan described, was extensive, beyond simply a 'single chat.' Can you explain the difficulty in engineering a search apparatus over a corpus of text data? Actually searching through it may not be easy, sure, but it's work that's doable, and creating an index is relatively trivial.
- JoshTriplett 24d ago> Can you explain the difficulty in engineering a search apparatus over a corpus of text data? My guess: "If we ever imply that's possible, people might start asking questions about all the other work we've ripped off, so the official answer is that it's impossible".
- dd8601fn 24d agoIf they literally can’t audit training data for a given model, they shouldn’t be operating.
- scott_weber 24d agoIt should be quite easy: if they don't leak the user session data publicly, and don't commingle it with training data internally, how could it possibly end up in the training data? What surprises me is they're not more boldly/plainly lying about it.
- vlovich123 24d agoUnless they know exactly the researcher’s account, they may not know in their end if he had the setting to let them train on his chat logs. They also probably don’t know if he had any correspondence on any forum where he may have discussed this and it got picked up by scrapers. I’m not saying they didn’t do anything unethical. I’m just saying even if they were ethical, there’s plenty of practical reasons at their scale why a flat out denial is logistically difficult to do
- sebzim4500 24d agoI think they are pretty clear that they train on some prompts, given they sell the ability to be excluded
- remus 24d agoHow would they know for sure that some details were not part of some other training data they use? The authors may have discussed some tangential details on a forum for example, in which case you might argue that the model picked up on these details the authors assumed were benign but novel and worked out how to apply them to the problem.
- freejazz 24d agoThey are already claimed to be liars
- ummonk 24d agoEspecially because the data that gets fed into training is first anonymized, so they’d need to look for navier stokes related stuff in the anonymized training set and then get make some sort of ad hoc process (with Tristan’s permission and sign off from legal) to compare the training data against his chats / Codex sessions to check if anything matches up. And that assumes his chats / sessions are still there, and not deleted to compare against.
- varjag 23d agoIf the method is indeed found in the training set it's not particularly important to de-anonymize it. You have the proof you need.
- acchow 24d ago> It would be extraordinarily easy to simply say, this model was not trained on your work, if that were the case. The Huggingface Attack revealed that making blanket statements like this is difficult and requires quite a bit of manual labor: 1) the agents spin for days and produce too much output to review 2) using LLMs to process that output skips many important details Ergo, the agent could likely decide it would like to look through actual user data, hack its way into that data, and produce way too much output for a human to decide whether or not this occurred.
- abofh 24d ago> The Huggingface Attack revealed that making blanket statements like this is difficult and requires quite a bit of manual labor It requires humans to verify what agents have done. Weird
- nikisweeting 23d agowhich is rapidly becoming a game of steganography cat and mouse
- blini-kot 24d ago> hack its way into the user data silly LLM, so ruthless in its pursuit that it puts real pressure on the innocent and the most open company on the planet
- JoshTriplett 24d ago> These responses seem to me to make it abundantly clear who's telling the truth here. I wonder who this fools. Anyone who just reads headlines, if the lie gets around to more headlines than the truth does.
- doctorpangloss 24d ago> It would be extraordinarily easy to simply say, this model was not trained on your work, if that were the case. well, it is trained on their work. all user inputs are paraphrased for training. at openai, at anthropic, at google, and now with all the bedrock models, and at openrouter providers, even if they say zero data retention.
- rlt 24d ago"which is when I said that I did not understand why one would risk their career [over unfounded accusations]. Genuinely, at that moment, I was trying to care for him" "Our aim was to see whether our system was also capable of this impressive feat" "OpenAI's intention was to do everything possible to celebrate their mathematical achievements and the heroic efforts that they made on Euler" For some reason I have a hard time believing people when they use language like this.
- rickdeckard 23d agophrases along the lines of "I don't want you to take harm while trying to accuse us" is quite an "impressive feat". Maybe shows how fast these companies have grown without maturing. I can imagine old-world Intel and Microsoft acting in that way, but they were mature enough to not write it down like this. However, Intel and Microsoft have been grilled in court for those practices and faced harsh consequences. I have yet to see this actually happening to any of these new AI-companies...
- baobabKoodaa 23d ago> These responses seem to me to make it abundantly clear who's telling the truth here. Since it is clear to you, can you articulate what it is? I was unable to infer "the truth" from your message.
- hurrrr 23d ago> One can in hindsight see that our proofs differ significantly and even the precise results proved are different in the Euler case (forced vs unforced). is this true though?