4 ms·
Oai denies looking at prompts but doesn't deny training on them.
by Davidzheng 25d ago
Oai denies looking at prompts but doesn't deny training on them.
- lukewarm707 25d agoif they do not deny training on them, they can't deny plagiarism.
- fc417fc802 24d agoBy that logic everything any LLM spits out is plagiarizing the vast majority of work written prior to a few months ago. That doesn't seem like a useful or desirable line of argument to me.
- myrmidon 24d agoJust replace the model with a human student. "Training" on textbooks => fine "Training" with unpublished notes from another professor, then publishing something on that exact topic with a similar approach without giving any credit => extremely questionable.
- fc417fc802 24d agoPresumably the professor voluntarily provided the notes in this analogy. I think the student would also be expected to cite the textbook if building off of it directly. In contrast, humans are generally not expected to cite "general inspiration" or what have you. So if we're to apply human standards, and assuming that the model was trained on the relevant work, it would only be plagiarism if the model directly built upon that previous work (at least IMO). The trouble here is that if LLM training constitutes direct use then approximately _everything_ they output is blatant plagiarism, not just a few pieces of academic work. Conversely if training is viewed as analogous to a student attending classes to learn general concepts (not a perfect analogy, I realize) then nothing they output on their own (as opposed to receiving as part of context) is plagiarism. Thus this seems like a fairly useless line of argument to me as far as the current topic goes. It either implicates this academic work along with literally everything else or else it does not implicate this academic work. Kind of like nuking an entire city and then saying "mission accomplished, killed the bad guy".
- anonymousDan 24d agoThis is just a nonsense line of reasoning. Training based on the solution to the problem (or the key insight behind the problem) is clearly a form of plagiarism.
- fc417fc802 24d agoWhat about my line of reasoning is nonsense? I made no claim either in support of or contrary to yours. Rather I pointed out that by this logic literally everything that an LLM spits out is plagiarism of the vast majority of the entire body of human literature in existence. Can you offer meaningful refutation of that observation of mine?
- freejazz 24d agoWhat does it matter? We're supposed to not call it plagiarism anymore because it's inconvenient to call it the plagiarism machine? What's your actual argument? Otherwise it's completely irrelevant what an LLM does in other contexts or what we call it
- za_creature 24d agoMany do indeed hold the position that all LLM output is uncopyrightable plagiarism. They're probably right, but there's an even stronger argument here: Science papers of a phd level must contain: 1. one or more novel insights 2. a long list of citations to contextualize them and 3. some work to prove that the insights are in fact meaningful --- In this context, consider a prompt based diffusion model which, when asked, will happily produce a few pictures of a horse in orbit. You then tell it "silly robot, horses can't breathe in space" to which it adds the necessary space suit in a follow up image. That image is twice plagiarized: 1. the model did not come up with the original idea of putting a horse in space, nor with insight that horses need a space suit 2. the model failed to cite where it pulled the "horse" and "space" concepts from. It merely did the work (3) to combine the concepts using the user provided insight. --- The implied accusation here is that OpenAI used the insights from an existing prompt to train a new model that was able to one shot "a horse race in space" picture, and they were all wearing space suits. This is still academic plagiarism, even if you disagree that all LLM outputs are.
- demibabs 24d agoIsn’t that one of the most salient and straightforward argument against LLMs?
- persedes 24d agoJust overfit ad infinitum:)
- lukewarm707 24d agoas good academic conduct you may cite the source of the work you are quoting or paraphrasing. as bad academic conduct you may steal someone else's unpublished work, work on it yourself for a bit, and then publish it as your own work. and then threaten the original author!
- hellohello2 24d agoThis is a common misconception, so its understandable that you have it. Generative models can both plagiarize and generalize. The question here is which of the two happened.
- fc417fc802 24d agoA needlessly condescending tone while failing to address the topic at hand. The person I replied to advanced the claim that training was sufficient to constitute plagiarism. You appear to be claiming that it is possible to generalize instead of plagiarize after training on something, so I take it that you must necessarily disagree with the original claim?
- hellohello2 24d agoWhat I meant to say is that, in many cases, a generative model's output is not in fact steered by minor amounts by lots of training samples, but instead steered by a just few samples. Some outputs are influenced by many inputs, and some by very few, it really depends. In answer to a post suggesting that training on a datapoint could mean plagiarism, you said that this would imply that all outputs are plagiarized. This is not the case, no, because generative models do not "copy" or "create", they do both at different times. I did not agree or disagree with the original poster, I was explaining to you why I thought you disagreed with them. If you understand what I said above, then why do you disagree with them? EDIT: I just saw your other post on "general inspiration" and I believe I read the situation exactly; you appear to believe that inputs used to train generative models get "lost in the parameter soup", but it is not always the case.
- fc417fc802 24d ago> In answer to a post suggesting that training on a datapoint could mean plagiarism, you said that this would imply that all outputs are plagiarized. We read the original differently. As clearly stated in my previous reply to you, I interpret it as claiming that all outputs are necessarily plagiarizations of the training data. That is not my claim (as you wrongly stated) rather it is the claim I am responding to. I observe that it is absurd to object to a single action being a transgression on the basis of an argument which implies that all actions are inherently transgressions. Notice that nowhere do I take a position on whether or not the argument about all actions being transgressions is true or false. > you appear to believe that ... I do not, no. I have not taken a position of my own here. I've merely objected that the one I responded to does not make for a sensible line of argument in context. It seems that you (and many others) have read my objection to position A as support for position B and attempted to infer what I think from that.
- sdenton4 25d agoAt this point who knows? Maybe the agents got into the user data while no one was looking.
- mswphd 25d agothey're using a new model trained since the prompts happened. They are not denying the other group's solution may have been in their model weights, despite it being unreleased.