7 ms·
Even MIT licensed code requires you to preserve the copyright and permission notice. If a human did what these language models are doing (output derivative wor
by jedbrown 3y ago
Even MIT licensed code requires you to preserve the copyright and permission notice.
If a human did what these language models are doing (output derivative works with the copyright and license stripped), it would be a license violation. When humans want to create a new implementation with clean IP, they have one team study the IP-encumbered code and write a spec, then a different team writes a new implementation according to the spec. LM developers could have similar practices, with separately-trained components that create an auditable intermediate representation and independently create new code based on that representation. The tech isn't up to that task and the LM authors think they're going to get away with laundering what would be plagiarism if a human did it.
- cj 3y agoHas anyone been able to create a prompt that GPT4 replies to with copyrighted content (or content extremely similar to the original content)? I'm curious how easy or difficult it is to get GPT to spit out content (code or text) that could be considered obvious infringement. Tempted to give it half of some closed-source or restrictive licensed code to see if it auto-completes the other half in a manner that is obviously recreating the original work.
- littlestymaar 3y agoI don't know about GPT-4 but you could get ChatGPT to spit Carmac's Fast Inverse square root with the comments and all (I can't find the tweet though…) Edit: it wasn't ChatGPT but Copilot see https://twitter.com/mitsuhiko/status/1410886329924194309 https://twitter.com/mitsuhiko/status/1410886329924194309
- jameshart 3y agoI can reproduce when prompted all the lyrics to Bohemian Rhapsody, but my doing so isn’t automatically copyright infringement. It would depend on where, when, how, in front of what audience, and to what purpose I was reciting them as to whether it was irrelevant to copyright law, protected under some copyright use case, civilly infringing, or criminally infringing copyright abuse. The same applies to GPT. It could reproduce Bohemian Rhapsody lyrics in the course of answering questions and there’s no automatic breach of copyright that’s taking place. It’s okay for GPT to know how a well known song goes. If copilot ‘knows how some code goes’ and is able to complete it, how is that any different?
- throwaway38249 3y agoOK, it can exist without breaking any laws, but if you can't release anything it helps you write, what's the point?
- jameshart 3y agoThere’s a clear separation between the training process which looks at code and outputs nothing but weights, and the generation process which takes in weights and prompts and produces code. The weights are an intermediate representation that contains nothing resembling the original code.
- vkou 3y agoThe neurons in my brain when I plagiarize are just arrangements of atoms that contain nothing that resembles orginal code/text passages/etc.
- jameshart 3y agoThe trained weights of a GPT model are a frozen, static, transmissible representation. They’re not equivalent to the live state of a brain.
- szundi 3y agoPretty equivalent to the snapshot of a live brain. Those inside it are even called neurons and neural network
- jameshart 3y agoNo, they are the weights that are used to configure a neural network. They’re a map of how to build a useful brain, not a neural state.
- __loam 3y agoMachine learning neural networks have almost nothing to do with how brains work besides a tenuous mathematical relation that was conceived in the 1950s.
- visarga 3y agoYou can say that if you want to nitpick, but there are recent studies showing that neural and brain representations align rather well, to the point that we can predict what someone is seeing from brain waves, or generate the image with stable diffusion. https://sites.google.com/view/stablediffusion-with-brain/ https://sites.google.com/view/stablediffusion-with-brain/ I think brain to neural net alignment is justified by the fact that both are the result of the same language evolutionary process. We're not all that different from AIs, we just have better tools and environments, and evolutionary adaptation for some tasks. Language is an evolutionary system, ideas are self replicators, they evolve parallel to humans. We depend on the accumulation of ideas, starting from scratch would be hard even for humans. A human alone with no language resources of any kind would be worse than a primitive. The real source of intelligence is the language data from which both humans and AIs learn, model architecture is not very important. Two different people, with different neural wiring in the brain, or two different models, like GPT and T5 can learn the same task given the training set. What matters is the training data. It should be credited with the skills we and AIs obtain. Most of us live our whole lives at this level and never come up with an original idea, we're applying language to tasks like GPT.
- Maxion 3y ago> When humans want to create a new implementation with clean IP, they have one team study the IP-encumbered code and write a spec, then a different team writes a new implementation according to the spec. Maybe at a FAANG or some other MegaCorp, but most companies around barely have a single dev team at all, or if they're larger barely have one per project.
- visarga 3y agoWhy can't AI do the same: copyrighted code -> spec -> generated code. ... and then execute copyrighted code -> trace resulting values -> tests for new code. AI could do clean room reimplementation of any code to beef up the training set. It can also make sure the new code is different from the old code at ngram-level, so even by chance it should not look the same. Would that hold up in court? Is it copyright laundering?
- koolba 3y agoIsn’t the language model itself the spec? Potentially for all of the inputs at once.
- jedbrown 3y agoLanguage models don't understand anything, they just manipulate tokens. It is a much harder task to write a spec (that humans and courts can review if needed to determine is not infringement) and (with a separately trained tool) implement the spec. The tech just isn't ready and it's not clear that language models will ever get there. What language models could do easily is to obfuscate better so the license violation is harder to prove. That's behavior laundering -- no amount of human obfuscation (e.g., synonym substitution, renaming variables, swapping out control structures) can turn a plagiarized work into one that isn't. If we (via regulators and courts) let the Altmans of the world pull their stunt, they're going to end up with a government-protected monopoly on plagiarism-laundering.