4 ms·
While I understand the concepts of derivatives and tainted code, this AI/human-dichotomy is not as good as that reasoning requires it to be. Every statement of
by hurril 2y ago
While I understand the concepts of derivatives and tainted code, this AI/human-dichotomy is not as good as that reasoning requires it to be. Every statement of code I commit is in fact a derivative of work that potentially had an incompatible license.
- whatevaa 2y agoYou can ingest millions of line of code in minutes/hours? If AI can ignore copyright, why can't a human? Are we already inferior to machines just because of our copyright system?
- hurril 2y agoThat is my point right there.
- zarzavat 2y agoI agree. For the people who think that every tiny piece of code has copyright, have you never written the same code for two different projects? There are some utility functions that I have written dozens of times, for different projects. Am I supposed to be varying my style each time so that each has its own unique copyright? I do not believe any court would interpret copyright to work in this way.
- chrisjj 2y ago> have you never written the same code for two different projects? But note this prohibition is of "code that was not written by yourself" rather that written by yourself twice.
- johnisgood 2y agoI am pretty sure that I have implemented a lot of functions before that would be almost (or completely) identical (not including variable names, indentation, etc.) to that of someone else's work. The smaller the function, the more likely it seems to be true, and some programming languages encourage you to write concise functions.
- zarzavat 2y agoYes my point is that the prohibition is based on a theory that tiny pieces of code have copyright and so all AI generated code is tainted. If you write code for one company/client then the IP belongs to that company as a work for hire. If you write the same function again for someone else then you would be infringing the first company's copyright - if you truly believe in this theory that little pieces of code have copyright. If you believe in this theory then you need to have a database of all the code you've ever written and be constantly checking every function you write for potential infringement of a previous employer's rights.
- paulmd 2y agoYou can’t have it both ways though, if an AI can’t hold copyright and is merely a tool being used by a human engineer who bears legal responsibility and ownership, then the concept of “code that was not written by yourself” is completely incoherent. It is impossible for code to not be written by yourself under that view because it is just a tool. The fact that it’s mad-libs autocomplete instead of normal intellij auto-complete is totally irrelevant.
- chrisjj 2y agoIt is coherent e.g. if the "AI" is used as a tool to copy another's code. > The fact that it’s mad-libs autocomplete instead of normal intellij auto-complete is totally irrelevant Completely relevent if the former had and gave no permission and the latter did.
- ncallaway 2y agoOkay, but if you’ve become very familiar with a particular code base, the later produce code that’s an exact match for a significant piece of code to the first code base, that would be copyright infringement. There’s a reason companies will use clean-room strategies to avoid poisoning a code base. So, I’d be okay with a compromise wherein the makers of Chat GPT aren’t liable for copyright for the act of training the model, but are liable for copyright for statements produced by the model. So, when the AI spits out exact copies of copyrighted works in response to a prompt, then copyright has been violated and the AI creator should be liable.
- hurril 2y agoAn exact match is of course a derivative, let's call it the identity derivative. So in: let new_code = f previous_work where f could be identity. But it could be other transformations such as: rename_all_the_things, move_items_around, object_oriented_to_functional, and what not. Probably a combination. But here's the kicker: all permutations of these functions are all derivatives. And we don't have to be naive here, thinking that I'm talking about using refactoring tools on an actual code base because I want to hide the fact that I want to steel some piece of code. I'm talking about the fact that I've seen so much code and this has in fact taught me basically everything I know. I have trained myself on a stream of works. Just like an AI. The difference is in scale, not in nature.
- chrisjj 2y ago> An exact match is of course a derivative Not if neither is derived from the other.
- hurril 2y agoSo this comes down to probabilities then. What is the probably that the fact that my TotallyNotLinux-kernel matches line-by-line with the actual Linux-kernel? I didn't copy it. Promise! But my function called println which is very very similar to some other, well it might not be a derivative.
- nicklecompte 2y agoThe are two salient differences here: 1) You are a thinking adult - or a thoughtful teenager :) - and therefore you understand copyright law well enough to take responsibility for copyright law, and in particular you are capable of having standing in legal matters. AI is not, but it is quite capable of creating infringing code (eg GNU stuff) that human reviewers wouldn't even know was infringing. So it is much better for any honest and competent organization to ban commercial LLMs entirely. (I am fine with in-house solutions with 100% validated training data...but those aren't very good yet, are they?) 2) I am a broken record on this, but the biggest problem with the "stochastic parrots" analogy is that transformer ANNs are dramatically dumber than parrots, or any other jawed vertebrate (I am not sure about lampreys). As applied to code generation: when I first tested ChatGPT-3.5, I was shocked to discover it was plagiarizing hundreds of lines of F#, verbatim, including from my own GitHub. Obviously that's outrageous in terms of OpenAI's ethics. But it is also amazing how dumb the AI is! Imagine a human programmer who is highly proficient in Python, and pretty good at Haskell, yet despite reading every public F# project in GitHub it can't solve intermediate F# problems without shameless copying. It is a completely misleading comparison to say that humans reading source code is anything like transformers learning patterns in text. The most depressing thing about the current AI bubble is watching tech folks devalue human intelligence - especially since the primary motivation is excusing the failures of a computer which is less intelligent than a single honeybee.
- ADeerAppeared 2y agoThere's a key dichotomy here: While it is technically possible for you to read some GPL'd code, memorize it, and then reproduce it later by accident. That's not how programmers work. What humans remember is not the copyrightable code, but the patentable algorithm. (And there are few algorithms that are simple enough to memorize on a cursory reading, novel enough to be patentable, but not so novel you'll remember it's source) AI does not work in algorithms. It's a language model. It deals purely in the copyrightable code. (Both figuratively; LLMs are structurally incapable of the high level abstract reasoning required, and literally by way of the training data)