3 ms·
Algorithms should not be subject to copyright IMO. If you publish code, AFAIK under EU law only blatant copies are an infringement. In case of art, copyright is
by chromanoid 2y ago
Algorithms should not be subject to copyright IMO. If you publish code, AFAIK under EU law only blatant copies are an infringement. In case of art, copyright is enforceable anyway, at least not in any different form than before AI.
Yes, current AIs just remix their training data in a rather direct way. But in the end how different is that to how humans create? I would suggest we should embrace this new way of creating things while finding laws to empower all creators and not only those who were hired by deep pockets.
- martin-t 2y ago> AFAIK under EU law only blatant copies are an infringement Laws generally don't encode what is right but a compromise between the state's interests, lobbyists and the general population making enough ruckus if too unsatisfied. > But in the end how different is that to how humans create? 1) Scale. Some strategies that are socially acceptable when done by individuals but not when done at a massive scale. For example because individuals have very limited time and can invest very limited effort. Looking at a website is perfectly OK. Making thousands of requests a second might be considered an attack. Human memory is limited. Similar principles apply to humans looking at code. 2) Source of data. Much of human "input" is viewing the real world (not copyrighted material) through their senses. Much of learning is from teachers or documentation, both of which voluntarily give me information. I don't know about you but when I wanna know how to use a particular function, I don't go looking through random GH repos to see how other people use it, I go to the docs. > finding laws to empower all creators and not only those who were hired by deep pockets That is not even the only issue. When I publish something under AGPL, my users have the right to modify the code, even if my code gets to them through some third party. LLMs allow laundering code and taking that right from (my) users.
- chromanoid 2y ago> > AFAIK under EU law only blatant copies are an infringement > Laws generally don't encode what is right but a compromise between the state's interests, lobbyists and the general population making enough ruckus if too unsatisfied. of course, but I actually think that this is the correct moral stand. Patenting algorithms is like patenting thoughts. > I don't know about you but when I wanna know how to use a particular function, I don't go looking through random GH repos to see how other people use it, I go to the docs. I always look into sources. Usually I look into the code I want to call first. But this is probably also because I mainly use Java. So if I read your AGPL code and implement something similar in another programming language after also reading other implementations of the algorithm, is that something I have to attribute you for? Isn't code just an executable documentation of an algorithm? Especially if the algorithm is well known, I don't see any injustice here - "Die Gedanken sind frei". If I copy your code via copy and paste, then this is an infringement, but just retelling a similar story should not be affected by copyright.
- martin-t 2y ago> Patenting algorithms is like patenting thoughts. OK, I agree there, I should have written "function" or "module" something similar. Something that takes nontrivial amounts of work and although it is based on some general principles which should not be patentable/copyrightable, their particular implementation is novel/unique enough that it would take nontrivial amounts of work to replicate the functionality without seeing the original. > is that something I have to attribute you for Depends how closely you follow my implementation. If you use my code as the only reference and translate it verbatim (whether manually or using a tool), then you should credit me. If you look at many implementations, form an _understanding_ of the algorithm in general, then write your own implementation based on that understanding, then probably not. The question is where LLMs stand. They mix enough sources that crediting all of them would be impractical and in practise they end up crediting none. But their proponents (who always call them AI, sometimes even using pronouns like "he" to refer to the models) argue that the models also form an _understanding_ rather than just regurgitating a mix of inputs. And I have to disagree, what I see is an imitation of reasoning/understanding which is sometimes convincing due to how complex statistics are being used inside the models. But they are still just statistical models of existing content and we see that every time somebody releases a new model, HN upvotes it to the top and a few hours later we inevitable see people giving it trivial questions which it fails to answer correctly. My other two points: - Even if an ML company made a model that is actually intelligent, the burden of proof should be on them, otherwise or until them, it's just a remix of existing work. BTW this reminds me an interesting comparison is remixes vs cover songs in music. - Code is famously harder to read than write. If a human takes time to understand a piece of code and reimplement it not verbatim, then he generally does not get ahead by much. An LLM can do this at scale and speed unattainable by humans. Let's say two products compete (purely on features and quality instead of marketing - for the sake of argument). One is written first, is novel and written fully by humans. The other is written by training a model on the first product's code and using the model to generate the same product, all within hours or days instead of months or years. The other puts in less actual work but gets the same result. It is clearly parasiting on the first, benefiting from their work without giving them credit or compensation. --- Bottom line is copyright is meant to protect authors who invest effort into creating. Whether it succeeds in that can sometimes be questionable. But using an algorithm (even a very complex one) to take a bit of everyone's work and redistribute is for free without crediting or compensating them does not benefit authors. I hate analogies but if I write banking software and send 0.000000001% of every transaction to my account, none of the individuals thusly affected probably care that much but I am still going to prison.