8 ms·
This is missing the largest argument in my opinion. The weights are the derivative work of the GPL licensed code and should therefore be released under the GPL.
by carom 2y ago
This is missing the largest argument in my opinion. The weights are the derivative work of the GPL licensed code and should therefore be released under the GPL. I would say these companies release their weights or simply not train on copyleft code.
It is truly amazing how many people will shill for these massive corporations that claim they love open source or that their AI is open while they profit off of the violation of licenses and contribute very little back.
- wakawaka28 2y agoIt would make no sense to release the weights under the GPL because machine-generated stuff is uncopyrightable. There should be an argument about the model generating derivative works without attribution as a consequence of how it works. But that machine-generated stuff is also uncopyrightable, even though it might be kept secret.
- carom 2y agoWhat about compiler outputs? Those were initially not copyrightable then it was legislated that they were. So there is some precedent there and I would not be surprised if we saw copyrightable weights in the future (as a "compilation" of the dataset).
- wakawaka28 2y agoIt could be legislated of course. But the difference is pretty drastic. Almost nobody is creating binaries without a compiler. It is a mechanical process, but essentially everyone uses the same mechanical processes to generate binaries. I haven't looked at this issue in a while but I think compiled binaries are treated in a way similar to that of recorded music. For example, the particular bit patterns from a synthesizer might be generated from sheet music, and that is akin to code vs. binaries. But the bit patterns are copyrightable only so far as they are equivalent to or the direct manifestation of a creative work. There are other problems with releasing model weights under the GPL. It just doesn't fit, in the same way as releasing non-software under the GPL doesn't make sense. Calling the output of generative AI copyrightable violates the spirit of copyright, as it is neither creative nor labor-intensive. We could quibble about that, but I think we can at least agree that the point is that this generative AI stuff requires very little skill to use in most cases and can't operate without prior art to train on. Other lame stuff has been copyrighted before, like paint splatters and stuff, but even that type of art appears to involve more skill than entering a few words into a generative AI.
- klaustopher 2y agoJust FYI, Felix Reda was a member of the European Parliament and was responsible there for the copyright reform and also involved in GDPR, massively stepping on the feet of big tech. Don't know if it was your intention to include them in a list of people wo "shill" for big tech, but they shouldn't be included. edit wording about the shill
- blackoil 2y agoSince weights are not distributed only used by Github to provide the service, they need not worry about GPL atleast. I don't know about AGPL.
- bayindirh 2y agoWhat about the emitted code which is actually derived from GPL code? What about BSL, SSPL, or other source available (for your eyes only) licenses? Copilot harvests all public repos, regardless of its license.
- bluesign 2y agoIANAL but searched a lot on this, this is very tricky subject legally. To simplify: - imagine all code Copilot trained on is GPL licensed. - we have a universal function `isInfringing(code)` that has access to all GPL code, and returns `true` if it is infringing some GPL code. for a given prompt; if `isInfringing(copilot(prompt))==false` we cannot claim copilot infringing on GPL code, even it is trained on GPLed code. so the problem starts here; does the piece of code copilot emits, if written by yourself also would be infringing ?
- luqtas 2y ago> so the problem starts here; does the piece of code copilot emits, if written by yourself also would be infringing ? why everyone on discussions tries to bring "if a human made it"? a generative AI operates way faster than anyone ever existed and ever will and probably a person aware of the license & acting respectful towards it, will create something more sensible/plausible to avoid plagiarism now having dozen/hundreds/thousands of humans substituted by a machine that makes money for some for-profit company is really fair? even if they were a non-profit, as someone pointed up, people who create the content that feeds the weights aren't recieving a penny! they already made money with it, they will make more & that is/will upgrading/e the state of gen. AI for sure legal battles on people copying code from permissive licenses should exist but it's feels a different discussion
- wseqyrku 2y agoConsider it's already paid back because it's cheap. The price is only the service free.
- pornel 2y agoGPL doesn't apply/doesn't have to be agreed to when the usage is allowed by the copyright law in another way. GPL can't override copyright exceptions like fair use (details vary by jurisdiction, but the principle is the same everywhere). Even the license itself states it's optional, and you don't have to agree it (if you don't, you get the copyright law's default). Author of the article is a former member of the Pirate Party and EU parliament, so they have expertise in the copyright law.
- haksz 2y agoI would say that the Pirate Party has expertise in nothing apart from perhaps protecting Internet freedoms. So the same persons that supported Napster and the Pirate Bay now want to circumvent copyright for open source software. An unholy alliance, but the recent comments from some Microsoft brass about everything on the Web being freeware seems to indicate that these are the talking points that Microsoft and its new allies will put out.
- koolala 2y agoIf the Web is freeware... I wonder what options remain for licenced online information.
- janosdebugs 2y agoContent gating behind login screens. Scraping content behind a login screen could constitute a contract violation and would give rise to a lawsuit independent of copyright.
- pornel 2y agoIn this article, Reda explains the current copyright laws in the EU, not a hypothetical policy of the Pirate Party. They're not a member of the PP any more AFAIK. I expect that people professionally dedicated to a copyright reform are very familiar with it, regardless of which way they want to reform it. The copyright laws were written before generative AI existed, so they may not be adequate or fair in the new reality, but that's the current state anyway. As Reda notes, the law is not specific enough to draw the difference between collecting and processing data for search engines (that may be using ML for retrieval) and using the same data with LLMs.
- denton-scratch 2y agoBut they train their models on everything, regardless of the licence. It follows that the resulting derivative work likely mixes stuff that is under incompatible licences, with the result that it can't be distributed at all.
- dist-epoch 2y ago> The weights are the derivative work of the GPL licensed code EU courts disagree: > Under European copyright law, scraping GPL-licensed code, or any other copyrighted work, is legal, regardless of the licence used.
- JW_00000 2y ago> The weights are the derivative work of the [GPL licensed] code This is not immediately obvious to me. A small though experiment: the Harry Potter books are clearly copyrighted works. If I generate a frequency list of all words in these books, i.e. a list of all words and how often they appear, that frequency list is derived from the original work, in the normal way we would use the word "derived". But is it a "derivative work", under the strict legal definition of this term?
- carom 2y agoThe frequency count is not a function. The trained model is. Arguably, they at deriving a new function from ones covered by copyright. It is up to the courts for an official decision though.
- tpmoney 2y agoSo what if we made a function. What if someone scans all the works of Harry Potter and generates a program/function that uses the frequency and pairing of phonemes in Harry Potter character names to create a “Wizard Name Generator” to generate random but plausible sounding names. Would we expect a court to find the name generator is infringing on JK Rowling’s copyrights? Certainly it’s possible for the generator to generate a name verbatim from the books, but does that make the generator a derived work and infringing? If the authors of the generator put their generator on the web as Harry Potter Name Generator, we might expect the courts to tell them they can’t use the Harry Potter name, but if they put it under “Wacky Warlocks Wizard Wonder Namer” is the mere fact that the underlying function uses factual data about a work under copyright sufficient to strike it down? What if it used name frequencies from multiple fantasy series? How many series would it have to use as a source before we say that the name generator is not infringing on copyrights? Can it ever not be?
- gus_massa 2y agoWhat about N-grams frecuencies? 1-grams (aka characters) have too few information and are probably fine, using them you can only identify the language of the original work. With a few more you can identify the author and the book. I don't remember the exact number, but if you have the frecuencies of 10-grams you can probably reconstruct big chuncks of the book.
- qjakdx 2y agoHe works for GitHub and has probably never written anything in his life: https://okfn.de/en/vorstand/ https://okfn.de/en/vorstand/
- deleted 2y ago[deleted]
- prosim 2y agoHe started at GitHub 3 years after the article was written. Don't think GitHub's interview process takes that long. ;)
- qakjfa 2y agoThat is a good point for this individual article! However, the broader issue that Microsoft has infiltrated OSS and its organizations successfully by hiring and donating remains. It would not surprise me at all if they now hire people with an ostensibly "freedom fighter" background for credibility. Look at how many people here cite his (former?) membership in the Pirate Party for credibility! Party membership means nothing. Politicians (in general!) change their minds, can be bought, etc. The Green Party in Germany started out as a peace party and has been used repeatedly to lend credibility to the Kossovo and other wars.
- alyma 2y agoGitHub was just the logical progression from okfn: https://blog.okfn.org/2022/03/03/microsoft-to-support-open-data-day-2022/ https://blog.okfn.org/2022/03/03/microsoft-to-support-open-d... Today, we are pleased to announce that Microsoft will once again be supporting Open Data Day by providing mini-grants to organisations to help them run events, the call will launch on Open Data Day 2022. They also supported "Open Data Day 2021". Sounds like a nice trojan horse to influence EU legislation through purported activists.
- deleted 2y ago[deleted]
- ElectricSpoon 2y agoI'm with you on that. Many argue that AI models don't "contain the code" but if they are trained on the copyrighted data, and generate something similar, then the AI model is akin to a lossy data compression format. Frequency signal data over an image are not the image, but no one argues a JPEG encoded copy of a PNG isn't the same image. I think the weights vs code are similar in that regard. As for releasing weights, probably more if we're talking about AGPL code.
- batch12 2y agoI think it's amazing that licenses are ignored to train a model, but companies then try to impose a license on the use of the same model. It would be nice if there there was a training BOM that came with a model. And if not included, all rights to control the use of a model were forfeit.
- dwaite 2y ago> I think it's amazing that licenses are ignored to train a model, but companies then try to impose a license on the use of the same model. There's existing analogies like encyclopedias and dictionaries. One interesting aspect to those sorts of consolidation works is that they may contain errors and other artifacts, specifically to identify duplications of their work vs new from-scratch work.
- batch12 2y agoI don't think those are good analogies. An encyclopedia contains references or summaries of a concept or idea, but not a compressed volume of all possible text. A closer analogy would be an unauthorized "collected works" of your favorite HN commenter packaged up and resold. It also feels similar to the recent article posted about photography and how during its early days pictures were used for advertising without the consent of those photographed. [0] [0] https://www.truthdig.com/articles/the-troubled-development-of-mass-exposure/ https://www.truthdig.com/articles/the-troubled-development-o...