13 ms·
In general: (1) training ML systems on public data is fair use (2) the output belongs to the operator, just like with a compiler. On the training question spec
by natfriedman 5y ago
In general: (1) training ML systems on public data is fair use (2) the output belongs to the operator, just like with a compiler.
On the training question specifically, you can find OpenAI's position, as submitted to the USPTO here: https://www.uspto.gov/sites/default/files/documents/OpenAI_RFC-84-FR-58141.pdf https://www.uspto.gov/sites/default/files/documents/OpenAI_R...
We expect that IP and AI will be an interesting policy discussion around the world in the coming years, and we're eager to participate!
- croes 5y agoFair use doesn't exist in every country, so it's US only?
- krzyk 5y agoIt exists in EU also (and it much mire powerful here).
- croes 5y agoThe EU doesn't have a copyright related fair use. Quite the opposite, that why we are getting upload filters.
- jay_kyburz 5y agoYes, my partner likes to remind me we don't have it here in Australia. You could never write a search engine here. You can't write code that scrapes websites.
- king_magic 5y ago@Nat, these questions (all of them, not just the 2 you answered) are critical for anyone who is considering using this system. Please answer them? I for one wouldn't touch this with a 10000' pole until I know the answers to these (very reasonable) questions.
- joepie91_ 5y ago> training ML systems on public data is fair use Uh, I very much doubt that. Is there any actual precedent on this? > We expect that IP and AI will be an interesting policy discussion around the world in the coming years, and we're eager to participate! But apparently not eager enough to have this discussion with the community before deciding to train your proprietary for-profit system on billions of lines of code that undoubtedly are not all under CC0 or similar no-attribution-required licenses. I don't see attribution anywhere. To me, this just looks like yet another case of appropriating the public commons.
- patrickthebold 5y agoWhat does "public" mean? Do you mean "public domain", or something else?
- 6gvONxR4sf7o 5y agoUnfortunately, in ML "public data" typically means available to the public. Even if it's pirated, like much of the data available in the Books3 dataset, which is a big part of some other very prominent datasets.
- kzrdude 5y agoSo basically youtube all over again? I.e bootstrap and become popular by using widely available whatever media (pirated by crowdsourced piracy) and then many years later, when it gets popular, dominant, it has to turn around and "do things right" and guard copyrights.
- deleted 5y ago[deleted]
- qihqi 5y ago(1) training ML systems on public data is fair use This one is tricky considering that kNN is also a ML system.
- visarga 5y agokNN needs to hold on to a complete copy of the dataset itself unlike a neural net where it's all mangled.
- stwrong 5y agoWhat about privacy. Does the AI send code to GitHub? This reminds me of Kite
- avery42 5y agoYes, under "How does GitHub Copilot work?": > [...] The GitHub Copilot editor extension sends your comments and code to the GitHub Copilot service, which then uses OpenAI Codex to synthesize and suggest individual lines and whole functions.
- stefano 5y agoHow do you guarantee it doesn't copy a GPL-ed function line-by-line?
- king_magic 5y agoI truly don't think they can guarantee that. Which is a massive concern.
- ipsum2 5y agoYup, this isn't a theoretical concern, but a major practical one. GPT models are known for memorizing their training data: https://towardsdatascience.com/openai-gpt-leaking-your-data-2dfb9e22c1b2 https://towardsdatascience.com/openai-gpt-leaking-your-data-... Edit: Github mentions the issue here: https://docs.github.com/en/github/copilot/research-recitation https://docs.github.com/en/github/copilot/research-recitatio... and here: https://copilot.github.com/#faq-does-github-copilot-recite-code-from-the-training-set https://copilot.github.com/#faq-does-github-copilot-recite-c... though they neatly ignore the issue of licensing :)
- visarga 5y ago> GPT models are known for memorizing their training data Hash each function, store the hashes as a blacklist. Then you can ask the model to regenerate the function until it is copyright safe.
- ipsum2 5y agoWhat if it copies only a few lines, but not an entire function? Or the function name is different, but the code inside is the same?
- proteal 5y agoIf we could answer those questions definitively, we could also put lawyers out of a job. There’s always going to be a legal gray area around situations like this.
- stephen82 5y ago> We expect that IP and AI will be an interesting policy discussion around the world in the coming years, and we're eager to participate! Another question is this: let's hypothesize I work solo on a project; I have decided to enable Copilot and have reached a 50%-50% development with it after a period of time. One day the "hit by a bus" factor takes place; who owns the project after this incident?
- lovich 5y agoYour estate? The compiler comparison upthread seems to be perfectly valid. If you work on a solo project in c# and die, Microsoft doesn’t automatically own your project because you used visual studio to produce it
- breck 5y agoYou should look into: https://breckyunits.com/the-intellectual-freedom-amendment.html https://breckyunits.com/the-intellectual-freedom-amendment.h... Great achievements like this only hammer home the point more about how illogical copyright and patent laws are. Ideas are always shared creations, by definition. If you have an “original idea”, all you really have is noise! If your idea means anything to anyone, then by definition it is built on other ideas, it is a shared creation. We need to ditch the term “IP”, it’s a lie. Hopefully we can do that before it’s too late.
- anmk 5y agoI'm sure natfriedman will be thrilled to abolish IP and also apply this to the Windows source code. We can expect it on GitHub any minute!
- breck 5y agoI used to work at Microsoft and occasionally would email Satya the same idealistic pitch. I know they have to be more conservative , but some of us have to envision where the math can take us and shout out loud about it, and hope they steer well. When I started at MS, my first week was heckled for installing Ubuntu on my windows machine. When I left, windows was shipping with Ubuntu. What may seem impossible today can become real if enough people push the ball forward together. I even hold out hope that someday BG will see the truth and help reduce the ovarian lottery by legalizing intellectual freedom.
- cnp10 5y agoTalking about the ovarian lottery seems strange in a thread about an AI tool that will turn into a paid service. No one will see the light at Microsoft. The "open" source babble is marketing and recruiting oriented, and some OSS projects infiltrated by Microsoft suffer and stagnate.
- sombremesa 5y agoAll I know is that if a lawsuit comes around for a company who tried to use this, Github et al won't accept an ounce of liability.
- deleted 5y ago[deleted]
- tlamponi 5y ago> the output belongs to the operator, just like with a compiler. No it really is not that easy, as with compilers it depends on who owned the source and which license(s) they applied on it. Or would you say I can compile the Linux kernel and the output belongs to me, as compiler operator, and I can do whatever I want with it without worrying about the GPL at all?
- user-the-name 5y ago> training ML systems on public data is fair use So, to be clear, I am allowed to take leaked Windows source code and train an ML model on it?
- dylannorthrup 5y agoOr, take leaked Windows source code, run it through a compiler, and own it!
- abn120 5y ago(1) That may be so, but you are not training the models on public data like sports results. You are training it on copyright protected creations of humans that often took years to write. So your point (1) is a distraction, and quite an offensive one to thousands of open source developers, who trusted GitHub with their creations.
- dylannorthrup 5y agoFair Use is an affirmative defense (i.e. you must be sued and go to court to use it; once you're there, the judge/jury will determine if it applies). But taking in code with any sort of restrictive license (even if it's just attribution) and creating a model using it is definitely creating a derivative work. You should remember, this is why nobody at Ximian was able to look at the (openly viewable, but restrictively licensed) .NET code. Looking at the four factors for fair use looks like Copilot will have these issues: - The model developed will be for a proprietary, commercial product - Even if it's a small part of the model, the all training data for that model are fully incorporated into the model - There is a substantial likelihood of money loss ("I can just use Copilot to recreate what a top tier programmer could generate; why should I pay them?") I have no doubt that Microsoft has enough lawyers to keep any litigation tied up for years, if not decades. But your contention that this is "okay because it's fair use" based on a position paper by an organization supported by your employer... I find that reasoning dubious at best.
- deleted 5y ago[deleted]
- deepnash 5y agoIt is the end of copyright then. NNs are great at memorizing text. So I just train a large NN to memorize a repository and the code it outputs during "inferencing" is fair use ? You can get past GPL, LGPL and other licenses this way. Microsoft can finally copy the linux kernel and get around GPL :-).