8 ms·
Who/what is the: 1. Author 2. Copyright holder 3. Copyright license ...of the code generated by this tool? Unless the answer to all of these is unambiguous
by _odey 3y ago
Who/what is the:
1. Author
2. Copyright holder
3. Copyright license
...of the code generated by this tool?
Unless the answer to all of these is unambiguously "the original", then you shouldn't be using such tool on any code, especially on your employer's intelectual property.
Sorry to be so negative about it but this is something that I see skipped over in all discussions related to AI. Just because it's AI does not make it immune to copyright law. You're giving away your code to a 3rd party company under their terms and conditions, and receivig some new code back, again under their terms and conditions. The fact that it uses AI under the hood is irrelevant, you're dealing with a business that produces you an output and you should know the terms before submitting anything to them, especially if you don't own that thing.
- transitivebs 3y agoGreat feedback. I'm not the author of this project, but in my understanding, it's the same as if you were to write the code yourself. The project doesn't publish anything and it works entirely locally aside from LLM calls (which could in the future be 100% local as well). So you remain the author and have complete control over the license of the generated code.
- jamil7 3y ago> and it works entirely locally aside from LLM calls So not entirely locally. Yes these could eventually also run locally but OP’s point still stands.
- _odey 3y agoThis is great to hear but I don't fully know how far reaching are openai's claws (since it does require an openai api key). If it runs 100% locally then yes, it would be safe to use.
- transitivebs 3y agototally agreed that diff projects need to be careful w/ sharing proprietary code w/ third parties. openai's official stance is that it will never use API calls as training data, and that in my understanding it may retain API call data for up to 30 days for compliance purposes, but that it legally won't store it beyond that (whereas chatgpt convos are meant to be stored and used for training purposes). as a next step, they could provide a swappable version of the LLM provider using something like https://github.com/imartinez/privateGPT https://github.com/imartinez/privateGPT, https://github.com/alexanderatallah/window.ai https://github.com/alexanderatallah/window.ai, etc. would love to have a standard develop here as the community matures around LLM usage
- Towaway69 3y agoHow does running it locally make a difference to copyright? If I run an OCR software 100% locally, do I get the copyright on the scanned result of Harry Potter? Don't understand: using a giant LLM locally negates all the copyrights contained in the LLM? In which country is that a law?
- largbae 3y agoI think GP means that at least you aren't dealing with a SaaS middleman that _also_ may intercept your code or claim authorship of the conversion.
- 6510 3y agoYou mean ownership nor license changes.
- nbrycnb 3y agoNo, you don't "remain the author". You have never been the author if you translate the real author's code base. The license is the one of the original, since this is unambiguously a derivative work.
- rezonant 3y agoRecommended edit: IANAL I find this reasoning a bit flimsy. Copyright doesn't have anything to do with publishing, and is it really that clear that you are the sole author of the derived work?
- complex_exp 3y agoYour comment is very valid. I'd just add that AI tools are clearly taking the "YouTube approach": they provide a large value added, ignore copyright for the moment, and hope to resolve it peacefully at some later point in time. This worked very well for YouTube.
- rezonant 3y agoI wouldn't describe the Content ID regime and the myriad lawsuits and backroom deals as "peaceful".
- complex_exp 3y agoYouTube wasn't killed and thrived as a platform throughout the process. Meanwhile YT ads funded the lawsuits and negotiations, with a surplus. It is pretty much a solved problem now. This is as peaceful as it gets when you genuinely infringe on someone's very valuable rights.
- rezonant 3y agoYeah they survived but I think we're worse off in a world of Content ID, copystrikes, erosion of fair use, theft of ad revenue by game companies. The list goes on. YouTube should probably be a lesson, not a model to copy.
- nbrycnb 3y agoYes, the lesson is not to contribute anything of value to open source or another person's platform. Creators need to use restrictive licenses, then all of these parasitical corporations will cease to exist.
- muxator 3y agoLimiting the scope to software, I'd say it's fair to distinguish MIT licenses from GPL. The latter provides way more freedom for the user (as opposed to corporations willing to profit without giving back). I am fair more comfortable contributing to an AGPLv3 software as opposed to a MIT licensed one. I can't talk about licensing for content creators (like youtube), because I do not have much experience about it.
- AbrahamParangi 3y agoYou’re making a few claims implicitly here: - the copyright status of the output of LLMs is ambiguous - this ambiguity represents a legal risk to the users of LLMs - given that ambiguity, nobody should use LLMs for work that is intended to be copyrighted It’s not clear to me if any of these are true individually, never mind all together. Maybe the copyright status is ambiguous but I think the probability that the output of LLMs is owned by anyone other than API caller is very low. You can copyright works, but you can’t copyright the ideas within those works. Possibly this represents a liability issue but I think the probability of that is vanishingly small. Just because a legal theory exists doesn’t mean it’s going to be enforceable - one of the reasons for the existence of Uber, youtube, etc. if it’s a fait-accompli it’s not a risk. Finally, it representing a risk of copyright liability is still just a business decision. I would guess that existing software projects are riddled with code snippets of ambiguous or incompatible licenses. All it takes is one person to copy a function that is GPL and label it otherwise and then it’s out in the world, technically running amok. It is probable that under the strictest definitions, every software company is engaging in some low level of copyright infringement. All-in, this means that while it’s a risk, it’s a risk you’re probably already taking and so you might reasonably conclude to ignore.
- hobs 3y agoAt least the first two are true if you are a lawyer, as every lawyer I know is waiting for AI legislation to make the two first points clear. Any lawyer will tell you if the first two points are unclear the third point is rock solid - don't use tools that have ambiguous copyright terms until AFTER the big legal fallout/legislation unless you are willing to bet the entire farm.
- sigotirandolas 3y agoFrom my non-lawyer "common sense" point of view, the reasoning that 1 and 2 implies 3 seems absurd in the face of the "no penalty without law" principle. If the law is so unclear that lawyers can't determine legality and are waiting for additional guidance from the lawmakers, shouldn't it be legal by default?
- thumbuddy 3y agoSurprise you just paid a few pennies and gave all LLM users and another company you don't know all of your IP, while exposing yourself to future litigation.
- hhh 3y agoPeople continue to parrot this point, but API use is not used to train models.[1] Don’t trust OpenAI? Use Microsoft. Don’t trust Microsoft? Run TII Falcon locally. 1: https://openai.com/policies/api-data-usage-policies https://openai.com/policies/api-data-usage-policies
- nprateem 3y agoMost companies' IP is mostly worthless. It's their brand and market position that's much more valuable. If you disagree, how many users are on your twitter clone right now? Exactly.
- shipscode 3y agoHmm not sure, I'll let you know after I use it and receive a cease and desist letter.
- ericd 3y agoWhat’s the ownership of code that’s translated into another programming language by a human, given that usually there’s not a complete 1:1 of all features, so some of it will need to be rewritten?
- oceanplexian 3y agoIt’s weird all of these fringe legal theories are coming out now in response to “AI”, which isn’t much more than a fancy search engine. No one was asking those questions when the backbone of any modern tech company was, ironically, thousands of copy/pasted lines of code from StackOverflow and any other references that come up in a Google search. The number of SDE’s I know who can write code without a web browser is vanishingly small.
- thatxliner 3y agoI don’t know what’s the actual legal status on LLMs right now, but here’s my thought: if you created artwork in Adobe Photoshop, does Adobe own your art work? If you made a world in Minecraft, does Minecraft own your world? What about using the AI tools in Adobe Photoshop? Or WorldEdit in Minecraft? I’m sure these question can be easily answered by looking it up, so why would LLMs be any different? You have control over what the LLM generates (prompts).
- AlecSchueler 3y agoThanks for saying this, this is my perspective as well and I've been kinda dumbfounded by all the questions about copyright from LLMs. They're just tools and computer programs right? Why would copyright be any different while using them than other tools. Never seen it questioned that Windsor & Newton could actually be the copyright holder for half the world's art.
- nucleardog 3y agoBecause the inputs to the models are copyrighted works? If I open up Photoshop and make a new file, I own my art. So why would me opening up an existing PSD and moving all the pieces around and then claiming it as my own be an issue? I don’t think anyone would question this at all if the model were trained only on your own data. It’s the part where a bunch of other people’s stuff was involved that makes it fuzzy enough to be an open legal question.
- AlecSchueler 3y agoI read your position as being the copyright should belong to the holders of the copyright for the training works? I can sort of see that but a) I'm also seeing it posited that the copyright could belong to either ht model or the company that produced the model (where does that logic come from?) and b) My own creativity is trained on the works of many others but my work is still my own, why is this different for an LLM?
- basch 3y agoIt’s different because rules are written and enforced by humans not robots. So it’s whatever the copyright office says at the moment. I think a same result though is that any work translated by the original author counts as a translation that maintains but doesn’t extend copyright.
- newswasboring 3y ago> Sorry to be so negative about it but this is something that I see skipped over in all discussions related to AI. I don't know which forums you are discussing these things in, but this is the first comment on all discussions about AI on HN.
- swader999 3y agoValid points, and it needs to be said. But it's a PR away from switching from an external API call to using a locally installed gippity brain to do the migration.