5 ms·
This article is really thin and disappointing. It's perfectly reasonable for the FTC to be involved here, and what constitutes unjust enrichment and anti-compe
by kj4ips 3y ago
This article is really thin and disappointing.
It's perfectly reasonable for the FTC to be involved here, and what constitutes unjust enrichment and anti-competitive practice when covered works are used as components in a training set is within their purview.
The USPTO will be involved in issuance questions, and the legal system will be involved in determining single cases.
On this issue, the following is my opinion (not a lawyer, not legal adv.).
A training set is a collection of representative input, which may or may not be covered works.
The training set is a simple aggregation, it does not alter the works contained in any significant way, so it is not transformative.
Therefore, the training set is as an omnibus of other works.
The weights of an system are a product of the input training set (if DL is used), and the parameters of the training environment. 1 is an omnibus of potentially covered works, the other is likely non-covered (but maybe patent-eligible in some cases). It is unclear if this is transformative, but it is certainly not covered by any of the existing exemptions (such as covers of songs).
Therefore, it is possible that the product of the above is a non-transformative derivative work based on one or more covered works, and therefore would have required a license or grant from the author(s) of those covered works. Note that authors often retain their ownership of UGC, and grant a service provider the ability to make that available to others in limited capacity for specific purposes. Some service providers (Like stack overflow) require that contributors license contributions under a specific license, and yet others transfer ownership as part of the terms of service.
If the training set was built from only entirely public domain works, then there wouldn't be a question, but the extension of copyright term over the years has robbed the public domain of anything even remotely useful for these kinds of cases.
On a side note, StackExchange contributions are CC BY-SA, which means that if the training input included anything from stackoverflow, the weights could be be affected by the Share-Alike requirement, which could lead to some interesting results if someone ever tried to request them. (However the attribution would probably be pretty easy, because you could pretty much say "all stackoverflow contributors" and be correct)
- TimPC 3y agoI think under any reasonable definition of transformative it's impossible for the weights themselves to not be transformative. The weights of a neural network are substantially different in type and nature from the omnibus of works that generate them. The more pertinent question is whether the outputs generated by the weights are all transformative. The weights themselves being transformative is not sufficient to make an entity hosting an AI model not responsible for outputs that are not sufficiently transformative and I think most AI models we've seen to date tend to expose direct copies or near copies of their input data in part or all of a response.
- freejazz 3y agoBeing transformative is also not, on its own, sufficient to be a fair use. It's just one, albeit important, factor in the analysis. If they are transformative, that also means the weights are inherently a derivative work.
- hn_acker 3y agoIs there a term similar to "transformative" but instead for describing results created from but not remotely similar to existing copyrighted works? A term meaning "created from works but doesn't copy expression"? For example, suppose I take the upper left pixel from a hundred thousand images, then arrange the collected pixels into a new image (using an unspecified combination of manual and/or automatic methods). It's not out of the question that model weights would be <that hypothetical term for "created from works but doesn't copy expression">.
- freejazz 3y ago>For example, suppose I take the upper left pixel from a hundred thousand images, then arrange the collected pixels into a new image (using an unspecified combination of manual and/or automatic methods). That's "de minimis" copying. >It's not out of the question that model weights would be <that hypothetical term for "created from works but doesn't copy expression">. Model weights would be a derivative work that might be a fair use because of its transformative nature. The issue with derivative works isn't necessarily the extent of the copying so much as that it seems logical that we'd let the creator of Seinfeld have control over whether or not a Seinfeld movie exists. I feel like many people here struggle with copyright because they try to treat it like a math problem, but it's not a math problem. it's a political and social problem that we try to resolve by applying principles and reason (like much of the law). so if you aren't applying reason and principles, but instead treating it like a computer system and trying to blow it up by stuffing the inputs, etc, you aren't really getting it and you aren't engaging with the actual law, which doesn't work that way and doesn't suffer the problems that computer systems do, in part, due to their lack of ability to reason. > "created from works but doesn't copy expression" This is what I'm getting at. Copyright only applies to expressions. So you can't be a copy of something without having copied its expression. For the same reason you can't be a derivative of another work, without having copied some of that expression. To put it another way, if you have a copyrighted work, and you subtract out the expression - there is nothing left of the copyrighted work. There might be things left, but they were never what was captured by the copyright, because copyright only pertains to expressions fixed in a tangible medium.
- lmeyerov 3y agoThe acceptance of search engines set a lot of precedence that make any blanket reasoning harder to justify - the underlying technologies and use cases are often similar or even the same. Google has been using both extractive and generative summaries in their results for a long time, and even had 'cached' mode.