5 ms·
Why is it unestablished? Is a document not copyrightable based on its contents? Weights are just a different kind of a document.
by throwaway98721 3y ago
Why is it unestablished? Is a document not copyrightable based on its contents? Weights are just a different kind of a document.
- Conscat 3y agoNo, a document's contents aren't inherently copyrightable. They have to be a creative work or a method of production, and part of that basically means it has to be human generated content (as opposed to computer or animal generated). AI weights might be considered a method of production, but that isn't clear yet.
- throwaway98721 3y ago[flagged]
- ketzu 3y ago> Was there no work put into their creation by someone? Putting work into something is not a sufficient cirteria for copyright. > All of it is just a stream of bytes that the computer can interpret somehow This is also not a sufficient or at all relevant cirteria for assigning copyright. Also, in the sense you presented, those files are not fundamentally different from random noise. Which is not a particularly useful reduction for this exercise.
- dragonwriter 3y ago> Was there no work put into their creation by someone? This is the “sweat of the brow” theory of copyrightability, which courts have rejected (for good reason based on the statute.) “Someone did work to enable this thing to exist” is not sufficient to make a thing copyright-protected. > There's no fundamental difference between an image, code, or weights. And neither images, code, nor weights that are mechanically produced with no creative input by a particular author are subject to copyright in their own right (depending on their relation to the source material on which the mechanical process rests, they may be covered by the copyright on the source material.) The best argument for weights being copyrightable (and it probably applies better to some models than others) is that the assembly of source material is a creative work subject to a compilers copyright, and that the model weights themselves are just a mechanical translation of that compilation subject to its copyright.
- jerf 3y agoCopyright is not for "documents", it is for works that have creativity in them. The legal bar for that level of creativity is low, so low that it is easy to come away thinking that anything that can be cast as a "document" must be copyrightable, but the bar is in fact not zero. In particular, taking other documents and shoving them through a process that generates a lot of other numbers with no human or creative interaction is definitely something I'd be concerned the courts would judge as not sufficiently creative to be copyrightable. The process itself would certainly consist of copyrightable code, but the output doesn't necessarily. This would be somewhat similar to the observation that there is no copyright to be had in a big table of files and their MD5 hashes (or other hashes), such as a Linux distro might use for integrity checking. Lots of copyright in the original file contents, copyright available on the process for producing these tables, but the tables themselves would likely be ruled not itself copyrightable as there is no creativity in that output. Note this also has absolutely nothing to do with the question of whether AI output is copyrightable, this is about the huge table of numbers that make up the neural net weights being copyrightable. (Though it would be sort of an interesting question for the legal system to grapple with as to how a non-copyrightable set of numbers could then produce something copyrightable. Call it a philosophical variation on the "copyright washing" argument; can copyright spring from a non-copyrightable source other than a human brain, thus somehow "flowing uphill"? Would a human brain be copyrightable? Stay tuned for those questions, I guess, or if not you, your grandchildren.) Per your other comments, "work" is not the bar, "creativity" is. "Size" is not the bar either. Merely being a much larger table of numbers than a list of hashes or a phone book is not the question. No human is in that table of numbers creatively saying "no, wait, this neural weight should be -1.5 instead of 2.0 to produce this creative effect". No human is even capable of working in the medium of neural net weights in a creative manner. If you want to go the "novel legal theory" route, you could play with claiming creativity in the selection of input material and claim the resulting neural weights has a copyright in compilation: https://en.wikipedia.org/wiki/Copyright_in_compilation https://en.wikipedia.org/wiki/Copyright_in_compilation That's a long way from a slam dunk though. Way out on a legal limb there. It isn't entirely clear to me what exact rights would result from such a claim either. It would be a landmark copyright court case for sure.
- AnimalMuppet 3y ago
- dragonwriter 3y agoWeights are the output of a mechanical process over the training set with no element of human authorship, just as much the output a model produces with a prompt is, which the Copyright Office has already declared outside of copyright. > Is a document not copyrightable based on its contents? Creative process is the bigger issue. > Weights are just a different kind of a document. And who sits down and writes this document of weights?
- raincole 3y ago> Is a document not copyrightable based on its contents? Yes, exactly. It's copyright 101. For example, if you write a random number generator, and print 10000 randon numbers in a document, it's not copyrightable. Even if you invented a specific random number generation algorithm, the document is still not copyrightable. Your code is copyrightable. Again it's just copyright 101. If any of above surprises you, maybe you should read a few copyright case studies.
- deleted 3y ago[deleted]