3 ms·
That's a really weird argument. Copyright is a legal system created by the Constitution and statutes and administrative rules. It cares about whether you are
by bjt 3y ago
That's a really weird argument. Copyright is a legal system created by the Constitution and statutes and administrative rules. It cares about whether you are "copying", and it cares about whether you're creating things that compete with the works of the original authors. It doesn't care about potential output spaces.
In this context, I don't see a principled difference between the model weights and really good compression. If I send you a gzipped copy of the latest bestseller book it's still copyright infringement. And it would still be infringement if I shipped it inside a software program that can _also_ reshuffle the words in a bajillion different ways, if there's a "copy" of the original work in there.
- scheeseman486 3y agoYou can eke out large chunks of works from Google Books with the right queries too.
- kmeisthax 3y agoYes, but Google Books is a search engine. It doesn't write books, it just tells you where a particular phrase might occur in those books. There's explicit caselaw allowing you to do this, extending back before Google was even a thing. For related reasons, Google Books also does not let you read the whole book - just the page the search match came from. OpenAI and other large language model developers are claiming they have a machine that can write books, but they also fed it shittons of books, and they can't account for where all that text went. At best they can say "well, it doesn't produce exact, verbatim copies of the training set all the time".
- scheeseman486 3y agoGenerative AI doesn't write books either, it's a machine spitting out results based on queries entered by a person. It has no personhood nor rights to ownership, in spite of what AI more crackpot advocates try to claim, it's more a very complicated pencil.