4 ms·
> Why wouldn't this be the case? One possible legal theory: because the algorithm was trained on a text corpus upon which the algorithm's owner has no legal cl
by throwawayjava 7y ago
> Why wouldn't this be the case?
One possible legal theory: because the algorithm was trained on a text corpus upon which the algorithm's owner has no legal claim.
In this particular case, I don't think that theory would hold much water.
However, consider, e.g., a model that produces encyclopedia entries and is trained on a half dozen existing encyclopedias. IMO, if that model is using techniques similar to SoTA and isn't producing utter garbage, then the owner of that model should have a very difficult time claiming that the output of their model is anything more than a sophisticated round-about way of copy/pasting from existing encyclopedias.
But still, in that case, the output is still covered by copyright. It's just that the owner of the training set -- not the owner of the algorithm -- is the one with the valid claim to copyright.
- mamon 7y ago>> One possible legal theory: because the algorithm was trained on a text corpus upon which the algorithm's owner has no legal claim. The same can be said about human writers: they learn to write based on thousands of "training examples" - the articles and books they read thorough their life.
- beerandt 7y agoA better example is the "copy and paste" news articles that saturate feeds everyday. The exact same set of facts, that were obviously reported originally by a single individual, then rearranged, reworded, and republished by 100's of "reporters"/"bloggers", (sometimes) with an attribute of origin.
- throwawayjava 7y agoNot at all. Or rather, Who knows? Maybe. But certainly, at least today, a SoTA model generating a quality encyclopedia certainly is not doing what human writers do, and is certainly effectively copy/pasting. Maybe in 50 years -- or 10 years with a major breakthrough on the level of general relativity -- that statement might be true. but it's certainly not true of today's deep NLP systems.
- micimize 7y agoThat would be a problem, but would be a data licensing issue, which is distinct. It's more analogous to "Blurred Lines" infringing on "Got To Give It Up" or w/e.