3 ms·
I am not a lawyer, so the following is only my opinion. > It certainly prevents reusing the work in that specific manner without authorisation. Well, technica
by usrbinbash 4y ago
I am not a lawyer, so the following is only my opinion.
> It certainly prevents reusing the work in that specific manner without authorisation.
Well, technically speaking, it doesn't prevent it. It just messes up the results of the work being used in that manner. And what if someone builds a training workflow that can just ingest such changed images and use them without being negatively affected by the changes?
> so it's a bit up to interpretation.
Interpretation that would likely have to be decided in court, and likely in a very drawn out and very very very expensive manner, with uncertain outcome.
A machine-readable tagging that simply says "noone is allowed to use this for training AI" sounds way easier to argue in court to me.
- kamray23 4y ago> It just messes up the results of the work being used in that manner. And what if someone builds a training workflow that can just ingest such changed images and use them without being negatively affected by the changes? That is what encryption does as well. You can certainly attempt to watch DVDs without permission. It won't be very enjoyable. And what if someone builds a viewing application which can just watch those DVDs anyway? You see, if this is legally protected, building that workflow is circumvention and very, very illegal. Defining "effectively restricts" is left intentionally up to interpretation because there is no clear line between messing up the result and preventing access.
- usrbinbash 4y agoDisclaimer (again): I am not a lawyer, so this is only my opinion. Encryption prevents usage of the data. This doesn't. The data can still be viewed without any special software, device, password, key, etc. A sufficiently robust ingestion engine could still use it for training. In fact, even an unprepared engine can train on it, it only messes up the outcome. Honest question: Is it harder to argue legally that I DRM-protected my work if I publish them in a form that needs an encryption key/software/device, or if I publish them for everyone to see after changing some pixels around? > if this is legally protected "If" is the important term here. It was mentioned above that "maybe" this counts in the same way as a Copyright protection measure. I don't argue against that. Maybe it does. That is for lawmakers, courts and similar legal experts to decide. My opinion as someone who isn't a lawyer, is that it would be EASIER to get courts to agree on that machine-readable tagging, simply disallowing usage of works for training, is similar to DRM measures, and ignoring them should be punished the same way as circumventing copyright mechanisms. The added bonus for users: Such tagging is easier to implement, easier to update, there exists prior law already covering it (see several european countries) providing legal guidelines. Plus, artists wouldn't have to mangle their works to implement them, and it is useable with all forms of data, not just images.
- kamray23 4y ago> it would be EASIER to get courts to agree on that machine-readable tagging, simply disallowing usage of works for training, is similar to DRM measures It should be. It'd be really great if it was. Sadly, lawyers wrote the DMCA. It has to be a measure which actually restricts access in an effective enough way, just saying "don't touch this" isn't a measure because it doesn't effectively prevent the usage. 17 USC § 1201(a)(3) and all that. If training on image sets isn't copyright infringement, "don't use this" doesn't count. It's a license, and you're not infringing on it. If it is copyright infringement, "don't use this" is the default and you can't use anything without explicit permission, effectively requiring datasets to only include CC0 images. Since the first one is way more likely to be true, you instead use the 17 USC 1201(a) which prevents circumvention of technical protection measures, by creating what is hopefully a technical protection measure. Is it foolproof? It was never meant to be. It's an attempt at best. But it's better than relying on a law which is incredibly likely to never apply to dataset scraping. Preventing the usage of the data for a specific application vs total restriction is the big issue here. Is it enough to qualify as a technical protection measure? Maybe, maybe not. The courts may agree, or they may not. But it's fairly established in conversation that dataset scraping isn't infringement, so tagging it doesn't really work.
- usrbinbash 4y ago> But it's fairly established in conversation that dataset scraping isn't infringement, so tagging it doesn't really work. Conversations aside, laws can adapt to include terms regarding tagging. As said before, there are legal examples for this, eg. in the EU: https://discoverdigitallaw.com/is-web-scraping-legal-short-guide-on-scraping-under-the-eu-jurisdiction/ https://discoverdigitallaw.com/is-web-scraping-legal-short-g... Quote: the new law, that must be applied by all EU countries until 7 June 2021 (Directive (EU) 2019/790 on copyright and related rights in the Digital Single Market or ‘DSM Directive’), in its Article 4 provides an exception from the rights of the database owner mentioned above in case of ‘reproductions and extractions of lawfully accessible works and other subject matter for the purposes of text and data mining’ unless ‘the use of works and other subject matter referred to in that paragraph has not been expressly reserved by their rightholders in an appropriate manner, such as machine-readable means in the case of content made publicly available online’. End Quote. Again, I'm not a lawer, but to me that seems like it's up to lawmakers to do their homework, and update existing laws to deal with the reality that a) data mining exists and is useful for lots of things b) people want to make their works available publicly, and therefore ... c) people publishing works need a workable, stable and reliable way to tell others whether they are okay with their work being scraped and used for analysis/training/etc. or not And as I said above, ideally such a solution doesn't require changing the published data in some way, and works for all kinds of data.