3 ms·
> Publishing text or images on the web does not make it fair game to train AI on. The “public” in “public web” means free to access; it does not mean it's free
by carlosdp 2y ago
> Publishing text or images on the web does not make it fair game to train AI on. The “public” in “public web” means free to access; it does not mean it's free to use.
US Copyright law (which assigns that very copyright) disagrees with this. "Fair use" means you can use it, in certain transformative ways, not merely "access" it.
- Findecanor 2y agoThe European Union's Artificial Intelligence Act allows copyright holders to reserve the right to "opt out", which US law still does not. I think it is still very unclear though whether this could have any effect, or if is going to be a dud. One weakness is that it does not specify how. There are a number of opting-out protocols out there, and with variety it follows of course that while some site might use some protocol/s, a robot might support another. For example I could find that Cara.app uses DeviantArt's "noai" tag only in its HTML headers, but when I read through the source of Spawning.ai's "datadiligence" library [1], I found that it checks for that tag only in HTTP headers. So even in those cases when the company scraping the web for content has intent to comply with EU law, there is still the risk that they could be violating it. [1]: https://github.com/Spawning-Inc/datadiligence/ https://github.com/Spawning-Inc/datadiligence/
- stefan_ 2y agoRead up some on that law because precedent is rather down on mass corporate use.