3 ms·
> Not all LLMs use crawlers and identify themselves Yeah, exactly why I want to block them. > This is not a sustainable solution as the number of crawlers con
by PreInternet01 3y ago
> Not all LLMs use crawlers and identify themselves
Yeah, exactly why I want to block them.
> This is not a sustainable solution as the number of crawlers continues to grow
Well, I want zero crawlers to access my content, so that seems pretty sustainable?
> An ‘all or nothing’ approach is unacceptable
To whom? For me 'nothing' works? I do not want your crawler to access my content for any reason whatsoever. That's not too hard, is it? I block your UA, I block your IP range, done, right?
> Robots.txt is all about managing crawling while the copyright discussion is all about how the data is used
Potato, potato. Don't crawl me, don't use my data for AI, or anything. Does that affect my search engine ranking? Don't care.
> Reinventing the wheel
> The meta tag is the way
> Foolproof solution
Ah, OK, so anything but your totally imagined META solution won't work. Good luck with that!
- ke88y 3y agoThe article is making the point that you need to use copyright -- not robots.txt -- to enforce restrictions on use of your content. I think that's probably correct -- robots.txt is a request. Violating that request is probably not a violation of the CFAA and certainly doesn't entail extra copyright protection. Gripe all you want and black hole IP ranges all you want; your content will get crawled. TBH, I'm not sure there is a way to enforce this preference on the open internet if crawlers are willing to violate robots.txt. (There are, of course, practical ways that mostly involve not putting your work on the open internet; eg, a custom sign-up flow with all content behind a login page will do the trick for any site that isn't sufficiently popular.) So, per the article, use copyright instead. And if you do not want your content used to train an AI, then you need to find a way to clearly communicate to the bots that your work is under copyright; otherwise, they'll use your content for training and you'll have to sue them post-facto. Which is all good and well, but again, the premise is that you don't want your content used for training in the first place! Which brings us to the operative question: is it true that CC BY-ND and CC BY-NC-ND prevent the use of data by LLMs? In fact, is a licenses which explicitly disallows use of content in LLM training is effective? I'm not sure that this is clear yet; e.g., if it turns out that using content in training data turns out to be cut-and-dry fair use then do restrictive licenses have any practical effect?
- WorldMaker 3y agoIf someone is crawling content in any way and aren't respecting robots.txt, that's a violation of the intent of robots.txt. It may be a "friendly request", but it's never the requestor doing something wrong if the requestee simply (badly) ignores the signs marked "beware of dog" or "no robots allowed". (Really the big thing missing here with robots.txt is that the internet doesn't have guard dogs to try to put some teeth painfully into your skin if you violate private property. Lack of guard dogs to enforce them isn't lack of signs and some of these companies need to get over themselves that they think they are better than posted signs.) Plus, if those signs aren't enough, websites are published materials and all content under current US laws and court precedents is considered that published materials fall under copyright unless noticed otherwise. We've moved to a "prove that there is no applicable copyright" system, not a "prove that there is" system. It's not up to me to make sure that my copyright statements are clear and obvious enough, it's up to you to prove that you find a public domain or license statement of some kind (such as CC licenses including CC0; CC0 exists because there needed to be an explicit declaration of "no really, this is public domain" under current copyright regimes). This kind of crawling seems completely oblivious to how current copyright law works and seems to me to be wantonly asking for more lawsuits (which again, given current evidence, crawlers would lose). > is it true that CC BY-ND and CC BY-NC-ND prevent the use of data by LLMs? Yes. Easily. LLMs are a derivative process, LLMs fail all the requirements for CC ND clauses. LLMs are also commercial processes and being used for commercial interests and just as easily fail CC NC clauses. I would even go far as to say that LLMs are bad at attribution, often lie about it, especially in situations of near direct quotes, and easily fail most reasonable interpretations of the CC BY clause. It seems cut-and-dry to me. I don't think there is a lot of gray area.
- ke88y 3y agoYour first two paragraphs are your opinion. Unless you're a SCOTUS judge or powerful Senator, I'm not sure that your opinion matters much. And I do agree with your opinion, btw, so you're in good company, but I think the courts mostly vehemently disagree with this position... Your third through fifth paragraphs misunderstand the central question. I can say whatever I want in a license. Whether or not those clauses are enforceable is a very different question. If OpenAI and GitHub say that using open source code/content is fair use, then the terms of the copyright license aren't necessarily relevant. E.g., if I say that my website can only be used for non-commercial purposes and then run for office and NYT publishes a copy of a two page racist screed on my website, then I can try to sue then for violating CC-by-... But it won't matter and I will lose. The terms of the license are not relevant since NYT clearly has a very strong fair use case. That's the whole point of fair use -- that the content can be used in ways that the rights holder does not consent to!