3 ms·
Technically correct (the best kind of correct), but if I set a thousand users on to a website to each download a single page and then feed the information they
by wulfstan 1y ago
Technically correct (the best kind of correct), but if I set a thousand users on to a website to each download a single page and then feed the information they retrieve from that one page into my AI model, then are those thousand users not performing the same function as a crawler, even though they are (technically) not one?
If it looks like a duck, quacks like a duck and surfs a website like a duck, then perhaps we should just consider it a duck...
Edit: I should also add that it does matter what you do with it afterwards, because it's not content that belongs to you, it belongs to someone else. The law in most jurisdictions quite rightly restricts what you can do with content you've come across. For personal, relatively ephemeral use, or fair quoting for news etc. - all good. For feeding to your AI - not all good.
- JimDabell 1y ago> if I set a thousand users on to a website to each download a single page and then feed the information they retrieve from that one page into my AI model, then are those thousand users not performing the same function as a crawler, even though they are (technically) not one? No. robots.txt is designed to stop recursive fetching. It is not designed to stop AI companies from getting your content. Devising scenarios in which AI companies get your content without recursively fetching it is irrelevant to robots.txt because robots.txt is about recursively fetching. If you try to use robots.txt to stop AI companies from accessing your content, then you will be disappointed because robots.txt is not designed to do that. It’s using the wrong tool for the job.
- catlifeonmars 1y agoI don’t disagree with you about robots.txt… however, what _is_ the right tool for the job?
- hundchenkatze 1y agoauth, If you don't want content to be publicly accessible, don't make it public.