3 ms·
>sites that won't let you read their robots.txt file if they don't like your user-agent string That's hilarious. So do you use borrow a browser's user-agent or
by ThisIs_MyName 11y ago
>sites that won't let you read their robots.txt file if they don't like your user-agent string
That's hilarious. So do you use borrow a browser's user-agent or do you ignore the robots.txt?
- chris_wot 11y agoI'd assume they want you to crawl their website. When they say you are ignoring their robots.txt file, tell them you actively prevented you from seeing it and you could only assume that mean they WANTED to be crawled. That would get them to fix the issue pretty quickly :-)
- rdancer 11y agoNo `robots.txt` indeed means weapons free, but crawler gleans useful info from it, e.g. which areas of the site are dynamically generated. More likely the site is trying to serve custom versions of `robots.txt` to different bots, with good intentions, and the code is buggy.
- Animats 11y agoThe strict interpretation is that if "robots.txt" returns 403 Forbidden, it's interpreted as "deny all". That's what the Python library does. We list those sites as "Blocked".