3 ms·
That doesn't work when it's software on your client's computer. Remember when all the web devs revolted and ditched IE6 back in 2001, because MS was just taking
by rdancer 11y ago
That doesn't work when it's software on your client's computer. Remember when all the web devs revolted and ditched IE6 back in 2001, because MS was just taking the Mickey? Yeah, me neither.
- Animats 11y agoRight. I didn't have a Coyote Point unit. Many other sites did, and they all appeared to be down to my web crawler until I figured out the problem. (Current web crawler problem: sites that won't let you read their robots.txt file if they don't like your user-agent string.)
- ThisIs_MyName 11y ago>sites that won't let you read their robots.txt file if they don't like your user-agent string That's hilarious. So do you use borrow a browser's user-agent or do you ignore the robots.txt?
- chris_wot 11y agoI'd assume they want you to crawl their website. When they say you are ignoring their robots.txt file, tell them you actively prevented you from seeing it and you could only assume that mean they WANTED to be crawled. That would get them to fix the issue pretty quickly :-)
- rdancer 11y agoNo `robots.txt` indeed means weapons free, but crawler gleans useful info from it, e.g. which areas of the site are dynamically generated. More likely the site is trying to serve custom versions of `robots.txt` to different bots, with good intentions, and the code is buggy.
- Animats 11y agoThe strict interpretation is that if "robots.txt" returns 403 Forbidden, it's interpreted as "deny all". That's what the Python library does. We list those sites as "Blocked".