4 ms·
How would anyone block a crawler? A crawler is just a headless browser.
by creese 6y ago
How would anyone block a crawler? A crawler is just a headless browser.
- tleb_ 6y agorobots.txt https://www.robotstxt.org/ https://www.robotstxt.org/ https://en.wikipedia.org/wiki/Robots_exclusion_standard https://en.wikipedia.org/wiki/Robots_exclusion_standard
- Xylakant 6y agoNote that robots.txt is a hint to well-behaved crawlers, not blocking them in any regard. You can block crawlers if you can identify them, but reliably identifying them is hard.
- tleb_ 6y agoWe should probably classify the crawler identifying problem as impossible and move along. Less resources wasted and easier automation for everyone. Assuming a crawler is malicious is narrow-minded.
- ddorian43 6y agohttps://developers.google.com/search/docs/advanced/verifying-googlebot https://developers.google.com/search/docs/advanced/verifying...
- Xylakant 6y agoThis helps to verify that a bot that announces itself as google bot is indeed a google bot. It doesn’t help identify a bot that pretends to be a user/browser.