4 ms·
Maybe a naïve question but what prevents Knuckleheads’ from ignoring the robots.txt and crawl the side anyway? And if it's so easy to do, how does Google have a
by p-sharma 6y ago
Maybe a naïve question but what prevents Knuckleheads’ from ignoring the robots.txt and crawl the side anyway? And if it's so easy to do, how does Google have a monopoly on crawling then?
- judge2020 6y agoIt's just rude to do so, and there are some technical issues with doing that as well (such as crawling admin panel which might trigger backend alarms/security alerts). Google also doesn't have a legal monopoly on crawling, only a natural monopoly thanks to a lot of websites independently choose to only allow Google and Bing because of the many issues with third-party crawlers (eg. crawling all pages at once, costing money/slowing down the site[0]). 0: https://news.ycombinator.com/item?id=26593722 https://news.ycombinator.com/item?id=26593722
- foobar33333 6y agoOn smaller sites, nothing usually. But on bigger sites you will be blocked. You will probably be blocked even if you do follow robots.txt
- sn_master 6y agoTheir IP range can be blocked. Google has known IP ranges and if the website admin wants can allow them based on that rather than robots.txt, and that would really mess up everyone else.