4 ms·
> Google ignores websites privacy. Even if the website owner doesn't want to expose any content for search engines, Google will come, get it, and show it on its
by behindai 7y ago
> Google ignores websites privacy. Even if the website owner doesn't want to expose any content for search engines, Google will come, get it, and show it on its search results and will make money off it.
Absolute, unadulterated nonsense. The Google spider identifies itself via HTTP headers. It's trivial to refuse service to specific clients.
Google spider ignores rules in robots.txt and show results from closed directories. I could write a distinct article about how we fight against it.
- kazinator 7y agoIf that is so why does the indexer bother fetching these files at all? It's a waste of bandwidth and cycles to be doing this extra fetch of robots.txt for countless sites.
- behindai 7y agoKidding me? The average size of robots.txt is 500 bytes for one web site! Average web page size is 2Mb. So if a crawler respectfully downloads a robots.txt it could significantly save more bandwidth not downloading protected pages.