4 ms·
You may allow Google to scrape your website but Google will not allow you to scrape theirs.
by __b__ 10y ago
You may allow Google to scrape your website but Google will not allow you to scrape theirs.
- LoSboccacc 10y agoyou know about robots.txt right?
- __b__ 10y agoSure. You know about reciprocity?
- tunap 10y agoYeah, it's a request like "Do not track" with no controls to enforce it besides pulling a plug out. edit: To _b_'s point: "Info Sharing" is merely the promise of the marketeers', info cloistering/monetizing/weaponizing(?) are the realities.
- tedunangst 10y agoAre you claiming that Google ignores robots.txt?
- __b__ 10y agoThought experiment: You have a very large site with large amount of content. You allow Google to crawl your website and store a copy of all your content in its cache. This amounts to hundreds of pages. That is, many pages of results in the Google search engine. Imagine for the thought experiment you later lose your website and all the content. Can you "crawl Google", e.g., site:website.com, to get it back? No. You will draw a ban. Google is in some ways like the Internet Archive but with some very peculiar rules. The Internet Archive encourages users to crawl it. Is crawling a special privilege that is only reserved for search engines catering to advertisers?
- LoSboccacc 10y agogoogle is not your backup service, news at 11
- __b__ 10y agoThe issue is reciprocity as to crawling. There will never be "news at 11" about that.
- tedunangst 10y agoSo ban the googlebot in robots.txt and tell google you'll let them back in when they let you crawl their site.
- __b__ 10y agoOr allow Googlebot in robots.txt and then crawl Google.
- tunap 10y agoNo, I only find examples of Yahoo & Bing being caught at it. Google hasn't been caught, and they promise not to do evil.