4 ms·
What happens if other sites are scraping content faster than Google can crawl it? In these cases, will Google really be able to guess which site is the origina
by aamar 16y ago
What happens if other sites are scraping content faster than Google can crawl it? In these cases, will Google really be able to guess which site is the original? For all they know, SO is scraping a lot of its content from other sites.
If this kind of uncertain-originator is any part of the problem, one solution might be for Jeff to temporarily block robots other than google/bing/etc. from retrieving new content, until say, ten minutes later. This gives the search engine a chance to figure out who the original is, while still (I think) remaining within the spirit of CC-SA. A Google API call (I'm high reputation, please crawl this new page now!) might be even better.
edit: clarified API suggestion.
- 6ren 16y agoThat's assuming the scrapers will respect robots.txt. Some of them would, and that would help. Assuming google does take into account who was first, a similar solution is for Jeff to submit his content directly to google for indexing, immediately it's published. EDIT I'm wrong; google only accepts top-level URLs for indexing, not new content: http://www.google.com/addurl/?continue=/addurl http://www.google.com/addurl/?continue=/addurl
- aamar 16y agoWell, you could ban misbehaving robots by IP address. That seems like a reasonable step for bad behavior and not in conflict with Jeff's general "sharing"/"openness" values.
- rabidsnail 16y agoYou can specify specific urls by submitting a sitemap (http://www.google.com/support/webmasters/bin/answer.py?hl=en&answer=156184 http://www.google.com/support/webmasters/bin/answer.py?hl=en...).
- mootothemax 16y agoThe only way of solving this I can think of is sending new URL notifications to Google - if not sending the entire HTML content as well. Without the content it'd be open to abuse, and I can't any way to scale it - but then I'm not a Google engineer ;)
- aamar 16y agoThat seems good -- but I can imagine Google preferring to crawl the page, rather than receive it by API (so it's more likely to be what the user's going to see). I think Google can handle the scaling problem; one not-great solution: ignore notifications except for those from people who are being scraped and need it. Still, it's kind of a shame that webmasters have to worry about any of this.
- mootothemax 16y agoThat seems good -- but I can imagine Google preferring to crawl the page, rather than receive it by API I'd agree with you, but can think of edge cases where naughty sites A, B and C submit new URLs within an arbitrary amount of time of the new content being published. In that case there'd be no way for Google to tell who published the content first other than to have a big list of original content publishers - and I think that list'd get messy fast.
- aamar 16y agoYou're right, there's an exploit there for the re-publishers. Great point.
- trezor 16y agoIf you, as a webmaster or developer, needs to inform search engines when you your content has changed, you just became part of the search-engine. If that is what is required, and everyone on the internet needs to build google-hooks into their server-side website logic to not be SEO'd to death, I think its safe to say that google just stopped being useful. I hope that in 2011 people will start seeing beyond the Google RDF which has gotten increasingly annoying throughout the last year: We need more competition in this space. Pronto. Making your website part of the google search-engine will not accelerate this. It's not a "good" solution. It's barely a solution at all.
- mootothemax 16y agoIt's not a "good" solution. It's barely a solution at all. I couldn't agree with you more; it's a stupid problem to have and one that shouldn't exist. It seems like the time is ripe for either a Google-killer to emerge in the next couple of years, or for Google to improve themselves in the same amount of time. Personally I believe that decent competition would be the healthiest way, but let's not kid ourselves; as soon as one search engine starts being used by a decent proportion of users, the SEO blackhats are going to try and game it just as much as they do with Google.