4 ms·
>it's childish of you to blame Google+ for it< It is childish to run a search company & create platforms for stealing content without checking for stolen conte
by aaronwall 15y ago
>it's childish of you to blame Google+ for it<
It is childish to run a search company & create platforms for stealing content without checking for stolen content during the ranking process.
All the hard work on scraper sites doesn't apply so long as the content is hosted on Google.com???
As a search company your primary job is to deliver relevant search results to user. But as a secondary goal (every bit as important as wrapping everything in ads is) you need to ensure that the right people who are putting in human and financial capital get rewarded for their efforts. If you don't then the ecosystem you create becomes a ghetto.
Sites like Mahalo requiring the Panda update were funded by Google. Sure Google may have (eventually) solved that problem, but they created it too.
And in terms of "Google as scraper site" this isn't Google's first time around with this "accidental problem" either, as I distinctly remember them doing the same thing with Google Knol.
http://www.seobook.com/google-knol http://www.seobook.com/google-knol
Worse yet, when they did it with Knol they even had a check for how related the article was to other content that they posted right on the article, but still chose to outrank the original source.
http://www.seobook.com/images/knol-similar.png http://www.seobook.com/images/knol-similar.png
And I publicly posted about the above Google+ scraper site issue last September.
http://www.seobook.com/the-doors http://www.seobook.com/the-doors
I know Google engineers read my blog & read that post, so PR spin that tries to blame another party is simply unacceptable over 3 months later.
- VikingCoder 15y ago> checking for stolen content during the ranking process Google should check for stolen content during the ranking process, huh? That's a cool idea. How exactly would they do that? They find 5 web pages that all have the same content. How would they figure out who actually wrote it? Perhaps if they had some form of trusted identity service, where you knew what a person's real name was. And if you could mark up your web page with attribution tags of some sort, so that they knew it was you who had said it. But what if multiple people all do this, with the same content? Well, then maybe they would have to invent some kind of social graph, so that they could see which sources you trust most. Is it Linus Torvalds or Deeke McSlayton that you think is the most likely author of this post on the Linux Kernel? And maybe when you search, they should rank results by how close someone is to you in your social graph. Oh wait... That's what Google+ already is.
- aaronwall 15y ago>> checking for stolen content during the ranking process<< >Google should check for stolen content during the ranking process, huh? That's a cool idea. How exactly would they do that? They find 5 web pages that all have the same content. How would they figure out who actually wrote it?< Duplicate content filtering is nothing new. Scraper sites (mostly Google funded) have been around even before Google+ started behaving like a glorified Mahalo-esq scraper site. Document first crawl date is a great data point. And if a document links to another document with nearly identical content on it, then the document that is being referenced should probably rank higher, especially if the document was created earlier. >Perhaps if they had some form of trusted identity service, where you knew what a person's real name was. And if you could mark up your web page with attribution tags of some sort, so that they knew it was you who had said it.< You don't need real names to allow the link graph to work. If the proposed service was NOT SELF-SERVING GARBAGE it could use the signal from that network to act as a signal to help rank the rest of the web, rather than sucking content into that network and OUTRANKING THE ACTUAL LEGITIMATE ORIGINAL CONTENT SOURCE. >But what if multiple people all do this, with the same content? Well, then maybe they would have to invent some kind of social graph, so that they could see which sources you trust most. Is it Linus Torvalds or Deeke McSlayton that you think is the most likely author of this post on the Linux Kernel?< And, once again, if the service wasn't SELF-SERVING GARBAGE it would create a signal that would apply to the rest of the web so that the ACTUAL LEGITIMATE ORIGINAL CONTENT SOURCE ranked before the house-hosted copy of it. If I pull a chunk of content from a document and post it to Google+ it is pretty easy for you to put that citation in the SERPs (sorta like what is done with the Google+ votes on AdWords, http://explicitly.me/how-to-get-a-celebrity-to-endorse-all-your-products-on-google http://explicitly.me/how-to-get-a-celebrity-to-endorse-all-y... but maybe in a less sleazy way). >And maybe when you search, they should rank results by how close someone is to you in your social graph< I have seen Google+ scraper pages that were not in any of my social circles outranking original content sources. That is part of what makes Google's behavior so outrageous. >Oh wait... That's what Google+ already is.< Except it's not. Google was already one of the most heavily linked to websites before launching a social network & now Google+ is a PageRank funneling scheme http://www.thegooglecache.com/white-hat-seo/how-to-buy-1178857-links-the-google-way/ http://www.thegooglecache.com/white-hat-seo/how-to-buy-11788... built around coming up with a cheesy excuse for Google to outrank original content sources for their content. If it were legitimate then the original content source would outrank the scraped version hosted on Google+ I realize that things are "not spam" when they are done by Google, but if the mechanics & impact are the exact same as what a spam scraper site does then the quacking animal is a duck.