3 ms·
>> checking for stolen content during the ranking process<< >Google should check for stolen content during the ranking process, huh? That's a cool idea. How ex
by aaronwall 15y ago
>> checking for stolen content during the ranking process<<
>Google should check for stolen content during the ranking process, huh? That's a cool idea. How exactly would they do that? They find 5 web pages that all have the same content. How would they figure out who actually wrote it?<
Duplicate content filtering is nothing new.
Scraper sites (mostly Google funded) have been around even before Google+ started behaving like a glorified Mahalo-esq scraper site.
Document first crawl date is a great data point. And if a document links to another document with nearly identical content on it, then the document that is being referenced should probably rank higher, especially if the document was created earlier.
>Perhaps if they had some form of trusted identity service, where you knew what a person's real name was. And if you could mark up your web page with attribution tags of some sort, so that they knew it was you who had said it.<
You don't need real names to allow the link graph to work.
If the proposed service was NOT SELF-SERVING GARBAGE it could use the signal from that network to act as a signal to help rank the rest of the web, rather than sucking content into that network and OUTRANKING THE ACTUAL LEGITIMATE ORIGINAL CONTENT SOURCE.
>But what if multiple people all do this, with the same content? Well, then maybe they would have to invent some kind of social graph, so that they could see which sources you trust most. Is it Linus Torvalds or Deeke McSlayton that you think is the most likely author of this post on the Linux Kernel?<
And, once again, if the service wasn't SELF-SERVING GARBAGE it would create a signal that would apply to the rest of the web so that the ACTUAL LEGITIMATE ORIGINAL CONTENT SOURCE ranked before the house-hosted copy of it.
If I pull a chunk of content from a document and post it to Google+ it is pretty easy for you to put that citation in the SERPs (sorta like what is done with the Google+ votes on AdWords,
http://explicitly.me/how-to-get-a-celebrity-to-endorse-all-your-products-on-google http://explicitly.me/how-to-get-a-celebrity-to-endorse-all-y...
but maybe in a less sleazy way).
>And maybe when you search, they should rank results by how close someone is to you in your social graph<
I have seen Google+ scraper pages that were not in any of my social circles outranking original content sources. That is part of what makes Google's behavior so outrageous.
>Oh wait... That's what Google+ already is.<
Except it's not.
Google was already one of the most heavily linked to websites before launching a social network & now Google+ is a PageRank funneling scheme
http://www.thegooglecache.com/white-hat-seo/how-to-buy-1178857-links-the-google-way/ http://www.thegooglecache.com/white-hat-seo/how-to-buy-11788...
built around coming up with a cheesy excuse for Google to outrank original content sources for their content. If it were legitimate then the original content source would outrank the scraped version hosted on Google+
I realize that things are "not spam" when they are done by Google, but if the mechanics & impact are the exact same as what a spam scraper site does then the quacking animal is a duck.
- VikingCoder 15y agoI really wish that all technology challenges were as easy to overcome, as you think they are. That would be a beautiful world. And then it would be easy to blame companies for having products that weren't perfect. You could discount all of their efforts to try to make something new, original, and helpful - if it weren't absolutely perfect in its first iteration. Let's all switch to the search engine that is perfect, and does everything you describe. Since it's so easy to do, clearly there must be a search engine like that. Which one is it? ...or maybe, it's actually a really hard problem.
- gcarswell 15y agoBut that's exactly what everyone is saying and the OP was trying to get at. The Google Search before this WASN'T broken with this 'imperfect product' like it is now. It worked really well and that's why it became #1. You guys wouldn't have to do all this social rigging if you hadn't broken the web's tried-and-true link graph in the first place by making sites' feel like they had to 'hoard' pagerank and link-condom every outgoing link citation with a no-follow flag. Unintended consequences for sure, as no-follow was well-intentioned and designed to help with blog comment spam, but unintended consequences can be a bugger. The result is still sad. I'm sure there's no rolling back now to the 'old' google with 10 blue links, but I sure miss it.
- VikingCoder 15y agoGoogle introduced a game, and people figured out how to exploit it. Now they do everything they can to stop exploitation. It's a classic predator/prey relationship. What do you think of today's announcement? http://googleblog.blogspot.com/2012/01/search-plus-your-world.html http://googleblog.blogspot.com/2012/01/search-plus-your-worl... Specifically, "unpersonalized results"?
- aaronwall 15y agoI am with gcarswell on this one. +1 to her! I don't think search is an easy problem. And (where it is not conflicted by self promotional interests) Google does an amazing job of indexing, filtering & scoring. It is nearly unbelievable how good Google is in many areas. But that is also the problem...Google set the bar for itself rather high through its own performance. A gold medal sprinter doesn't get a pat on the back for running a 23 second 100-meter dash. When most companies put out a product they have to make it better than existing products to win marketshare. They can't put out something sort of average and then just arbitrarily promote it in the results through bundling (the way Google has with things like Checkout, places, product search, flight search, and +) In the past Google did a much better job at search when it wasn't actively subverting itself. If I searched on Google 3 or 5 years ago I didn't see Avis-rent-a-car ranking near the top of the search results for "Las Vegas hotels" ... it was only after Google decided to displace & monetize the organic results that Avis started ranking on hotel searches. http://www.seobook.com/images/hotel-price-ads.jpg http://www.seobook.com/images/hotel-price-ads.jpg And who's fault is that? Google's. If the monetary incentive wasn't there, I am sure Google wouldn't be polluting their own search results, especially as they have put thousands of man-years in trying to do the opposite. ~~~ I also think the suggestion that a person should switch their default search engine is a bit inauthentic for 2 reasons: 1.) Google is spending BILLIONS of Dollars per year buying search & browser distribution + buying default search rights on 3rd party browsers 2.) even if I change my default search engine that personal choice wouldn't impact how other users search & for those other users the copy of copyright content hosted on Google.com will often still outrank the original source.