6 ms·
Ask HN: Has anyone else noticed Stack Overflow clones in Google search results?
Has anyone else noticed Stack Overflow clones in Google search results? They come up frequently for me. I can't help but wonder who's behind these. It can't be hurting Stack Overflow's SEO.
So far I have saved 5 different domains, and it looks like 2 have vanished.
- [dead] http://www.codeitive.com/0izVUjjXVP/selective-foreign-key-usage-in-django-maybe-with-limitchoicesto-argument.html
- [dead] http://www.codedisqus.com/0QmqWVgjgg/hide-label-in-django-admin-fieldset-readonly-field.html
- http://w3facility.org/question/image-servingurl-and-google-storage-blobkey-not-working-on-development-server/
- http://goobbe.com/questions/3109325/how-can-i-disable-a-third-party-api-when-executing-django-unit-tests
- http://www.ciiycode.com/0HyN6eQxgjXP/django-admin-inline-popups
I actually wrote Stack Overflow support about this in April 2015, but so far nothing has changed. Here's the thread:
Me: "Hello,
There a lots of spam results on Google. As a web developer, I am frequently googling for how to resolve some programming issue. Often, I get these spoofing sites that link to stack overflow.
I would suggest adding a stricter robots.txt or perhaps blocking some of these bots that are scraping your site.
Here is an example.
http://www.ciiycode.com/0HyN6eQxgjXP/django-admin-inline-popups
Thank you."
Response: "Hello,
Thank you for reporting this content. I've passed the information along to the person at our company who handles such issues. It's the diligence of users like you that helps us stay valuable!
Please note, bringing these sites into compliance (or getting them to no longer serve our content) is often a long and arduous process. You may not see immediate results. However, rest assured that we're working on it.
Thank you again,
Stack Exchange Team"
Thoughts?
- nmbdesign 11y agoHave been noticing this for the last few months too, very weird.
- kjhughes 11y agoSee "A site (or scraper) is copying content from Stack Exchange. What do I do?" http://meta.stackexchange.com/q/200177/234215 http://meta.stackexchange.com/q/200177/234215
- jrockway 11y agoStack Exchange makes all the content available as a downloadable file, so they must be expecting clones. https://archive.org/details/stackexchange https://archive.org/details/stackexchange They then have a bunch of licensing terms if you reuse the content, which must be what they're referring to by "bringing these sites into compliance is often a long and arduous process".
- ivanca 11y agoThey even have a tool/site to query the data: https://data.stackexchange.com/stackoverflow/query/new https://data.stackexchange.com/stackoverflow/query/new
- plorkyeran 11y agoThey've been around for years. A while back a few were sometimes beating SO in google results, but google eventually fixed that. Everything on SO is cc-by-sa licensed, so as long as the source is attributed (and it is on some of the sites), it's legal. The motivation is simple: the clone spam sites have ads on them. They're not actively trying to attack SO; they're just leeching value.
- rchrd2 11y agoI guess it's a win-win situation. Stack Overflow's ranking goes up, because of all the sites referencing it, and the spammers make money. But as a consequence, the users have to sift through spam results.
- volaski 11y agoI don't understand why StackOverflow allows this, yeah creative commons is cool and all but it IS NOT COOL for actual users. I get so annoyed every time i search for something on Google and it leads to an its clone site. It's not like StackOverflow has better search than Google (which is ridiculous). I still have to search on Google if I want quality search result instead of StackOverflow. With more power comes more responsibility. StackOverflow's policy feels too irresponsible in this regard
- rhinoceraptor 11y agoJust append your query text with 'site:stackoverflow.com'.
- fezz 11y agoThat or use the SO docset in Dash or StackStash https://itunes.apple.com/en/app/stackstash-stackoverflow-offline/id602112351?l=en&mt=8 https://itunes.apple.com/en/app/stackstash-stackoverflow-off... https://news.ycombinator.com/item?id=7932477 https://news.ycombinator.com/item?id=7932477
- JasonPunyon 11y agoWe use that license because it protects the content from us. No matter who comes along to run Stack Overflow in the future, Stack Overflow can't do something like put up a paywall and lock it up. Someone else'll just be able to host a copy.
- msie 11y agoYes, look at what happened to imdb.
- volaski 11y agoAs an end user, I don't care about that at all. If StackOverflow does something like put up a paywall and lock it up, then it will die off and some other site will arise that will replace it, just like some other sites that came before StackOverflow which put up a paywall and faded away. Also, StackOverflow can always change the policy if they want (which probably won't happen for the reason I mentioned), so the license as an excuse doesn't really make sense to me. Especially when it comes at a cost of horrible user experience. Lastly it doesn't seem like StackOverflow is doing much to improve search on the site itself and that's what makes this even worse. I wouldn't be complaining if I could find more relevant StackOverflow results on StackOverflow than searching on Google. How is it that I can find more relevant results on a generic search engine than the site where the contents came from?
- fiatjaf 11y agoI don't understand why Google can't figure it out and remove these clones. They can do much harder things. Why couldn't them outrank sites based on equal text content or -- much better -- huge presence of ads.
- rhizome 11y agoThey have made inroads in the past, but lately copycats have been cropping up in results (iswwwup.com is one I've been seeing a lot) again. I imagine there's a ranking algo update in the future that might fine-tune this more. To be sure, it's going to be an arms race, since the only purpose of these sites is adsense siphoning.
- TeMPOraL 11y agoI suppose it's mostly a social/legal problem. Google could get rid of 90% of those scrappers by basically hardcoding a preference for StackOverflow if it has the same content than another site. Same with sites like Wikipedia. But then obviously people will call foul play (probably scrappers and SEO people will be the loudest to complain). So whatever solution Google makes has to be general enough people won't call it unfair, and that makes the problem much more difficult.
- Animats 11y agoGoogle doesn't have a good way to establish provenance, and has trouble distinguishing copies from originals. It's a common complaint of blog operators that some bigger blog copied their stuff and got a higher ranking on Google. Google could check when it saw something, but that won't work against fast scrapers. For that, you need trusted timestamps. One solution to this would be to have a few time-stamping services. You send in a string, probably a hash, and it adds a timestamp, signs it, and sends back a signed result. Then provide a WordPress plug-in to use this service, hashing and time-stamping each blog entry, and putting the result in the HTML in some standard way. (Perhaps <span signed-provenance-timestamp-hash="xxxxx"> blog entry </span>). A few mutually mistrustful services for that would help; blogs with serious forgery problems could use multiple time-stamping services. Search engines then need to look at timestamps as a rating indicator. If two results are very similar, the earliest one wins.
- Osmium 11y agoSince establishing provenance is such a big problem for Google, perhaps it might be a good idea for Google to offer a time-stamping service itself?
- dsjoerg 11y agoIt's not directly a problem for Google — it's primarily a problem for sites that create original content.
- ezequiel-garzon 11y agoNot quite. The quality as a search engine could improve dramatically by letting true authors let Google know the content is coming before it's been available anywhere on the web.
- what_ever 11y agoWell, I think Google would prefer to send the traffic to StackOverflow instead of it's clone as it's better source?
- 11y ago
- onion2k 11y agoIt's not just StackEnchange. I noticed the other day there's a Twitter account and website called "@explodingAds" that tweets HN user comments and mirrors them on it's website. I imagine taking content verbatim is just a quick way to build a corpus of search indexed pages that generate page views and ad impressions.
- wyclif 11y agoThere's also a lot of Trello clones.
- jeffmould 11y agoSort of related, but in the last year it seems to me that Google's results in general have been increasingly getting worse. Now it seems that at least two or three spam results are always present in the first page of results. Most of these results also happen to be duplicate content of bigger sites though. I have tried reporting on numerous occasions to Google, but just usual "we will investigate" response and never a word back.
- Bjartr 11y agoIn the cases where the information you're looking for doesn't need to be recent (in my case exercises for a particular sport) you can change the date range of your search to e.g. only include results from prior to 2005. Obviously this isn't an option for recent information, but I've found it useful on occasion to filter out blogspam.
- deleted 11y ago[deleted]
- dredmorbius 11y agoThis is common across numerous contexts. I've found what appear to be bots Tweeting my reddit and HN posts (I don't mind), several Reddit clones of various levels of sniffitude, Google+ content harvesting, and some Diaspora syndication (to be expected), again, of various levels of sniffitude. To the extent that this simply distributes data around, doesn't claim it for its own, and credits source, I'm OK with this. Better even if it follows site-specific licensing. Among my visions for the Web would be content syndication where such schemes would actually directly benefit authors and creators, regardless of where their content is served.
- Phil_Latio 11y agoWhen I look at w3facility.org, it seems like Googles algorithms do not properly handle the case when a scrape-site provides a source-link to the original content. Google recommends the latter to protect against duplicate content penalties when you use some external content to enrich your site (for example a short section from Wikipedia, imdb actor info, etc).
- douche 11y agoI see this happen a lot also with the MSDN forums. The interesting thing there is that some of the mirror sites are still carrying topics that have been deleted or otherwise disappeared from the real MSDN forums. More than once, in the obscure subset of Microsoft tooling that I work in, the only hits that are still alive are on somewhat suspicious looking .ru sites, so in some sense, I am glad these sites do exist - otherwise, I'd be completely SOL trying to figure out why the badly-documented API I'm relying on is barfing up an opaque HRESULT.
- aaron695 11y ago> I would suggest adding a stricter robots.txt or perhaps blocking some of these bots that are scraping your site. The only people who care about robots.txt are some of the big companies. Even Baidu ignores it (As they can, it's purely there as etiquette) Blocking bots is hard.
- inguinalhernia 11y agoScraper spam exists for virtually every website that is even somewhat popular, there is almost nothing site owners can do about it but rely on Google, Bing, DuckDuckGo, etc, to figure out the originator. Google should be able to figure it out and rank the true origin source appropriately, and usually Google is pretty effective at that. But sometimes they're not. Perhaps what I find ironic is that many of the user "answers" on StackOverflow are basically clone spam themselves, copy/pasted from other websites by some user of the site, usually without sourcing the origin. I have personally found my own unique solutions and code copied verbatim and pasted to answers on the StackExchange network multiple times, outranking my original work without a reference to the origination of course, and I'm sure others have experienced something similar. Perhaps that's what you get with a user generated site, maybe Wikipedia experiences something similar. Related, some of you may recall a few years back, that StackOverflow basically complained to Google in a public fashion about not ranking well enough and got a boost from them, whereas obviously the average Joe and an average website has no such option nor recourse. Here was the discussion on HN: https://news.ycombinator.com/item?id=2152286 https://news.ycombinator.com/item?id=2152286
- ConceptJunkie 11y agoI haven't seen this yet, and I do searches pretty frequently that result in StackOverflow.com hits... and I make a point of choosing them first, because they are usually the best results. However, a few days ago, I did see a single hit for expertsexchange.com for first time in years, at least that I noticed. Back when Google used to let you blacklist domains from your search results, before StackExchange was ever around, or when it was still very new, I used to block them, and eventually I'd assumed they'd gone away, but maybe not. I hope StackOverflow and/or Google are able to do something to put the kibbosh on this kind of thing, because despite the complains SO is still a huge and valuable resource.