11 ms·
Digg banned from Google?
- niggler 14y agoI'd wait for an official Digg statement, because Digg may have requested to be removed from google search results (the net result would be the same)
- jcampbell1 14y agoMaybe they nuked themselves: http://digg.com/robots.txt http://digg.com/robots.txt
- shanelja 14y agoI noticed that too, they are also missing a sitemap.xml
- kyrra 14y agogetting 404 on that file.
- fmavituna 14y agoI'm pretty sure getting 404 from robots.txt will not get you nuked from Google :)
- ErikAugust 14y agoYou do not need a robots.txt or a sitemap XML to be listed.
- franze 14y agocurl -I http://digg.com/robots.txt HTTP/1.1 404 Not Found Content-length: 1443 Content-Type: text/html Date: Wed, 20 Mar 2013 16:27:13 GMT Server: nginx/1.1.19 Connection: keep-alive an HTTP 404 robots.txt should not be an issue - but maybe there was a robots.txt with something else in there instead, and now it's gone. but even so, if there was an all blocking robots.txt site:digg.com would still show the URLs of digg.com, as crawling is optional for indexing. (maybe they used the non-documented Noindex: / directive, but i doubt it) if you want to achieve such a clean removal, in most cases you must request a complete removal via google webmaster tools. so yeah, somewhere, someone might have screwed up, most likely on diggs side, maybe on googles side, maybe a combination. UPDATe: now it's curl -i http://digg.com/robots.txt HTTP/1.1 200 OK Cache-Control: public Content-Type: text/plain Date: Wed, 20 Mar 2013 17:05:29 GMT Etag: "c47ccf1a49c24cc5842430aa75c72ef491292412" Last-Modified: Wed, 20 Mar 2013 16:51:48 GMT Server: TornadoServer/2.2 Content-Length: 24 Connection: keep-alive User-agent: * Disallow: which is the best practice robots.txt (allow all, or rather: hey, i have a valid robots.txt file but the second statement is not a valid robots.txt disallow directive, so this means allow all as there is no other disallow directive (note: i once coded https://npmjs.org/package/robotstxt https://npmjs.org/package/robotstxt which tried to reimplement googles spec of the robots.txt https://developers.google.com/webmasters/control-crawl-index/docs/robots_txt?hl=en https://developers.google.com/webmasters/control-crawl-index... so i have some experience in reading robots.txt files )) still, it wasn't a robots.txt issue hmmm, just a hypothesis: maybe just maybe someone thought it was a good idea to remove www.digg.com via google webmaster tools from google (as their main domain is digg.com and not www.digg.com and they definitely tried had some www.digg.com URL indexed, as even bing some some of these URLs indexed http://www.bing.com/search?q=site%3Awww.digg.com&go=&qs=ds&form=QBRE&filt=all http://www.bing.com/search?q=site%3Awww.digg.com&go=&... ) but google is (maybe) set to treat www.digg.com and digg.com as the same (via the settings panel in google webmaster tools), removing www.digg.com then could result into removing digg.com as well (something similar happened to a client of mine years ago), so could be the issue, could not be the issue, we would need more data (access to GWT) to verify this.
- ryusage 14y agoAs another data point: I just clicked that link in Chrome, and I get the robots.txt with * Disallow like others are saying. Weird that some people are getting a 404.
- highace 14y agoIt's cloaked so only google can see it - try changing your user agent to match googlebot's.
- cpeterso 14y agoUsing the googlebot User-Agent string gives me: User-agent: * Disallow: This robots.txt should allow all bots to search the entire website. However, I think Google also penalizes websites that serve different content to googlebot than to non-bot user agents.
- blauwbilgorgel 14y agoThat's weird because http://blog.digg.com/robots.txt http://blog.digg.com/robots.txt does not cloak. Why cloak one and leave the other? Sitemap: http://blog.digg.com/sitemap-pages.xml Sitemap: http://blog.digg.com/sitemap1.xml User-agent: * Disallow: /private Disallow: /random Disallow: /day Crawl-delay: 1
- TobbenTM 14y agoIt's up for me with: User-agent: * Disallow:
- RyanMcGreal 14y agoIt seems to be fixed now: User-agent: * Disallow:
- kiallmacinnes 14y agohttp://digg.com/robots.txt http://digg.com/robots.txt <-- 404 Not Found I know a robots.txt 404 shouldn't de-list you, but I would have expected digg to have one? Maybe they requested de-listing and removed their robots.txt? Who knows!
- davenseo 14y agohttp://digg.com/robots.txt http://digg.com/robots.txt looks fine too me
- randomdata 14y agoThere is one there now: User-agent: * Disallow:
- kiallmacinnes 14y agoYup - It's back now.. But it 100% didn't exist when I posted earlier :)
- MiguelHudnandez 14y agoThe article is down, but here is the text from Google's cache: --snip-- Something interesting has just come across one of my networks (hat tip to datadial), just a few days after Digg have announced that they are building a replacement for the much loved Google Reader, they have (coincidentally?) disappeared from the primary google index. [image unavailable] Is it an SEO penalty for links? That seems to be the number one reason that brands are getting booted from google’s index these days… Some conspiracy theorists will no doubt be proclaiming its something to do with their announcement to build a replica of the now defunct Google reader, but personally I really cant see that having any effect. Could there? Doing a site: search for Digg certainly demonstrates that they are no longer in the index: [image unavailable] Its likely (only if it is link based however) that it would be down to what individuals who submit content do after the fact – ie. sending spammy links at their posts to try and build the pagerank, and create “authority” which they then pass back to their own sites. Digg has long been listed in every “linkwheel” sellers handbook, and if that is the reason then what does it mean for every community site on the internet? Will we have to manually aprove all new links soon at this rate? Come on Google – WTF – let the internet know what you’re doing please. --end snip-- Without the images, maybe I am missing some context. But it seems hyperbolic, considering digg is not serving a robots.txt anymore [1]. It is probably just a blunder on Digg's part. [1]: At the time of my comment, digg was serving an xhtml document with status 404 at /robots.txt. Now it appears to be a valid robots file. PS: I am enjoying the irony of having a copy of the article, even though the site is down, because of Google's cache. Need to pontificate about how Google is potentially evil but can't keep your server running? Don't worry, people can read it via Google's cache.
- pyre 14y agoWhy don't people just post things via a Coral Cache link when it's not a major site (that should obviously be able to handle the traffic)?
- username111 14y agoBecause it isn't proper etiquette, it is their content and they should get the views for it. Some people may not be able to handle it and caches are always nice but pageviews are also a nice thing when you run a website even if they don't generate any revenue.
- Alex3917 14y agoIt's possible (though not necessarily likely) that they finally got banned for their toolbar, which is basically designed to scam Google.
- sjs382 14y agoI'm not familiar with their toolbar. How does it scam Google?
- mikeryan 14y agoAre they still using it? the current Digg is a whole new site rewritten from the ground up from betaworks I'm not seeing the toolbar at all.
- Alex3917 14y agoHmm it appears they killed it in 2010, though I still see some browser plugins available.
- uptown 14y agoWhat year are you posting from?
- qompiler 14y agoGoogle has rules concerning SEO. A warning for the rest of us not to try to cheat our way to a better ranking.
- deleted 14y ago[deleted]
- dhimes 14y agoThey are in Yahoo EDIT: I fucking hate HN's 'Delete' function.
- blauwbilgorgel 14y agoDoing a site:digg.com/news/ search on Bing shows a lot of pages like these: http://digg.com/news/gaming/ing_bank_i_ilanlar http://digg.com/news/gaming/ing_bank_i_ilanlar and even more duplicate tag and rss pages for "site:digg.com/tag/" and "site:digg.com rss". These /news/ pages 302 redirect to many different sites (some are bound to contain spam or be of lower quality). 302 redirects for these links is bad practice. Some link shorteners (ab)use 302 Found (instead of 301 Moved Permanently) to hoard content that doesn't belong to them. The content for these links can't be found on digg.com, so they too use the wrong redirect and associate themselves with all pages they link to. Besides that: Digg.com acts like a single page webapp for most of its content. There are no discussion pages or detail pages for the stories. The content that does appear is near duplicate to other content on the web, especially with popular stories, where many blogs just copy the title and the first intro paragraph.
- tracker1 14y agoIt's funny, but I've been pretty careful to add the nocache,noindex headers for many of the pages on the site I work on where a given page is no longer active... also, I would think that digg would maybe consider a rel=nofollow for any links that aren't "popular" yet. That may be the best way to handle the link spam.
- ChuckMcM 14y agoI saw a similar pattern in the links that remain, it's sad to see the link authority of a site get plundered, and it seems like something inside Google's indexer realized that was going on and deleted it. As a search engine its something you have to do if you want to consider site authority in your ranking model.
- AznHisoka 14y agoIt doesn't matter if they were banned - they didn't care about SEO in the first place. All those old URLs they had for years? All disappeared as they wanted to start fresh.
- dredge 14y agoFrom my experience, when Google de-indexes a site, they also suppress any PageRank the Google Toolbar would have shown for it; that doesn't appear to be the case here, digg.com is still PR8. Having toolbar PageRank[1] and yet no cached page[2] is not something I've seen before. [1] http://toolbarqueries.google.com/tbr?features=Rank&sourceid=navclient-ff&client=navclient-auto-ff&iqrn=UgnC&ch=8f25a5d62&q=info:http%3A%2F%2Fdigg.com http://toolbarqueries.google.com/tbr?features=Rank&sourc... [2] https://www.google.com/search?q=cache:digg.com https://www.google.com/search?q=cache:digg.com
- jasongill 14y agoYour experience may be limited, because this is very common. PageRank and indexed status operate independently of each other; it's not uncommon to see a site that was deindexed still maintain PR for quite some time (and often, indefinitely). However, if your site is deindexed, PR means "nothing" because your links no longer provide juice. Valid PR while being deindexed is one of the (many) tricks that Google has added in the last couple years to try to reduce the usefulness of getting PR for blackhat purposes
- rodion_89 14y agoTheir robots.txt file clearly asks to not be crawled at all. User-agent: * Disallow: http://digg.com/robots.txt http://digg.com/robots.txt
- blauwbilgorgel 14y agoIncorrect. To configure your robots.txt to not be crawled at all use: User-agent: * Disallow: / Allow indexing of everything with: User-agent: * Disallow: It seems they are at this very moment struggling to change things.
- rodion_89 14y agoYou are totally right. Disregard everything I said.
- chenster 14y agoDisallow nothing -> allow everything.
- anotherbadlogin 14y agoDigg deserves to be eliminated from Google. It's a pile of spam now. Pathetic.
- Karunamon 14y agoBetteridge's law of headlines strikes again.
- Matt_Cutts 14y agoThis has nothing to do with Reader. We were tackling a spammer and inadvertently took action on the root page of digg.com. Here's the official statement from Google: "We're sorry about the inconvenience this morning to people trying to search for Digg. In the process of removing a spammy submitted link on Digg.com, we inadvertently applied the webspam action to the whole site. We're correcting this, and the fix should be deployed shortly." From talking to the relevant engineer, I think digg.com should be fully back in our results within 15 minutes or so. After that, we'll be looking into what protections or process improvements would make this less likely to happen in the future. Added: I believe Digg is fully back now.
- jmount 14y agoThe Star Trek TOS "The Ultimate Computer": Captain James T. Kirk: "And how long will it be before all of us simply get in the way?"
- walshemj 14y agoSo was it the well known missing robots.txt killing a site problem (as described in the original article) or a manual action gone astray? When you looking at protections you need to ask what happens of this is a mom and pop site and not a well known sv company with high level Google contacts?
- jmount 14y agoI can't speak for Google or Matt Cutts. But it sounds like he said it was a Google bug.
- luckynic 14y agoIs this a stunt to put Digg.com on the map again?
- relix 14y agoIf this would happen to a less popular site, what chances does a site-owner have of getting attention to this problem, and getting it fixed?
- eiderv 14y agoWhat I love most about this situation is the multiple theories, when in essence Google puts it very simply: WE F*ED UP. Sorry!
- joetek 14y agoGoogle admitted the mistake: “We’re sorry about the inconvenience this morning to people trying to search for Digg. In the process of removing a spammy link on Digg.com, we inadvertently applied the webspam action to the whole site. We’re correcting this, and the fix should be deployed shortly.” http://thenextweb.com/google/2013/03/20/google-seems-to-have-de-indexed-digg/ http://thenextweb.com/google/2013/03/20/google-seems-to-have...
- nicholassmith 14y agoIf you assume conspiracy then you're most likely ignoring the simple fact of life in the technology world. Someone made a screw up.
- hugbox 14y agoAnd nothing of value was lost.