5 ms·
Hello all I work at Google as a Webmaster Trends Analyst to help webmasters with issues like this one. Looking into this, the first thing I noticed is that th
by pierrefar 15y ago
Hello all
I work at Google as a Webmaster Trends Analyst to help webmasters with issues like this one.
Looking into this, the first thing I noticed is that the blog.jquery.com seems to be blocking Googlebot from fetching its pages, but the site responds normally for web browsers: it returns an HTTP 500 error headers for requests using a Googlebot user agent. You can see this yourself using a public tool like Web Sniffer to fetch the page spoofed as Googlebot ( http://web-sniffer.net/?url=http://blog.jquery.com/2011/06/30/jquery-162-released/&uak=9 http://web-sniffer.net/?url=http://blog.jquery.com/2011/06/3... ) or using Firefox with the User Agent Switcher and Live HTTP Headers addons.
Unfortunately this is a very common problem we see. Most of the time it's a mis-configured firewall that blocks Googlebot, and sometimes it's a server-side code issue, perhaps the content management system.
Separately from that, I also notice that the blog.jquery.it URL is redirecting to the blog.jquery.com, suggesting they are fixing it on their end too.
If an jquery.com admins want more help, please post on our forums ( http://www.google.com/support/forum/p/Webmasters?hl=en http://www.google.com/support/forum/p/Webmasters?hl=en ).
Cheers,
Pierre
- yaakov34 15y agoDo you have any insight into what makes jquery.it rank so highly? I thought a pagerank of 8 was very difficult to attain. Was this done with just some SEO, or did some mega-popular site link to them by mistake? This is the real problem here, because people do have a lot of trust that very highly-ranked results on Google won't hurt them.
- deleted 15y ago[deleted]
- josegonzalez 15y agoI'd imagine that if they are a perfect duplicate of jquery.com, and jquery.com is constantly returning a 500 error page to the googlebot, then all the content on jquery.it would be "unique". There is so much of it, and all highly relevant content, that it's pretty much a no-brainer that it's pagerank would be so high. That said, I find it funny that there is so much vitriol against Google in this case when, as has been noted above, it's likely a misconfiguration on jquery.com's part.
- yaakov34 15y agoVitriol? Certainly not from me - I think I've been perfectly civil, and I am trying to understand the situation. I don't think your explanation holds. Pagerank is based on reputation (created by links and other means which Google isn't very specific about), more than on contents. jquery.com has a pagerank of 8, so it can't be all inaccessible to Google. The GP says that blog.jquery.com is inaccessible. So how does jquery.it get a pagerank of 8? Your explanation would look correct if jquery.com had a pagerank of 1 (due to misconfiguration) and jquery.it sneaked in with a rank of 2. But this isn't the pattern.
- bhartzer 15y agoJust look at the backlinks to jquery.it and it should be pretty clear why that site ranks in Google for jquery-related searches.
- yaakov34 15y agoCan you share some details about what you found? I am not much of a webmaster, but naive searches with Google don't find any massive backlinkage to jquery.it. Maybe they are already cleaning the index, or maybe I didn't search it right. EDIT: 1 more datapoint: bing doesn't return anything from *.jquery.it for the search in question.
- fhars 15y agoOne factor may be unique fresh content. If jquery.com only returns 500 to googlebot, jquery.it is one of the most content rich sites with fresh content on jquery 1.6.2. When I tried it some hours ago, jquery.com had first place for the query "jquery", but was only on the second page for the more specific query "jquery 1.6.2", while the jquery.it page was still on the first page for that query, but had lost some of its freshness bonus (or my search history had pushed it down, you do no longer really know what is googles ranking and what is your own confirmation bubble these days), appearing on place 7, two places below the OPs post. The other top places where more or less shady news aggregators and download sites. [edit: right now I see the download site at heise.de as the highest ranking download site on place 2 on google.de, which is about as reputable as it can get, so google does actually do something useful here, given the misconfigured jquery.com server.]
- yaakov34 15y agoI seems that only blog.jquery.com is misconfigured. jquery.com is OK. But it makes sense that this would be a part of the problem, since the blog has the freshest content, and the bad result does come from the blog. However, this still doesn't explain why jquery.it has such a high pagerank. And blog.jquery.it has a pagerank of only 1. I can't find any interesting-looking links to jquery.it that would explain the high rank. It's also strange that Google returns blog.jquery.it 1st for the search "jquery 1.6.2", since those strings are found at lots of highly ranked blogs (pagerank>1 for sure). We could be dealing, once again, with Google's preference for finding search terms in the domain name. But this puzzle is still not coming together for me. This is why I'd like wiser persons than me to try to get to the bottom of this.
- jeresig 15y agoHi Pierre - Thanks for the details on this! We dug into our Wordpress and realized that W3 Total Cache was configured to block anything that had the word 'bot' as part of its user agent string (sigh). That's now fixed and live. As to the redirect - that's actually a bit of devious magic on our end. Since jquery.it is hotlinking to our JavaScript and CSS we just added a bit of JavaScript to automatically redirect them to the right jQuery.com page: if ( /jquery\.it/.test( window.location.hostname ) ) { window.location = window.location.href.replace('y.it', 'y.com'); } Thanks again for your pointers - looks like the jQuery 1.6.2 blog post is already showing up in the search results and jQuery.it is not in the first page of results (although "jquery 1.6.1" hasn't updated yet and jQuery.it is still there - I assume that it'll be remedied as the Google spider makes more progress). Thanks again!
- yuhong 15y agoOf course, the best solution is to get the jquery.it domain name transferred to the genuine owner of jQuery, but that will take time.
- jeresig 15y agoNaturally, we're talking with our legal representation about that - we'll have to see what we can actually do. I think we're "ok" for now though as users are getting to the right place in the end.
- pierrefar 15y ago>configured to block anything that had the word 'bot' as part of its user agent string Yep, that'll do it :) Glad you found it and fixed it quickly! One follow-up thought for you and anyone in this situation: Set up a Google Webmaster Tools account ( http://www.google.com/webmasters/tools/ http://www.google.com/webmasters/tools/ ) and make reviewing it part of the webmasters' daily routine. In this particular case, the Crawl Errors page in the Diagnostics section would have flagged this problem very quickly. A couple of pro tips for Webmaster Tools: 1. Be sure to look at the "date detected" column because it is accurate and somehow people miss it. 2. Set up email forwarding for messages Webmaster Tools sends to you: http://www.google.com/support/webmasters/bin/answer.py?answer=140528 http://www.google.com/support/webmasters/bin/answer.py?answe...