3 ms·
If you change your browser's UserAgent string to Googlebot, then your client will be treated as a first-class citizen, by many of these sites. Google always win
by devrandomguy 9y ago
If you change your browser's UserAgent string to Googlebot, then your client will be treated as a first-class citizen, by many of these sites. Google always wins, so let's all be Google.
- Xeoncross 9y agoWorks great until you stumble on one of the big sites who will auto-ban you for not having a valid Google IP address.
- devrandomguy 9y agoIs that really a thing? That must be such a hazard for their developers. I usually have a test for sites that I work on, that scrapes a few URLs as Googlebot, to verify that they are getting an optimized view (no JS, structural-only css).
- Normal_gaussian 9y agoGod. That's the reason so many sites look great in the results and are confusing interactive messes when I get into them.
- devrandomguy 9y agoIf it's any consolation, the site is well tested in Noscript mode.
- SomeCallMeTim 9y agoYou sir are a god among Web developers. If only they all did this. So many sites I get to and they're a blank page or an absolute disaster....
- SomeCallMeTim 9y agoYou sir are a god among Web developers. If only they all did this. So many sites I get to and they're a blank page or an absolute disaster....
- ChuckMcM 9y agoYes. Googlebot only crawls from legit addresses (even when their developers are trying new things) so it's an easy scraper/scammer signal to key off of.
- dmn001 9y agoNo. Most websites don't do this.
- Spivak 9y ago* Reverse DNS records. Webmasters shouldn't be verifying Google's bots by hard coding IP addresses. https://support.google.com/webmasters/answer/80553?hl=en https://support.google.com/webmasters/answer/80553?hl=en
- imron 9y agoShouldn't != never happens
- deleted 9y ago[deleted]
- dredmorbius 9y agoCIDR blocks and ASN advertisments are cheap. Update those periodically (hours / days / weeks). The adverts don't change particularly quickly.
- dmn001 9y agoIt's extremely rare to be ip-blocked by any website just for using the Google's user agent from a non-specific range. IP's get re-used and you can switch to a new one easily, so it's really not common or good practice for this to happen.
- kbenson 9y ago> IP's get re-used and you can switch to a new one easily, so it's really not common or good practice for this to happen. On the flip side, some people can't change their IP addresses easily, and getting IP banned (even if rare because of the reasons you stated) is actually a major hassle when it actually happens for those people. :/
- ipsum2 9y agoI just tried this, doesn't work on WSJ. Tried the user agents listed here: https://support.google.com/webmasters/answer/1061943?hl=en https://support.google.com/webmasters/answer/1061943?hl=en
- bduerst 9y agoIsn't that what the entire article is about? That WSJ no longer gives access to Google users (and bots)?
- mirimir 9y agoNo. It allows Google bots to see full articles, but shows only the first paragraph or so to non-subscribers. Even if they're coming from Google search results. However, I don't see cache links on Google :( Edit: Oops, I'm wrong. The article does say that the Google bot only sees the first paragraph or so.
- Jimmie_Rustle 9y agoNo. You're wrong, as it states in the article: "The reason: Google search results are based on an algorithm that scans the internet for free content. After the Journal’s free articles went behind a paywall, Google’s bot only saw the first few paragraphs and started ranking them lower, limiting the Journal’s viewership."
- mirimir 9y agoYes, you're right. I got confused by all the discussion about Google checking for cloaking by comparing results using different user agents. So maybe this is why there's no Google cache. Also, if Google can only index the first few paragraphs, the results are much less comprehensive.