31 ms·
How Google’s Web Crawler Bypasses Paywalls
- dude_abides 11y agoOr simply use incognito mode and click on Google search result.
- mrmcd 11y agoDid you read the article. It talks about how that trick no longer works on a lot of sites because they are now checking User-Agent strings too.
- lstamour 11y agoActually I noticed sites have simply changed policies -- if you're a regular visitor your cookies will identify you and block content. The Incognito mode trick works for WSJ and others that would still check the referrer header. Allowing Googlebot access and checking the referrer header are two different things. Also, Google has published IP addresses it uses, so this extension might not last long...
- slig 11y ago> Also, Google has published IP addresses it uses, so this extension might not last long... They do not [1], but you can find out by doing a reverse DNS query. > "Google doesn't post a public list of IP addresses for webmasters to whitelist" [1] https://support.google.com/webmasters/answer/80553?hl=en https://support.google.com/webmasters/answer/80553?hl=en
- lstamour 11y agoSorry, that was what I meant. :)
- crazysim 11y agoDoesn't this kind of also hurt SEO? I'm would guess Google has some automated system to detect and apply a negative signal to sites that provide different content to a Googlebot user agent than a non-Googlebot user agent. I guess these sites are counting that the other signals outweigh that negative hit. Otherwise, why would expertsexchange be obligated to provide the answers at the very bottom? Did something change?
- eli 11y agoI'm 99% sure I've encountered a Googlebot crawling pages with the UA of a regular browser, presumably for exactly this purpose.
- chinathrow 11y agoThey do this via an iPhone like UA, though googlebot is still in there too.
- eli 11y agoI'm pretty sure that's just to see if the site is serving different content to mobile vs desktop. I think they also sometimes hit pages with no mention of googlebot.
- x0 11y agoIf you do an nslookup on the IP, it should come up with crawl-xx-xx-xx-xx.googlebot.com, where xx-xx-xx-xx is the IP.
- skocznymroczny 11y agoexpert sex change
- deleted 11y ago[deleted]
- anewhnaccount2 11y agoIf this is true, what WSJ is doing is called "cloaking" and should cause it to get de-indexed: https://support.google.com/webmasters/answer/66355?hl=en https://support.google.com/webmasters/answer/66355?hl=en
- vonklaus 11y agoconversely, everyone should actively cloak and use random generated numbers to dynamically serve variants of their content similar to how mapmakers use trap streets. That way, like, another company wouldn't be profiting directly off of their work and threatening to sort of, destroy their entire business if you disagreed.
- jonknee 11y agoWhat? Any site can very easily not be in Google if they choose to. It's a very dumb decision for a news site, but you're free to do it.
- vonklaus 11y agoIts a false choice without a conpelling alternative. Like saying, anyone upset with the status quo should vote. I was joking a bit, but I also wasn't. Google has end to emd control over some users internet experience, and much of it in other cases. They own: * 100s of thousands of servers * domain registrar * ~50% of web browsers in US. * code CDN, FontService * define web standards * hundreds of millions of emails. * CA implementation * ISP infrastructure * Develop software for a large part of thr mobile ecosystem. * decide what you see when you go to search (most search copy google, buy results, or both) * also many of the web beacons and advert targeting. * oh, and the largest collection of video and images in the world. So when you say, just do what they say or get deindexed, and you present it as if that is reasonable(not just you but the collective you) I just think I must be insane. I mean, assuming google is good (i fo mostly) doesn't mean I would let them become the entire internet. Real question, if google were to disappear vs. the "too big to fail banks" that would have gone under, where a case could be madr for a few certainly failing, what would have bigger impact today? Tl;dr everyone cares about single point if failure except at the macro system level: finance, banking, healthcare, etc
- slig 11y agoIf they're now blocking clicks from Google, doesn't that mean that they're cloaking and violating the Google's Webmaster Guidelines [1]? [1]: https://support.google.com/webmasters/answer/66355?hl=en https://support.google.com/webmasters/answer/66355?hl=en
- elaineo 11y agoGoogle is not okay with cloaking, but they will whitelist publishers if the publisher specifically includes a parameter that declares if the site requires registration or subscription. This is done in the sitemap. https://support.google.com/news/publisher/answer/74288?hl=en https://support.google.com/news/publisher/answer/74288?hl=en
- slig 11y agoI didn't know that, thanks. But reading, it seems that option is about Google News, not the main Google Search. > News-specific tag definitions > Yes, if access is not open, else should be omitted > Possible values include "Subscription" or "Registration", describing the accessibility of the article. If the article is accessible to Google News readers without a registration or subscription, this tag should be omitted.
- ikeboy 11y agohttps://support.google.com/news/publisher/answer/40543?hl=en https://support.google.com/news/publisher/answer/40543?hl=en seems to specifically ban this. WSJ is in violation, not fitting any of the categories there.
- pacquiao882 11y agoWSJ is big enough to negotiate their own terms with Google Search.
- morgante 11y agoThey're really not. They need Google a lot more than Google needs them. One of these things is going to happen: (1) They end this "experiment." (2) They stop serving Google the full content. (And see their rankings drop accordingly.) (3) They get delisted for cloaking.
- sylvinus 11y agoWell, that trick won't last long either. It's trivial to verify that an IP indeed belongs to Google: https://support.google.com/webmasters/answer/80553?hl=en https://support.google.com/webmasters/answer/80553?hl=en
- deleted 11y ago[deleted]
- dogma1138 11y agoSeems to work if you deploy a proxy on Google's app engine and use it to access WSJ ;)
- slig 11y agoA App Engine wouldn't have a IP with a reverse DNS *.googlebot.com, would it?
- dogma1138 11y agoNope, but it does resolve to something .google and if you check netblocks.google.com it will appear there so they might not be limiting it to googlebot only at this point. https://support.google.com/a/answer/60764?hl=en https://support.google.com/a/answer/60764?hl=en
- dsl 11y ago_netblocks.google.com includes only the servers that handle Gmail and corporate email.
- greglindahl 11y agoI don't think that any server that random people can start a proxy on is in Google's SPF record.
- dogma1138 11y ago
- deleted 11y ago[deleted]
- matt_wulfeck 11y agoI like wsj but I only read maybe 1 article every other day. They need a more reasonable price point, especially since the market will almost bear no price at all. That being said I do enjoy their content, save for maybe the op-eds.
- acomjean 11y agoI'm surprised that most online papers won't sell you one day's worth online for a buck or so. Like buying a real newspaper. They all seem to want to sell subscriptions, which are perpetual and probably difficult to cancel..
- matt_wulfeck 11y agoEven a dollar a day. I can pay Netflix $10 a month and stream unlimited HD video, but wsj wants $30 to read the first few paragraphs of a few articles a day? The pricing here is much too aggressive
- Nadya 11y agoVideos are more relevant to rewatch than articles are to reread. One needs to output more articles than videos, because articles must be "fresh" or you'll lose an audience. Nobody is printing 40 year old news - people are still watching 40 year old movies. Playing devil's advocate here. Pricing for many online goods is almost completely arbitrary and varies with little accord to service/product quality or even what that service provides. Another related example of arbitrary pricing: people will pay $2 for a soda from a vending machine but won't pay $1 for a useful app on their phone. There's something going on there... The sooner the $1 app's figure out what makes people buy $2 sodas is the day they become rich. And the sooner content providers figure out why people will pay $20/month to stream media (Let's say $10 to Netflix, $10 to Spotify or something) and charge people $20/mo for their articles... things will turn around for them.
- majewsky 11y ago> Nobody is printing 40 year old news - people are still watching 40 year old movies. Because news operate at a different scale than movies. People are still reading 40-year-old books. A 40-year-old news article is not relevant anymore because it only fits within the momentary context in which it was created, whereas a non-fiction book or even an essay can span a broader context and thus stay as informative for future readers.
- jgh 11y agoI just tried clicking on "Harper Lee, Author of ‘To Kill a Mockingbird,’ Dies at Age 89" from wsj.com's homepage and got the paywall. I then pasted the headline into google and clicked on it from Google results and did not get hit by the paywall.
- e40 11y agoMine did, which surprised me, and I did exactly the same thing.
- gsibble 11y agoI also checked the paywall and the old trick still works for me. Odd.
- zem 11y agothey're probably doing some sort of a/b testing by selectively letting some clicks through
- adamrights 11y agoThis is basically true ^^
- simonswords82 11y agoAny idea what the deal is with SEO impacts of WSJ taking the idea of blocking everyone who isn't a google bot?
- creativityhurts 11y agoI did the same with the same article and got paywalled from Google results.
- throwaway21816 11y ago>Archaic news source does something to hurt their market penetration to internet Great idea here guys
- mbroshi 11y agoAm I alone in feeling like this is akin to a tutorial on how you can shoplift without getting caught? WSJ, for better or worse, does not want to give you content without your paying for it. If you take that content without paying, you are stealing. Just because you have figured out how to get past their security does not mean it's not stealing. (See the second precept here: https://en.wikipedia.org/wiki/Five_Precepts https://en.wikipedia.org/wiki/Five_Precepts)
- deleted 11y ago[deleted]
- elaineo 11y agoI don't disagree with you. But, given the nature of this forum, I think that the information content has merit. What people choose to do with the information is another story...
- azakai 11y agoLegally it might, I don't know. Morally, I'm not sure. 1. If they allowed all bots but disallowed all non-bots, that would raise questions of what defines a "bot". Couldn't I write a personal bot that fetches the story for me? As a browser addon, even? 2. It's even more complex since allowing bots means they allow tools that provide the information to third parties, as the bots are not intended for private use by the bot maker. So the door is already open. 3. But in practice, it seems they favor certain bots. Is it ok that the WSJ lets Google do things Google's competitors cannot? Try to use the "web link" trick from HN on any other search engine, and it doesn't work in my experience. That seems anti-competitive and discriminatory in favor of the existing dominant entity in this space, Google.
- mankyd 11y ago> If they allowed all bots but disallowed all non-bots, that would raise questions of what defines a "bot". Maybe, but I think its a pretty easy distinction. They aren't even allowing all bots - they're allowing a white list of them. You're not just writing your own bot to get around it, you're pretending to be someone else's bot. > But in practice, it seems they favor certain bots. Is it ok that the WSJ lets Google do things Google's competitors cannot? That's the really important question. I personally have no context for answering except to say that I can see both sides argued. If you view their website as a physical store / private establishment, then I assume that they have every right to establish who has access to what and under what conditions. Of course, that hampers a lot of legitimate use cases along the way.
- mchahn 11y agoBypassing the paywall is more unethical that blocking ads. It is one thing to have control over your own browser but another to steal something from another site. Also, isn't it illegal to bypass computer security?
- nkrisc 11y agoHow is modifying your own request headers any different than choosing to not display content returned in the response body? Their server can choose to do what it wants with your request and you can choose what to do with the response it sends. Are User-Agent headers legally protected identities?
- drostie 11y agoUnfortunately, yes. The law legally protects anything which you might use to gain unauthorized access; that includes e.g. a password field. (That is, it is indeed breaking into a system if you type in the correct password, say by reading it on a post-it on someone's monitor, but you are not supposed to know that password.) This sort of thing makes the entire business of law complicated and impossible to automate. Then again, lots of sci-fi dystopias are dreams of an automated law that somehow destroys the fabric of society, so...
- effie 11y agoThe difference is substantial; in principle, one can modify his request to intentionally bypass the authorization mechanism occurring on the server. One cannot mislead anyone/anything by displaying his data in a customized way in private on his computer.
- nkrisc 11y agoFair point. Intent matters. However must it be demonstrated that there is some specific malicious intent behind changing a request header versus simply changing your user agent for the heck of it? In a related point, some news sites load a modal and prevent scrolling over an article asking you to sign up. However if the full article is included in the response and I read it by simply viewing the response body (HTML) is that circumventing security? Actual example. In this case modifying how the response is rendered in my browser, I can bypass their intentions. Obviously the better way of doing this would be to not send the entire article content until they've determined I should be able to view it.
- zaroth 11y agoAnd congratulations, you have likely just "exceeded authorized access" and committed a felony violation of the CFAA punishable by a fine or imprisonment for not more than 5 years under 18 U.S.C. § 1030(c)(2)(B)(i). From the ABA: "Exceeds authorized access is defined in the Computer Fraud and Abuse Act (CFAA) to mean "to access a computer with authorization and to use such access to obtain or alter information in the computer that the accesser is not entitled so to obtain or alter." To prove you have committed this terrible felony, the FBI will now demand that Apple assist in disabling the secure enclave of your device in order to access your browser history. But remember, they only need to do this because they aren't allow to MITM all TLS and "acquire" -- not "collect" -- every HTTP request your machine ever makes. </s>
- jrockway 11y agoUser agent strings have a long history of being intentionally misleading. IE 11 claims to be "Mozilla/5.0". Chrome claims to be "Safari/537.36". The User-Agent string is all lies, and has been ever since the first site started doing UA sniffing.
- zaroth 11y agoIt's intent that matters. Setting user-agent in order to properly render a page is legal. Setting a user-agent string to gain access to otherwise unauthorized content is probably not.
- profeta 11y agointent can be many things. i disable javascript. so i can't even see their page (and hence i don't know my access is being denied since i got a 200 http response, which means "OK") so i try different user agents with the intent of reading the content they are providing. Just like microsoft case.
- jsprogrammer 11y agoUser-agent string is not an authorization mechanism.
- eps 11y agoCorrect me if I'm wrong, but wasn't there a long standing Google's policy that the version of the page served to their crawler must also be publicly accessible. That would then be the reason why WSJ articles were accessible through the paste-into-google trick, rather than because WSJ was incompetent and failed to "fix" the bypass. So does it mean that Google will no longer index full WSJ articles or does it mean a change in the Google's policy?
- morgante 11y agoYou are correct, Google requires that you let users see the first click for free if you want to index content behind a paywall. [1] Since this is billed as an "experiment" I'm guessing that WSJ is just testing the waters. If they roll it out to everyone, they will have to serve only snippets to Google or risk getting delisted. [1] https://support.google.com/news/publisher/answer/40543?hl=en https://support.google.com/news/publisher/answer/40543?hl=en
- zem 11y agoi thought of doing that when the "search google" trick stopped working, but i decided it crossed the point where i would feel like i was unfairly circumventing their clear desire not to serve me the content. i've just added wsj to my mental ignore list and count it as a few more minutes gained to do something else.
- ivan_ah 11y agoYeah, same here. Every time I get to a tab where I see a paywall I just close that tab, probably saving 5-10 mins of my life!
- metafunctor 11y agoI'm pretty sure Google will soon stop indexing WSJ. Why index something if the vast majority of users cannot access the pages behind the links? EDIT: The "paste a headline into Google" trick still works for me, though. If this continues to be the case, they will keep indexing, of course.
- ikeboy 11y agoIt doesn't work for me. They did say they're "testing" it, so maybe A/B testing conversion rates. And yes, this violates Google's policies laid out explicitly at https://support.google.com/news/publisher/answer/40543?hl=en https://support.google.com/news/publisher/answer/40543?hl=en
- Joky 11y agoIndexing is fine, a great feature would be if Google was able to show it only to the user that can access it.
- jsprogrammer 11y agoWhy should Google manage WSJ's paywall?
- yeukhon 11y agoIt doesn't. It 1) penalize WSJ, 2) personalize to WSJ subscribed Google users.
- rhino369 11y ago>Why index something if the vast majority of users cannot access the pages behind the links? So people can find it? I'd be pissed if Google de-indexed something like IEEE because it has a paywall. Assuming the internet has to be freely available is a mistake. Especially with the continued growth of an adblocked internet. We could be facing an internet with significant paywalls in the future. I'd support a "free" search term to weed out paywalled results. Furthermore, Google shouldn't be making normative judgements about what people should see. It's an abuse of their monopoly.
- jasonwilk 11y agoI've noticed that this has stopped working on WSJ if you've already hit the paywall and try to google the article to bypass.
- coverband 11y agoMy Windows anti-virus deletes the linked sample code automatically upon download, marking it as "Trojan:Win32/Spursint.A". Did anyone have the same experience? (I was actually more interested in using it as a template for writing a simple Chrome extension.)
- mattmaroon 11y agoYep. I then pasted it but it didn't work on wsj.com. Oh well.
- ikeboy 11y agoNew workaround: paste the article title into archive.is. I don't know what they're doing but they have a workaround of some sort.
- LunaSea 11y agoThat's actually really interesting! Could anyone chime in and explain how they might work around this issue?
- ikeboy 11y agoMy guess is they have a login, but they could just be using some workaround like described in OP. I've used them to save Facebook posts before, and the pages were logged in to some "Nathan" IIRC. They probably have a bunch of hacks for specific sites that needed fixing.
- mikedmiked 11y agoThis is what I do. I actually made a bookmarklet with the following pasted into the URL, so you can do it in a single click: javascript:void(open('https://archive.is/?run=1&url='+encodeURIComponent(document.location)) https://archive.is/?run=1&url='+encodeURIComponent(document....)
- chinathrow 11y agoSo soon they have to block anyone with a fake Google UA and whitelist the well known 66.249 IP range. Trivial.
- deleted 11y ago[deleted]
- lloyddobbler 11y ago"Remember: Any time you introduce an access point for a trusted third party, you inevitably end up allowing access to anybody." See also: http://www.apple.com/customer-letter/ http://www.apple.com/customer-letter/ :)
- m52go 11y agoI came back here to post this line! It's so perfect.
- elaineo 11y agoHaha! Thanks for catching that :)
- melted 11y agoExcept of course what FBI is proposing wouldn't give access to "anybody". Just unfettered access for FBI, CIA and NSA by way of gag orders and national security letters, forcing Apple to break security of their hardware and not speak about it publicly. Of note is the fact that it wouldn't even give direct access to FBI, since they don't have the firmware keys (yet).
- interpol_p 11y agoYeah but if you don't trust those companies to be secure and to maintain that security throughout all their employees and contractors (see: Snowden). Then you must assume giving someone access is the same as giving everyone access.
- kbenson 11y agoIn this example, the FBI is the "trusted third party", but by giving them access, we inevitably open access for everyone, as the system is no longer strongly secure. The trusted third party in the quote isn't asking for access for everybody either, but in the end that's what happens.
- melted 11y agoApple isn't giving access. Apple would be required (by court, unless they manage to fight this off) to install a signed custom build of the OS in order to give access to that particular device. FBI would not have this build, nor a key to create their own signed custom build.
- dheera 11y agoGreat, build a better mousetrap and they will build a better mouse. The next better mouse will be IP-based whitelists of Google's crawler.
- warrenmar 11y agoYou can also access WSJ for free at the library.
- creativityhurts 11y agoIt reminds me of this: "Trying to save a quarter..." https://www.youtube.com/watch?v=j4nRHHPpnVc https://www.youtube.com/watch?v=j4nRHHPpnVc
- jdunck 11y agoIf Google (or any other crawler) wanted to play nice with paywalls, they could issue a public key for their bot, and put a signature in their User Agent string that the domain could then verify. Those signatures could obviously leak, but on a per-domain basis. Perhaps the domains could have a secure way of bumping the valid key generation if they had a leak.
- AnthonyMouse 11y agoThere are two problems with this. First, they don't want to. In fact, if a search engine can figure out that a link is going to lead to a paywall, they'll probably want to reduce the ranking of the result, because the user is not going to want results they can't actually look at. Second, it would be a massive antitrust violation because it would prevent access by competing crawlers. The only way around that is to allow access to anyone who claims they're a crawler, which was the original problem.
- LunaSea 11y agoThe current situation with the WSJ could already be considered an antitrust violation. It's whitelisting one crawler and leaving the other ones out.
- desdiv 11y agoGoogle (and every other major search engine) already provide a way, i.e. reverse DNS lookup, to authentic bot ownership: https://support.google.com/webmasters/answer/80553?hl=en https://support.google.com/webmasters/answer/80553?hl=en AFAIK no content provider actually does this check though.
- jupp0r 11y agoIt's not bypassing at all. Googles crawlers are deliberately let in because a paywall that nobody runs into is useless.
- systemz 11y agoSo their next move is check if IP is from Google
- philip1209 11y agohttps://cloud.google.com/ https://cloud.google.com/
- philip1209 11y agoI wonder how many Google Cloud customers use the servers to run spoofed Googlebot crawlers from the Google IP range in order to bypass paywalls and scrape large sites (like LinkedIn) without hinderance.
- hueving 11y agoBased on the comments here, am I to understand that constantly browsing the web with my user agent string set to a googlebot string, I am committing a felony? How would I even know which sites I'm gaining unauthorized access to? That is completely idiotic if there is a string you can put in a Mozilla browser config that is literally illegal to browse the web with.
- effie 11y agoI do not think using any User-Agent alone constitutes a crime. There are valid non-criminal reasons why one would like to use Googlebot's or other User-Agent. I think it is the intent to bypass the paywall + success to do so that may be regarded as offense or even crime, but I'm not sure.
- majewsky 11y agoGood luck trying to argue to a judge that the law is forbidden from being idiotic. ;)
- chrishn 11y ago> Remember: Any time you introduce an access point for a trusted third party, you inevitably end up allowing access to anybody. coughNSAcough
- jrochkind1 11y agoI thought Google specifically disallowed returning different pages based on User-Agent targetting googlebot, and this included paywalls. Are they running afoul of Google policies and going to get pinged by Google? I can't find the text from Google now (when can you ever find any docs at google?), but I am very certain I remember reading from them that you may not return different content to GoogleBot based on User-Agent.
- 0xCMP 11y agoIt's broken already. Tried to access an article about new china rules for online news and it pay-walled me. They're probably looking for clients coming from googlebot.com now.
- Illniyar 11y agoAren't you supposed to verify if a visitor is a googlebot by reverse lookup of the IP address? I.E.: https://support.google.com/webmasters/answer/80553?hl=en https://support.google.com/webmasters/answer/80553?hl=en User-agents are notoriously unreliable.
- mikestew 11y agoSo does HN now choose to not post articles from the WSJ? I was comfortable with the "google it" trick, and frankly was a little annoyed with constant "paywall, wah!" comments when what should be by now a well-known workaround was available. But that workaround no longer works.
- mark-r 11y agoThey've been testing the new wall for a while now. I know I made one of those "paywall wah" comments when the Google workaround didn't work for me. Then the next time I tried it worked fine, so it must have been random selection.
- Gratsby 11y agoIf you hit a paywall or a "sign up to access this content" message from a google search result, report it. Google will remove them from the search results, they will lose their largest traffic source, and they will address the issue. Or they won't because they have enough paying customers.
- mangeletti 11y agoThis is not meant to be purely controversial, but I thought long and hard about WSJ back a few months ago when HN mod (always forget his name) said to stop complaining about HN links being posted because paywalls were ok. I agree paywalls are ok. But some things are not ok. Take a look, for instance, at the WSJ.com home page with an ad blocker turned on (note all the missing letters and scrambled up titles). They want me to pay, and they want me to see ads, and they want to track my behavior? Should I send them my DNA also? Organizations like WSJ are exactly the disease that causes ad blockers to proliferate and ruin the web for all the decent publishers. They're at war with my privacy (by breaking their site intentionally when I visit with a blocker on). They want it all, ads, tracking, your private data, and subscription revenue, not to mention... # Agenda-Driven Content I mean, we're basically talking about NBC or Fox here, just on the web. Imagine every morning when you woke up you turned on the television and tune to some "news" show. After talking about the weather, they start talking about a lost pickle that is thought to be potentially alive and moving about with free will. Over the next two years, talk about the same pickle extends to every other TV show. Before you know it, everybody in the nation is talking about the same pickle. Years go by, and that pickle has become a part of our society, and that's not because people are born with an innate care the well-being of pickles, but because "news" shows taught them to be. That's not a good position to be in. I have to believe I'm not the only one in here that doesn't watch any TV. So, why do we all treat the same media giants differently on the web? We crave their content so much that we build browser add-ons to get to their content, etc.
- Laaw 11y agoIf they make sending your DNA a requirement of consuming their content, then yes, you send it to them if you want their content. That's their right, as owners of something, to dictate its use. You aren't entitled to WSJ.com, NBC, or Fox.
- mangeletti 11y agoI'm not a collectivist, nor do I believe I'm "entitled" to anything, including human rights. I'm simply stating that "I don't support organizations like WSJ"... as you can clearly see, if you read my comment. I don't propose anyone ban them. I simply "don't know why we [as a society] support them", also clearly in my comment. My point, which your political diatribe disallowed you to notice, was that we're fighting awful hard to consume their garbage. Meanwhile, there is plenty of content out there there is free, not because of socialism or entitlement, but for the same reason as blockers became popular... because information is so readily available now that nobody is ready to send their DNA in.
- mikemikemike 11y agoThis is an odd debate. Let's say a restaurant declares "veterans eat free." This blog post is like a friend telling you "Hey if you tell this restaurant you're a vet they'll give you a free meal." No one said it's legal or ethical. It's lying to trick someone into giving you something at their expense. I think the relevant point, underscored by the author's last sentence, is it doesn't matter who you open a back door for - it opens the possibility for anyone to barge through.
- elaineo 11y agothat's a good analogy.
- GigabyteCoin 11y agoI was under the impression that the "hack" whereby you searched for the article on Google and clicked through to that article (effectively skipping over the paywall) was a demand of Google's and not an oversight by the paywalled website. I thought that google deemed providing search results which were behind paywalls as a "bad experience" for their search users, and would penalize websites for doing so. Is this no longer the case?
- elaineo 11y agoGoogle doesn't demand anything. If your paywalled website is not accessible by Google's crawler, then Google will not index it. Publishers want Google to index their pages and drive potential paying visitors, which is why they open the loophole themselves. For the second point, Google does require that publishers specify "registration required" in their sitemap.
- GigabyteCoin 11y agoIf you're showing Googlebot one thing, and visitors who visit your website through google another, that's essentially "cloaking" (a blackhat SEO technique). At least it used to be.
- tete 11y agoDoesn't Google usually try to punish websites that show users something different and even mentions that somewhere? Not an SEO Expert here, but wonder how and whether Google will end up handling that. I mean making an exception could also be considered abuse of power in some countries of the world. Don't have any strong opinion yet on that, just saying that because of how the EU exercised certain laws in recent years.
- pmontra 11y agoThey'll start allowing only some IP addresses search engines agreed with them.
- spitfire 11y agoIs there a version of this available for Safari?
- kenshaw 11y agoBasically, the article is stating to change the User-Agent to GoogleBot or Bing or whatever other crawler UA you'd prefer. While that's doable, that's something that is easily detectable and prevented, as all of the big crawlers can be validated against DNS. Additionally, I would like to point out that I wrote a Varnish extension for the express purpose of validating User-Agent strings through DNS lookups, and is available here: https://github.com/knq/libvmod-dns https://github.com/knq/libvmod-dns It was built because we had specifically a problem with bad bots crawling a large site (multiply.com) and this was one of the easiest ways to filter out the bad bots from the good, and to enforce robots.txt policies on a per bot basis. It works very well, as you can do any kind of DNS caching internally and prevent this kind of behavior, if that's your goal.
- amelius 11y agoFix: replace the user agent string by a cryptographic challenge/response scheme.
- obelisk_ 11y ago1. Google's Web Crawlers are not "bypassing" paywall. It's the paywall that let's crawlers through. I.e. exactly the reverse of what the author implies with their headline. 2. The idea that this is somehow new is wrong. The way for a server to identify crawlers have "always" been to look at the user-agent, and, when done right, IP, verified either by net block owner or by doing PTR lookup and then checking that the A or AAA record for the claimed host points back at the same IPv4 or IPv6 address. Meanwhile, I do agree that paywalling is a more recent phenomenon, at least with regards to the extend it is popular among sites today, but the concept of presenting different data to crawlers and visitors arose much earlier and is something Google have been aware of and has made sure to delist such sites when found, whereas in fact Google has since then moved abit in the direction of allowing it in that they do so for Google News if declared as explained by others ITT. So in my view, it seems that the author is jumping to incorrect conclusions based on an incomplete understanding of what's actually going on here. What then about the HN readership, how come this article became so highly voted and I don't see these issues raised by anyone else? Or maybe I'm just crazy?
- tomkwok 11y ago> Google's Web Crawlers are not "bypassing" paywall. It's the paywall that let's crawlers through. I.e. exactly the reverse of what the author implies with their headline. Don't nitpick. It's just a shortened version of How To "Be" a Google’s Web Crawler to Bypass Paywalls. You get it. I get it. Everyone gets it.
- deleted 11y ago[deleted]
- yyin 11y agoDoes WSJ check visits from a Googlebot UA against a list of known Google IP addresses?
- daveheq 11y agoPossible in Firefox? Some people won't use Chrome.
- f137 11y agoI wonder if anybody tried to do as suggested? I copied the files to Chrome as per instructions, and the paywall was still in place.
- mildweed 11y agoSolution: Content providers register a (yet-to-be-written) Google News API account, get an API key, with which Google indexes the site and the site recognizes as legit.