6 ms·
Ask HN: How does archive.today bypass paywalls?
I can understand how it does it for simple paywalls that can be bypassed using a user agent pretending to be Google.
But how does it work for firmer paywalls with, e.g. NYT?
- gsora 3y agoSometimes the paywall is nothing more than a couple HTML div’s. You can programmatically remove them and then return the page.
- m348e912 3y agoI wondered the same thing, I just figured they were whitelisted by the paywalled sites but I don't know exactly why. Perhaps archive.today has accounts to these paywalled sites although I am not sure that's the case. There is another site, 12ft.io, that used to bypass many of these site's paywalls but they got pushback and now longer offer bypass for sites like WSJ and NYT.
- tpmx 3y agoMy guess: stolen usernames/passwords. Or actual paid subs.
- Bender 3y agoDoes that mean they would have to rewrite each site to remove the username specific links and content out of the mirror to keep the origin sites from banning their accounts?
- ipaddr 3y agoIdentify as google-bot
- simpli 3y agoIdentifying as Googlebot no longer gives access to most paywalled news.
- navjack27 3y agoI mean you could kind of do the same thing locally if you use archivebox. I think it's just the way it scrapes the web page.
- stonogo 3y agoI believe the question here is, what is "the way it scrapes"?
- deleted 3y ago[deleted]
- guilhas 3y agoI would guess most paywalled sites will allow Bots to scrape content to improve site discoverability. Whether it is Archive.org, Archive.today, Google, Google news...
- chmaynard 3y agoPerhaps they subscribe (gasp!) to sites with paywalled content, sign in, and begin scraping.
- throwaway67743 3y agoMostly used agents, sites want to be crawled to appear in search results, but from time to time it does actively login to sites (linkedin was one that stopped working, presumably suspended as it likely violates t&c)
- nora-puchreiner 3y agoNYT has one of those simple paywalls indeed: https://gitlab.com/magnolia1234/bypass-paywalls-chrome-clean/-/blob/master/sites.js#L2068 https://gitlab.com/magnolia1234/bypass-paywalls-chrome-clean...