4 ms·
Off topic, but I'm always amazed by Archive.md/.is/whatever. To this day I don't understand how they manage to bypass a lot of paywalls. The mystery about the
by reddalo 8mo ago
Off topic, but I'm always amazed by Archive.md/.is/whatever. To this day I don't understand how they manage to bypass a lot of paywalls.
The mystery about the owner makes it even more intriguing.
- LordHeini 8mo agoI think archive has mostly news, random articles and such. And as they say nothing is more worthless than yesterday's news.
- ventegus 8mo agothetimes.com has a paywall if you visit it from the UK, and full content if you are in the US. entonces, US-based archive.org "bypasses" this paywall as well: https://web.archive.org/web/https://www.thetimes.com/culture/books/article/sicilian-man-leonardo-sciascia-rise-mafia-struggle-italy-soul-caroline-moorehead-review-lbsbd2p5w https://web.archive.org/web/https://www.thetimes.com/culture...
- jama211 8mo agoI just assumed they copied it into their own db
- amouat 8mo agoI assume they just pretend to be the Googlebot so the site just gives the text.
- dewey 8mo agoWon’t work for any popular site. You can try that easily by using extensions to set the user agent. If you are not checking the public list of IPs that Google publishes for the crawler you are doing it wrong.
- moffkalast 8mo agoGiven to how many people its existence must be incredibly infuriating, it's so odd that it's not being chased down with more haste than pirate bay was. I mean I'm glad it's not, but kinda surprised.
- dewey 8mo agoThe music or movie industry lobby is much more aggressive I’d assume.
- nosafemode 8mo agoThere has been some dns resolver issues, some DNS resolvers wont return the address to the sites like archive.is or sites like Annas Archive
- silcoon 8mo agoMaybe they have a paid account? I don’t think there’s much magic behind
- blast 8mo agoPublications could use watermarking to encode the name of the account an article is being served to, but they don't seem to. I wonder why.