4 ms·
This is absolutely something that's happening. There are even paid scraper API's that offer "Wayback Machine fallback" as a feature.
by packetslave 17d ago
This is absolutely something that's happening. There are even paid scraper API's that offer "Wayback Machine fallback" as a feature.
- bsimpson 17d agoIt's an open secret that you can often circumvent paywalls by searching Wayback.
- gambiting 17d agoEvery single paid article linked on HN has the way back machine link as the very first comment.
- ValentineC 17d agoThe links are usually to Archive.today (aka archive.ph and a bunch of other domains), not Wayback Machine (which is run by Internet Archive).
- petcat 17d agoehh it's a distinction without a difference. The point is that alternative links are available to circumvent paid access for anyone that wants them.
- organsnyder 17d agoThey're different sites, with different goals, run by different people.
- petcat 17d agoThat provide the same functional service.... Hence, distinction without a difference.
- fluffybucktsnek 17d agoGiven that the root of the discussion is about Internet Archive being hit with huge traffic and not the functionalities provided by Wayback Machine, it very much is a distinction with a difference.
- petcat 17d agoBot traffic or human traffic doesn't matter. The goal is to read websites without having your own access. So Internet Archive, Archive.today, Archive.ph, etc. are all just means to the same end.
- HDBaseT 17d agoI think you have the wrong impression of the Internet Archive. The internet archive is not designed to circumvent anything. It is not designed to "grant access without having your own access".
- fluffybucktsnek 17d agoInternet Archive's traffic may not matter to you, but that's the main topic of this discussion, regardless of what you care or use website archival tools for.
- publlus_enigma 17d agoI suspect you may be conflating two different things. Archive.org exists to preserve historical snapshots of the public parts of websites, and not to bypass subscriptions or pay walls.
- celsoazevedo 17d agoThey are 2 different services, run by different people, one goes out of their way to bypass paywalls while the other doesn't, one is banned by Wikipedia and the other isn't, etc. I think it's a distinction worth making. Not to mention that the Wayback Machine itself isn't exactly a good tool to bypass paywalls as most paid sites don't let them archive paywalled content anyway.
- sandcat_ 17d agoThat isn’t the point being discussed. The point being discussed is that it’s bad form to abuse a service (archive.org) that is provided for free, for the public good in order to run commercial scraping operations.
- petcat 17d agoIt's bad form to scrape the scrapers?
- sandcat_ 17d agoYes, arguably, and for reasons I already gave. I’d genuinely spend a bit more time reading and thinking rather than replying. Your replies are pithy but you’re missing details and frankly making cognitive mistakes. (Apologies if this seems harsh, I don’t mean it as an insult, but this thread has blown up entirely unnecessarily- and yes, I know I’m not helping either!)
- petcat 17d agoYou seem to think that scraping websites "for the public good" is somehow different than scraping websites for any other reason. The end result is exactly the same.
- DaSHacka 17d agoThe minuscule traffic generated by the wayback machine, which serves to preserve the content for years to come, is completely incomparable to the scrapers that hammer every single href linked on a website.
- sippingabonedry 17d agoYou missed the XCancel flamewar yesterday. The consensus is we should be allowed to scrape data and bypass login walls if we dislike the site owner, or feel we are owed free access by arbitrary criteria, it's sort of an unwritten rule. Unless of course it's Google or Meta properties we're talking about because that might impact RSUs.
- eek2121 17d agoCorrect:Also, archive.* has actively edited archived sites to promote their agenda. Why folks continue to use them confuses me. One would think the big wikipedia purge would curb such behavior.
- DaSHacka 17d agoIronically, your framing of the situation is infinitely more disingenuous to push a personal agenda versus anything the archive.today guy did.
- normie3000 17d ago> Why folks continue to use them confuses me. I use them. I haven't ever heard mention that the content is edited. Do you have a source?
- uqers 17d agoSee https://arstechnica.com/tech-policy/2026/02/wikipedia-might-blacklist-archive-today-after-site-maintainer-ddosed-a-blog/ https://arstechnica.com/tech-policy/2026/02/wikipedia-might-... https://en.wikipedia.org/wiki/Wikipedia:Archive.today_guidance#Why_are_we_doing_this https://en.wikipedia.org/wiki/Wikipedia:Archive.today_guidan...? Besides tampering with content, the site was also using visitors to DDOS a blog that mentioned the owner of archive.today.
- sam_lowry_ 17d ago[dead]
- Mogzol 17d agoSee the "Background" section of the Wikipedia RFC on banning archive.today links: https://en.wikipedia.org/wiki/Wikipedia:Requests_for_comment/Archive.is_RFC_5 https://en.wikipedia.org/wiki/Wikipedia:Requests_for_comment... They bulk replaced one string (a name) with another one across many archived pages, and added malicious code to all archive pages that would rapidly send requests to gyrovague.com in an attempt to DDOS them.
- Anonyneko 16d agoWhich sucks because these just don't work for me for some reason (Finland, no luck with VPNs either).
- zymhan 17d agoOnly some of them, it is not universal.
- ghostly_s 17d ago"often"
- koolala 17d agoOne site was doing that which archive in their name but wasn't apart of archive.org
- sam_lowry_ 17d agoarchive.is or archive.today? Why being shy in the era of stealing AI?
- LoganDark 17d agoarchive.today uses clients to perform DDoS, I would not recommend using their site.
- schnebbau 16d agoLet's all just believe this baseless assertion shall we. You have to include more than that, this isn't common knowledge.
- LoganDark 16d ago<https://en.wikipedia.org/wiki/Wikipedia:Archive.today_guidance#Why_are_we_doing_this? https://en.wikipedia.org/wiki/Wikipedia:Archive.today_guidan...> for those without a search engine.
- DonHopkins 16d agoOr rather for those like schnebbau pretending they don't have access to a search engine.
- x______________ 16d agoSure it is! This has been going on for years and global attention was gained at the beginning of this one.[0] Wikipedia deprecates Archive.today, starts removing archive links (arstechnica.com) 616 points by nobody9999 6 months ago | hide | past | favorite | 368 comments 0 https://news.ycombinator.com/item?id=47092006 https://news.ycombinator.com/item?id=47092006