4 ms·
> Here’s what’s going on. The Internet Archive’s Wayback Machine has been hit by waves of high-volume automated traffic, and we’ve put protections in place to k
by simonw 19d ago
> Here’s what’s going on. The Internet Archive’s Wayback Machine has been hit by waves of high-volume automated traffic, and we’ve put protections in place to keep the service running.
I'm pretty certain this is scrapers that are trying to workaround blocks on accessing original sites by hitting the Wayback Machine copy instead. Appalling behavior.
In addition to the load it puts on this vital non-profit piece of Internet infrastructure, we've also already seen some sites opt out of the Wayback Machine to prevent their content from being scraped via this alternative route.
- throwawayk7h 18d agoPerhaps it would be sensible for the wayback machine to not show paywalled articles for the first, say, 3 months.
- packetslave 19d agoThis is absolutely something that's happening. There are even paid scraper API's that offer "Wayback Machine fallback" as a feature.
- bsimpson 19d agoIt's an open secret that you can often circumvent paywalls by searching Wayback.
- gambiting 19d agoEvery single paid article linked on HN has the way back machine link as the very first comment.
- ValentineC 19d agoThe links are usually to Archive.today (aka archive.ph and a bunch of other domains), not Wayback Machine (which is run by Internet Archive).
- petcat 19d agoehh it's a distinction without a difference. The point is that alternative links are available to circumvent paid access for anyone that wants them.
- organsnyder 19d agoThey're different sites, with different goals, run by different people.
- petcat 19d agoThat provide the same functional service.... Hence, distinction without a difference.
- fluffybucktsnek 19d agoGiven that the root of the discussion is about Internet Archive being hit with huge traffic and not the functionalities provided by Wayback Machine, it very much is a distinction with a difference.
- petcat 19d agoBot traffic or human traffic doesn't matter. The goal is to read websites without having your own access. So Internet Archive, Archive.today, Archive.ph, etc. are all just means to the same end.
- HDBaseT 18d agoI think you have the wrong impression of the Internet Archive. The internet archive is not designed to circumvent anything. It is not designed to "grant access without having your own access".
- fluffybucktsnek 18d agoInternet Archive's traffic may not matter to you, but that's the main topic of this discussion, regardless of what you care or use website archival tools for.
- zymhan 18d agoOnly some of them, it is not universal.
- ghostly_s 18d ago"often"
- koolala 18d agoOne site was doing that which archive in their name but wasn't apart of archive.org
- sam_lowry_ 18d agoarchive.is or archive.today? Why being shy in the era of stealing AI?
- LoganDark 18d agoarchive.today uses clients to perform DDoS, I would not recommend using their site.
- schnebbau 18d agoLet's all just believe this baseless assertion shall we. You have to include more than that, this isn't common knowledge.
- LoganDark 18d ago<https://en.wikipedia.org/wiki/Wikipedia:Archive.today_guidance#Why_are_we_doing_this? https://en.wikipedia.org/wiki/Wikipedia:Archive.today_guidan...> for those without a search engine.
- DonHopkins 18d agoOr rather for those like schnebbau pretending they don't have access to a search engine.
- x______________ 18d agoSure it is! This has been going on for years and global attention was gained at the beginning of this one.[0] Wikipedia deprecates Archive.today, starts removing archive links (arstechnica.com) 616 points by nobody9999 6 months ago | hide | past | favorite | 368 comments 0 https://news.ycombinator.com/item?id=47092006 https://news.ycombinator.com/item?id=47092006
- luckylion 19d agoWhat sites would they be targeting? Generic "just give me anything"? Whenever I check regular sites on IA, the coverage is spotty -- they'll have the homepage and a few important pages, but it quickly fizzles out. Very understandable, you can't store all 15000 pages of any random website and update them etc etc, but that makes them pretty useless for indirect scraping because you usually don't want a tiny taste, you want everything.
- ajaysingh4651 18d ago[flagged]
- toomuchtodo 19d agoIt is. They will most likely eventually need to move to a walled model for Wayback due to scraper aggressiveness (like Reddit deprecating anonymous old.reddit.com), or behind Cloudflare for aggressive bot and scraping protection. Hard to defend against abuse of a public resource when its intent is public access with as little restriction as possible. https://en.wikipedia.org/wiki/Tragedy_of_the_commons https://en.wikipedia.org/wiki/Tragedy_of_the_commons (no affiliation)
- ronsor 19d agoReddit has no excuses for the anonymous old.reddit.com removal; they're simply greedy. On the other hand, the Internet Archive is a non-profit offering a free public resource.
- toomuchtodo 19d agoExamples provided as technical examples, strong feelings are out of scope for this thread.
- itintheory 19d agoAs someone who operates a large non-profit public data driven website, I have some VERY strong feelings about scrapers. We looked into various commercial solutions (Datadome, HUMAN) and based on our traffic estimates from logs we'd be looking at at least 250k/yr for bot mitigation. Anubis is offering a temporary reprieve, but after reading the recent kernel.org article [0] it's increasingly clear that this is a temporary bandaid. The cheapest solution is to require a login and rate limit by API key. I also have strong feelings about the tragedy of the commons. [0] https://people.kernel.org/monsieuricon/creepy-crawlies https://people.kernel.org/monsieuricon/creepy-crawlies
- toomuchtodo 19d agoNo strong feelings here is what I meant. Certainly, that energy is best directed into aggressive countermeasures and defense in depth of public goods. https://hn.algolia.com/?dateRange=all&page=0&prefix=true&query=author%3Adang%20%E2%80%9Cstrong%20feelings%E2%80%9D&sort=byDate&type=comment https://hn.algolia.com/?dateRange=all&page=0&prefix=true&que...
- bradly 19d agoJust yesterday from my one of my sessions with Sol: > Hacker News and the Rails forum are blocking the text fetcher, so I'm using the browser workflow to inspect the pages directly
- TeMPOraL 19d agoAs it should. Unfortunately I sometimes have to browbeat Claude into acting like an agent of the user is supposed to. Usually it works, though last time it refused to recognize my moral argument (on the grounds that it's not bound to my interests exclusively and needs to protect the interests of its maker too).
- bradly 19d agoDo you think there should be a way for a site to tell an agent it isn't allowed access? I'm not sure where I land on this exactly tbh, so no judgment cast.
- aaron_m04 19d agorobots.txt?
- bradly 19d agoHas it been settled whether robots.txt applies to user-driven chat sessions and if things like the crawl delay should be applied to say an end-user, an ip address, a harness provider, etc? My understanding is robots.txt is more for training exclusions, but less so for agent work.
- xena 19d agoAI bros think they should be exempt from robots.txt. Administrators of big services beg to differ. No solid consensus has arisen. I bet it's gonna take a lawsuit or two to see how it shakes out.
- pantsforbirds 19d agoWe used to always "scrape" the wayback machine for any sort of news article we actually paid to consume. I was absolutely shocked by major news sites making very important edits to an article without any sort of editorial notice! Sadly this sort of thing is probably not really possible anymore, but I can't really blame anyone for making this sort of decision. I can't imagine how much more traffic they get now vs 2021 when we were doing this.
- subarctic 19d agoWhat if they charged money? Is it something you'd pay for?
- msephton 19d agoI'd pay for it, but only if they implemented the changes the community of users have been requesting for years.
- carlosjobim 19d agoNo matter what they did, you'd have a new excuse for why you won't pay.
- msephton 18d agoAh, the old ad hominem attack. How refreshing. But anyway, no, I wouldn't keep finding reasons. I donate to them every year already. Somebody asked if I would be willing to pay and my answer was "yes, but". It would need to be improved because certain aspects of it suck right now, not only the error this post is about. They only need go as far as their forums and github repos to see the community feedback.
- jakderrida 18d agoIs it really an ad hom if he doesn't know the hom? Their reply is 100% based on the content of your post.
- RobotToaster 19d agoDo they offer bulk torrent downloads as an alternative?
- echelon 19d agoI would love to be able to download every page of a given domain as an archive, and I'd pay to do this.
- msephton 19d agoThey provide a free cli tool to do this.
- petcat 19d agoisn't that what wget -m does? what is there to pay for?
- carlosjobim 19d agoYou'd pay the domain owner for it? How much?
- QuantumNomad_ 19d agoOnce upon a time some people explored backing up the Internet Archive. However, that experiment ended. They mention there were some learnings and they then say: > The Internet Archive continues to explore methods and code to decentralize the collection, to have a mirror running in various ways - these include IPFS, FileCoin, and others. The INTERNETARCHIVE.BAK project also added general mirroring and tracking code to a number of projects that are still in use. https://wiki.archiveteam.org/index.php/INTERNETARCHIVE.BAK https://wiki.archiveteam.org/index.php/INTERNETARCHIVE.BAK I would really like to know if any sort of thing like that is still ongoing and if it’s accessible to people in general. Would be nice to mirror some data from IA to my local drives, for example via BitTorrent or IPFS, to have it for offline exploration and personal archive. I know that individual items have torrents. And I’ve downloaded a few that way but always it ends up only using the “web seed” (i.e. the BitTorrent client is retrieving the files from IA via HTTP) because there are no one seeding some random single item I found. Plus, those torrents are unreliable sometimes because they include meta data files that were since updated but the torrent was not updated and so the web seed is giving the updated files that don’t match what the torrent says their hashes should be. So then you have to jump through some extra hoops to fix that and then resume the download, and all the while the HTTP connections to IA servers time out because their servers are overloaded. So when I say I wonder about possibilities of using BitTorrent I mean to retrieve whole collections of many items instead of individual ones, and with actual other peers instead of just having it put load on IA HTTP servers.
- jader201 19d ago> I'm pretty certain this is scrapers that are trying to workaround blocks on accessing original sites by hitting the Wayback Machine copy instead. Appalling behavior. Appalling, yes. But also expected. I'm surprised they haven't been the target of scrapers for years. But sites putting their content behind login walls and other anti-bot mechanisms has certainly exacerbated this. But again, this isn't at all a surprising progression. > we've also already seen some sites opt out of the Wayback Machine to prevent their content from being scraped via this alternative route. To be fair, another big motivation was likely users on sites like HN using archive.org (and similar sites) to get around their paywalls. In fact, I'd be surprised if this wasn't a big motivator. Again, it sucks, but it's not at all surprising to see it progress like this. I wouldn't be surprised to see similar blocks on other archive sites eventually.
- autoexec 18d agoI've personally been using the Wayback Machine more often because I increasingly find myself being blocked from websites who are trying to keep out scrapers even though I'm just a regular person with JS disabled (along with a bunch of other stuff)
- eek2121 18d agoSites are getting too overzealous with blocking IMO. I got blocked for several hours by huggingface simply because my download didn't complete and I had to retry. It gave me error 429, suggested I login, and the login page wouldn't load because error 429. A popular tech news site blocked my phone because of Apple Private Relay. That didn't last long because their traffic fell off a cliff when that happened. Many sites are throwing more captchas at the problem, without understanding that captchas don't actually help with LLMs, they just hinder normal users and primitive scripts. LLMs solve captchas just fine. Some big sites have put up improved paywalls. I'm fine with subscribing to a quality site, however, WSJ and all the other big media sites routinely spit out regurgitated garbage that can be had for free elsewhere (and due to political spin, their garbage is less valuable than the free versions of said content). Some folks are declaring the internet dead. I wouldn't go that far, however, I will say that a reckoning is going to happen, especially when advertisers figure out that most ads served on basically every website are no longer viewed by humans.
- Roark66 18d agoYou download from huggingface without logging in? They are known for throttling not logged in users horribly.
- matt_heimer 18d agoI wonder if the entire internet is going to slowly move behind logins and allow lists for specific trusted crawlers at some point. Open access doesn't seem sustainable. But I might just grumpy about spending another hour this week adjusting rules to prevent bots.
- intrasight 18d agoRather than logins or regional filters, how about they just be a content provider to local libraries and perhaps use an app like Libby.
- TZubiri 18d agoNah, there's actual value in hitting historical versions and with agents the gap between "how long has this product been offered by this company" and "I should go to wayback machine and do a binary search to find the earliest snapshot that contains this product offering " has closed.
- Kodiack 18d agoI run some small websites, including a tiny forum that’s been a goldmine for scrapers. I had to significantly tweak some firewall rules and configuration after scrapers behind residential proxies suddenly accounted for over 99% of requests. However, I also relaxed rules for automated traffic that was well-behaved, and I went out of my way to ensure that the Wayback Machine was able to hit everything. I should kick a small donation their way. They provide an incredibly valuable service and I love the benefit that I get from them just for personal side projects.
- mkatx 18d agoThis is the way to go! Cut the cat and mouse, win win ish.
- sippingabonedry 18d agoHow do you separate the Wayback Machine from malicious bots that pretend to be the Wayback Machine? Are you whitelisting their IP blocks? Because I get a ton of scraper requests that forge Googlebot, Bing, and Yandex user-agents that are totally not coming from their IP ranges. In fact, sometimes they all come from the same IP...
- Kodiack 18d agoThey have their own ASN, which I’ve explicitly allowed requests from. https://www.peeringdb.com/asn/7941 https://www.peeringdb.com/asn/7941
- junon 18d agoTIL they have their own ASN. This is helpful, thanks.
- fc417fc802 18d agoI thought at least google (and possibly others) provided a way to verify the user agent?
- 18d ago
- ezekiel68 18d agoYou might be right but -- why would they have watied until these recent weeks?
- e40 18d agoI say name and shame!
- hedora 18d agoMy use of wayback has skyrocketed recently due to anti-bot measures. I often cannot get past captchas, and archive.org is one of the fallbacks I try. However, archive.is, etc are more reliable. I wish the internet archive acted more like a library system, where multiple organizations could mirror the content. They are a big single point of failure, and I’m shocked Trump/SCOTUS haven’t intentionally burnt the archives down yet.
- account42 18d agoIronically, archive.is itself has a captcha that doesn't like my home FF install.
- simonjgreen 18d agoNearly every time a link is posted to HN to a site behind some form of wall, a high voted comment on the post will be a link to an archive site bypassing the owners wall. Bot owners are not the only ones routinely circumventing the choices of content owners.
- kragen 18d agoI think the Wayback Machine offers an official API for bots to call.
- bothers 18d ago[dead]