14 ms·
An update on Wayback Machine access
- xbar 18d agoThank you for the Wayback Machine. It is immensely powerful for good.
- Onavo 19d agoWhy not just offer a paid endpoint for the crawlers? It's not like the demand is going to go away anytime soon. It serves nobody except CloudFlare and hardware companies when one side set up blockers and the other side spend money putting VPN SDKs in consumer TVs. I am also curious how the (Russian?) paywall bypass mirror archive.is is doing given that they are probably subject to similar amounts of traffic.
- croes 19d agoIt’s one thing to archive other companies content, it’s another to sell the access to it
- Onavo 19d agoThat's for the lawyers to sort out, they have a lot of flexibility as a US nonprofit. The case law isn't that clear cut for this.
- simonw 19d agoInternet Archive was almost destroyed by a copyright lawsuit from book publishers within the last few years. I expect they aren't excited to take on any additional risk of similar lawsuits right now.
- quotemstr 18d agoThey brought it on themselves by marketing a read-for-free product
- celsoazevedo 19d agoThey need access to sites to archive them. It's already hard to do it as it is, imagine if they start selling access to content. They'd be shooting themselves on the foot, independently of what the law says.
- xp84 19d agomajor [citation needed] on that. There are very limited exceptions to the massive power of copyright -- and they're mainly granted to libraries in the form of narrow waivers. And just the cost of fighting the most powerful copyright holders can bankrupt you -- especially if you're a relatively modestly-funded nonprofit.
- faefox 19d agoYeah, who does the Internet Archive think it is, (insert literally any AI company here)?
- bonoboTP 19d agoWhich AI company is selling access to reliable verbatim copies of websites? I don't mean "it may regurgitate a paragraph", but as a reliable service where you can repeatably get website content snapshots to a reliability level that makes such a use case viable? Using the information for training purposes is not the same thing. Not legally the same and otherwise.
- drdexebtjl 19d agoSites would just block the Internet Archive crawler as well.
- xp84 19d agoMy guess? Because even with a paid endpoint, the type of unscrupulous yahoo that is DDOSing IA today would probably still abuse the free endpoints because they can. The revenue that might come from a paid endpoint could help to scale up, but with how slow IA usually seems, I suspect there is an upper limit to how much traffic they can serve without a LOT more revenue. This is a major "this is why we can't have nice things" situation in my opinion. IA is one of the most valuable gems of the Internet. The only thing that even comes close to preserving our shared history. The damage being caused (both by the effective DDOSing and by the knock-on impact that abuse has in encouraging publishers to remove their content from the archive) is incredibly serious.
- katatue 18d agoIA might be large enough to earn consideration, but generally scrapers just don't care about being good citizens. I work in the GLAM space and we offer OAI-PMH interfaces for the harvesting of our collections data - which doesn't stop companies from preferring to scrape our website for worse (less complete, less structured, less standardized) data instead.
- mitxela 18d agoDo they know it exists? On my site, some types of blocked bots are getting plain-text instructions saying why I'm blocking them and what they can do instead - and it seems to have worked in some cases.
- imglorp 18d agoMicropayments would solve so many Internet problems. It's not too late to adopt. Content creators could charge by page instead of depending on malware/ad/surveillance revenue. Spam is cut if there's a charge per mail. Scraping abuse goes away, along with a bunch of DDOS garbage. The impact is a few cents per page or mail, negligible for a human. But if you're consuming a trillion pages per day, you'd reconsider.
- novok 18d agoMicropayments are blocked by government money laundering regulations increasing the costs significantly to make them untenable.
- Analemma_ 18d agoMicropayments would solve all the problems except for the problem that people absolutely loathe micropayments. Like, vein-popping furiously hate them. Whenever the topic of micropayments for internet content comes up, a bunch of people start talking about payment processors and their floor on prices, and so on. That's not wrong, but it can be designed around and I think it's a scapegoat to avoid confronting the fact that users despise micropayments and we'd rather blame credit card companies for the lack of adoption.
- mindcandy 18d agoMicropayments would solve so many problems for the internet. And, cryptocurrencies would solve so many problems for micropayments. But, it's a non-starter because any proposal gets flooded with people popping veins about how crypto can't solve anything.
- mitxela 18d agoIf it's so easy to solve, go solve it. Set up a test site with micropayments. Maybe scrape CNN and see how many people will pay you micro for a copy of CNN.
- imglorp 18d ago
- KPGv2 18d ago> Why not just offer a paid endpoint for the crawlers? Because then you're definitely violating US copyright law. There are four prongs of fair use analysis, and one of them is the "nature of the use." In this case, you'd be turning into a commercial use.
- Ajedi32 18d agoWhat if you're not charging for the content, but as compensation for the network bandwidth / server resources consumed by serving that content? The idea isn't to profit from content (the IA is a nonprofit anyway), just to allow the IA to continue to serve its purpose as an archive of public data without being overwhelmed by bots.
- oasisbob 18d ago> It serves nobody except CloudFlare and hardware companies when one side set up blockers and the other side spend money putting VPN SDKs in consumer TVs The problem with this perspective is that it ignores the victimization which is happening to all sorts of sites right now. On one hand, you have content owners/suppliers which are trying to place restrictions on how much free bulk use is allowed. When scrapers go to exotic lengths to evade the blocks, eg by using thousands of ephemeral IP addresses to collect an entire corpus, saying stuff like that makes it sound like it's all a wash. "Oh, what a silly situation... How did we ever end up like this? It's not good for anyone ..." No, there is a victim trying to defend themselves from rampant theft of resources, and a corporate asshole which doesn't care about the effects of their actions.
- simonw 19d ago> Here’s what’s going on. The Internet Archive’s Wayback Machine has been hit by waves of high-volume automated traffic, and we’ve put protections in place to keep the service running. I'm pretty certain this is scrapers that are trying to workaround blocks on accessing original sites by hitting the Wayback Machine copy instead. Appalling behavior. In addition to the load it puts on this vital non-profit piece of Internet infrastructure, we've also already seen some sites opt out of the Wayback Machine to prevent their content from being scraped via this alternative route.
- throwawayk7h 18d agoPerhaps it would be sensible for the wayback machine to not show paywalled articles for the first, say, 3 months.
- packetslave 19d agoThis is absolutely something that's happening. There are even paid scraper API's that offer "Wayback Machine fallback" as a feature.
- bsimpson 19d agoIt's an open secret that you can often circumvent paywalls by searching Wayback.
- gambiting 19d agoEvery single paid article linked on HN has the way back machine link as the very first comment.
- ValentineC 19d agoThe links are usually to Archive.today (aka archive.ph and a bunch of other domains), not Wayback Machine (which is run by Internet Archive).
- 18d ago
- CqtGLRGcukpy 19d ago> We’re getting better at telling abusive bots apart from the people who depend on the Wayback Machine every day. If you think you were blocked in error, email info@archive.org with your operating system, browser, and IP address, and we’ll look into it.
- swingandamiss 19d ago[flagged]
- dallen33 19d agoYeah cuz X is fucking shitty, why would I want to give them any traffic?
- qwerpy 19d ago“It’s ok to do bad things to people/things I don’t like” Feels good when you get to dish it out doesn’t it?
- mitxela 18d agoif kidnapping is so bad why do we put the Unabomber in prison
- xp84 19d agoThen... don't? If it sucks so much why do you need to read the tweets? Great take: "This private website is owned by a man I don't like, so I refuse to pay for it - or even give it the possibility to monetize my traffic with ads!" Still quite mainstream take: "... so I'll use an adblocker on it" Immature take: "This private website that I hate and boycott is also an important part of our culture, but the posts on it are too important and valuable to ignore, so I'll use a proxy to scrape it"
- ImPostingOnHN 18d agoYou're confusing the site for the content on it. Some of the content is good, the site sucks and is run by a guy who seig-heils crowds. Even if the content sucked, your post has big "you want to improve `X`, yet you participate in `X`"[0] energy. 0 – https://kitzy.com/content/assets/images/we-should-improve-society.png https://kitzy.com/content/assets/images/we-should-improve-so...
- 18d ago
- tech234a 19d agoI wonder if they'll end up behind Anubis at some point. I'm surprised it hasn't happened already.
- stickfigure 19d agoPlenty of threads on HN about this, Anubis does not work.
- autoexec 18d agoIt always seems to keep me, a normal human, locked out of any site that uses it.
- phendrenad2 18d agoPlenty of threads saying it works, too.
- danbolt 18d agoI’ve read a few different experiences with hosts having had success with Anubis to cut down on excessive scraping. One that comes to mind is the Dolphin project.[1] I’ve read a few of those threads, but often it’s people at cross-purposes with the goals of Anubis. Is there a chance you could clarify the “not working” bit? [1] https://dolphin-emu.org/blog/2025/06/04/dolphin-progress-report-release-2506/ https://dolphin-emu.org/blog/2025/06/04/dolphin-progress-rep...
- stickfigure 18d agoFrom a few weeks ago: https://news.ycombinator.com/item?id=49500040 https://news.ycombinator.com/item?id=49500040 Basically, the cost of an optimized solution is orders of magnitude lower than the cost of an in-browser solution. Anyone dedicated can easily afford to solve workloads higher than your users will tolerate. You might stop casual scrapers, but you're not going to stop someone who cares. AI scraping companies care.
- 18d ago
- BeetleB 19d agoWow, but I wonder if there's more to it. I've not been able to access web.archive.org from my work computer - I always get the 429 error. But I then pull out my phone and can access it just fine. All along I was assuming my company was blocking it. Still weird that it happens every time from my work PC and never from my home one.
- dotmanish 19d agoCould be due to some scrapers from either your work ISP block, or the larger block which lends IPs to multiple workplaces.
- flexagoon 19d agoI assume that's because the IP range of your company network overlaps with a range used by some scrapers, and if it doesn't happen on your phone even in the corporate network, then IA probably checks some extra signals like the user agent in addition to the IP
- BeetleB 18d agoNo - my phone is not connected to work's WiFi. Wonder who the bad actors in my company are...
- flexagoon 18d agoDoesn't have to be someone at your work, it could be a block on a whole ISP network or at least an IP subnetwork that is shared between many clients You can try emailing the address mentioned in their post so they adjust their filters to match just the bot networks more precisely
- ButlerianJihad 18d agoYou should file a support ticket with your manager and the IT security or support desk. Show them the evidence of 429s that are blocking your assigned tasks during working hours. Also include the screenshots and files that you downloaded on your personal device in order to access your work-related materials. Be sure and thank them for adequately configuring the MDM on your personal mobile device so that you could do these work-related tasks. You should definitely also file an expense report to request reimbursement of your personal mobile bill, any data charges incurred, and the hours of networking or collaborating with external colleagues, while you were working on these work-related projects with your personal device.
- timpera 19d agoI really appreciate the Archive team's efforts to make the Wayback Machine more responsive, and have donated a few times to support them. Unfortunately, the restrictions have been way too strict for the last few months: from my residential IP, simply moving the mouse on the calendar for a specific URL is enough to get stuck on 429 error messages for a while; and from corporate ISPs (for example, on airport WiFi), you often can't access the WM at all. I hope they'll find a way to relax those.
- mitxela 18d agoI wonder if your browser is prefetching every link you move the mouse over.
- timpera 18d agoI don't think so, but moving the mouse over any date in the calendar makes a request for the list of snapshots taken on that day.
- account42 18d agoI have had similar experiences just opening archived pages with a couple of embedded images. It's at a level where just using the site normally is painful.
- mrweasel 18d agoGenerally speaking I feel like detecting the bots might be a lost cause. For someone like the Internet Archive I don't know how to deal with it, for smaller sites, cache everything, static pages whenever possible. Sadly I see rate-limiting usage in general becoming a thing. With residential proxies and more sophisticated bots either pretending to be Chrome or directly piloting Chrome, it's going to become impossible to tell a real user from a bot. Only solution is to pretend that everyone is a bot and design for it.
- markalby 18d agowe’ve been using Datadome at work for this because it’s nice to get someone else to think about the constant bot cat and mouse, and we all benefit from rules and fixes created from other client data. Not an ad- that service is eye watering expensive but I think it makes sense as something to offload.
- UltraSane 19d agoWhy not put it in S3 with downloader pays?
- Kayvanian 19d agoAs a public resource the hope is for Wayback to be free to access. I imagine putting up a paywall would be their last resort.
- charcircuit 18d agoS3 price gouges on bandwidth.
- mitxela 18d agoAnd storage. $23 per TB per month, and $90 per TB downloaded, is highway robbery.
- lousken 19d agoAI companies should pay billions to wayback machine for access
- KPGv2 19d agoI think that'd raise serious copyright concerns, if the Wayback machine started selling other people's intellectual property.
- roblh 18d agoShouldn’t it follow that it’s illegal for the AI labs to profit off of all of that stolen copyrighted data too?
- deleted 18d ago[deleted]
- Joel_Mckay 18d ago[flagged]
- mitxela 18d agoIt should, but it doesn't.
- KPGv2 13d agono. The topic at hand is distribution of copyrighted material, and you're talking about reproduction and possibly preparation of a derivative work. At last in the USA, copyright law defines specific things copyright owners have exclusive rights to: - reproduction - preparation of derivative works - distribution of copies to the public - public performance - public display The most immediate issue with AI companies is whether they've made infringing reproductions. The other possibility is the preparation of derivative works: does an AI response count as derivative of something it's consumed? Sorry I'm not going to do the analysis for you, though. I'm no longer a bright, chipper IP law scholar.
- jMyles 18d ago
- unkeen 19d ago[flagged]
- stronglikedan 18d agoYes, that's one acceptable alternative, and another commonly accepted alternative is API's. Although, I'm not sure why you included the asterisk.
- maxrev17 18d agoUnkeen on the apostrophe that’s why! Gotta keep HN proper and correct guize
- tomhow 18d agoWe detached this subthread from https://news.ycombinator.com/item?id=49716735 https://news.ycombinator.com/item?id=49716735 and marked it off topic.
- xyst 19d ago[flagged]
- alex1138 18d agoI mean there are people who have reported that with their own personal website Facebook's crawlers were essentially DDOSing them
- gooeyblob 18d agoWhat reason do you have to doubt the claim?
- plorkyeran 18d agoIf you're in a room with a TV then literally yes, there's a good chance there's an abusive bot in the room.
- ignoramous 18d agohttps://archive.vn/WWENg https://archive.vn/WWENg
- MattCruikshank 18d agoThere was a feature on Amazon Web Services for a while, and I wish it was still there... Downloader pays. I make some content and upload it. When you want to download it, you pay Amazon the egress fees. And maybe I get to charge just a bit more, to help me with the Ingress, storage, content creation, etc. I mean, I know that there's going to be problems with rate limiting, etc. And yes, we have those problems with LLM tokens today. But this just feels like such a useful thing that it baffles me that it doesn't exist already.
- mitxela 18d agoNote the Amazon egress fee is one hundred times anywhere sane's egress fee.
- MattCruikshank 18d agoMy desired usage pattern stands... Someone who publishes content shouldn't be punished for everyone else wanting to access it, and shouldn't have to resort to product placement, advertising, sponsorship, or begging to fund it. I don't know, maybe WebTorrent should have been the answer? For upcoming, viral content? But for the deep archives, like the Wayback Machine? I feel like I'd happily pay for egress, and a bit to support them. If it was automatic and built in... I wish Flattr or something like it had thrived...
- mitxela 18d agoThere was MegaUpload. It got shut down because it was used almost exclusively for piracy.
- vlyan 18d agounrelated: if a website gets hit with "This URL has been excluded from the Wayback Machine", do existing snapshots get purged or may they still be preserved somewhere?
- msephton 18d agoThey get marked as inaccessible, but still exist in IA data
- vlyan 18d agois it possible to access somehow? it seems the site got excluded because of robots.txt set by some domain squatter, not manually.
- msephton 18d agoGet a job at IA?
- alightsoul 18d agoThey have all the WARC file cataloged on their site outside the way back machine
- tech234a 18d agoSee also: https://wiki.archiveteam.org/index.php/List_of_websites_excluded_from_the_Wayback_Machine https://wiki.archiveteam.org/index.php/List_of_websites_excl... Note that Archive Team is separate from the Internet Archive.
- basilikum 18d agoMad props to the people at the Archive. You are the heros we need in a formerly open internet that is surrendering to evil big corps and closing down free access. The Internet Archive is in a really bad spot being attacked from multiple sides at once. But — while service has not been consistent — they have maintained open access. I can still access anonymously from Tor without Cloudflare or some other centralized gatekeeper showing me the middle finger. If you got some money to spare, consider donating to them. They need it.
- superxpro12 18d agofully expect them and wikipedia to get assaulted by AI companies to monopolize data source access in the near future. The future is bleak :\
- mrguyorama 18d agoI'm expecting a full court legal attack on Wikipedia at some point. It's just too good for information, and companies would rather you use their chatbot to regurgitate that info now that they have their own copies. Google was already built to pretend as if they had some magic system giving you "Answers" that were 95% just the text of the infobox they used to have for wikipedia on the side. It's going to suddenly be extremely valuable that wikipedia didn't settle for having a small rainy day fund and instead ceaselessly grabbed every fucking donation they could for two decades so they can fight such a legal battle.
- hubraumhugo 18d agoThere is a HN article on abusive AI crawlers on the front page almost every week, but we rarely talk about the path forward. Web scraping has been around for as long as the internet, and it was fine because we had established best practices (rate limiting, self-identification, robots.txt, etc.) that the industry agreed upon. Now we have AI labs and their crawlers that don't care about any of this gentlemen's agreement. So how do we go from here? Is adding more difficult Anubis and Cloudflare bot protection really the solution? How many millions of human hours and billions in infra costs are we willing to spend on this arms race? Some approaches that I think are promising: - A robots.txt V2[0] as a standard way for website owners to state how bots and AI crawlers can use their online content and where to go (e.g. distinguish search from AI training use cases, point to a downloadable file instead of crawling everything, etc.). - Something like Web Bot Auth[1] as a non-centralized standard for self-identifying bots and agents cryptographically. This would allow websites to allow or deny bots very precisely. - what else? [0] https://datatracker.ietf.org/doc/draft-vaughan-machine-reada https://datatracker.ietf.org/doc/draft-vaughan-machine-reada... [1] https://datatracker.ietf.org/doc/html/draft-meunier-http-mes https://datatracker.ietf.org/doc/html/draft-meunier-http-mes...
- maxrev17 18d agoYeah it’s kinda crazy to me that what was once a back alley python script is now accepted as ‘fine, free for all’. The new era of bros really are smth else.
- edelbitter 18d ago- Find some new way for Cloudflare to acquire paying customers. If their business did not depend on the status quo, they would be exceptionally well positioned to roll out the technical & organizational frameworks that that make massive botnets a thing of the past.
- msephton 18d agoI've been getting this error a lot. Asking users to email them with details of their OS, browser, IP address is just crazy. Their support is supposedly already swamped and they are asking for more!? Changes made by IA shouldn't become my responsibility.
- jolmg 18d ago> Asking users to email them with details of their OS, browser, IP address is just crazy. It's surely to serve as data to help tell humans apart from bots. > Changes made by IA shouldn't become my responsibility. They're a free service. It's ultimately not their responsibility to service you either.
- msephton 18d agoImagine if Apple or Microsoft introduced a bug and said, ah yes we know about it we did that on purpose and we know it affects a huge number of people, if each of you could email us these details that'd be great. It's just such an insane request. IA have broken it and have no real idea how to make it better so they are going to whitelist IPs or browsers or entire operating systems? Wild.
- jolmg 18d agoNo, it's more like you're requesting something from them and they're telling you they may need some technical, non-personally-identifiable info from you to fulfill your request.
- kjs3 18d agoWhat is 'insane' here is the shear level of entitlement displayed here, including lumping a niche, free, volunteer supported service in with billion dollar, for profit corporations and demanding they pander to your inflated expectations. Wild.
- msephton 18d agoIt doesn't matter who or what the service is, how much they have, or whatever else. They created a problem and now users have to pay for the inconvenience by emailing(!) specific details that could be captured automatically through web logs: OS, browser, IP address. It's ridiculous.
- pelican0 18d agoIs it established that the scraping scourge of late is primarily driven by AI companies? Anyone aware of any relevant studies? Beginning to think that the difficulty to browse most websites nowadays due to throttling, is yet another negative externality of AI development that society is forced to bear.
- userbinator 18d agoIt's not. There is no evidence, just propaganda.
- emaro 18d agoIt's shame that the AI arms race causes such collateral damage. Free resources were always exploited, but the stakes ($T) and capabilities around AI allow unprecedented abuse. I wish we could go back... :/ I really don't see any solution to this; the scrapers probably wouldn't even mind destroying sources like IA too much, which would leave them as the only "authorative" source of knowledge in the end. Best way is likely regulation incl. hefty (!) fines, but politics are too slow and too fragmented to be effective. So... Enjoy it while it lasts, I guess.
- zdragnar 18d agoI'm a little more skeptical that this is "AI is big so it is worse" issue. Yes, AI is big in scale, but this has been the case for almost every popular free service. They either start: - charging (news / journalist services) - gate-keeping (X forcing log-ins) - enshittifying (lots of ads and degraded service) The fact that the way back machine is incredibly useful but most people didn't know about it or use it very much doesn't change the fact that it has basically become very popular... only with LLM agents rather than humans. Ads alone aren't enough to support human traffic for many sites with human traffic.
- thimabi 18d agoI wonder why doesn’t the Internet Archive require logging-in prior to accessing the Wayback Machine. It would probably help them distinguish humans from bots, at a very little cost to humans.
- extralongdivisi 18d agoGatekeeping information is not the solution
- hamandcheese 18d agoWhy not? If its the difference between the information being available at all, then I choose login any day of the week.
- extralongdivisi 18d ago> If its the difference between the information being available at all That's the point. The solution should avoid information not being available. Requiring login will incentivize bots to create spam accounts and move the battle to a new frontier, hurting real people in the process.
- TechSquidTV 18d agoThis really only inconveniences people, not bots.
- mitxela 18d agoDo you have a library card?
- extralongdivisi 18d agoI can walk into a library, pull a book from a shelf, sit down, and read it front to back. No library card needed. Only need one to take a book home. Completely different scenario. Edit:clarification
- int32_64 18d agoAre any AI companies using residential proxies to scrape?
- xena 18d agoYes. It's impossible to tell which because the split is residential proxies, dataset curators, and AI companies all being separate actors. However I fucking guarantee you it's out there and people are too cowardly to be honest about it so they don't get sued out of existence.
- oasisbob 18d agoOh yeah, absolutely.
- xacky 18d agoThe anti virus industry needs to crack down on crawler and proxy malware, plus ISPs FINALLY need to replace CGNATs with iov6 to stop crawlers banning everyone behind a NAT.
- sicktriple 18d agothe root of the root of all evil: NAT
- HDBaseT 18d agoHow exactly does IPv6 "stop crawlers". If anything, it will make it harder to block due to the vastness of the IPv6 address space.
- fulafel 18d agoCGNAT makes all ISP users appear to come from one v4 address, so blocking by v4 address becomes unworkable.
- mrweasel 18d agoAnti-virus, maybe, the ISP definitively needs to step up and just shut off people internet when large amounts of bot traffic is detected. IPv6 is going to do nothing, because right now you're getting scrapped/attacked/DDoS/whatever with a single request from millions of IPs at once.
- brador 18d agoThe only solution is to make visitors do compute. Compressing files for the archive to access other files would be perfect for this. Cross verify hashes to prevent cheating. Ez.
- ilamont 18d agoShouldn't the solution be to gate bulk access for automated services for a price? Not just the wayback machine, any personal or corporate website? My blogs are getting slammed and there are issues with cloudflare or captchas.
- iamacyborg 18d ago> Shouldn't the solution be to gate bulk access for automated services for a price? Fine in theory but determined scrapers will use residential proxies in bulk.
- hamboomger 18d agoThis! But only wayback machine, other services I'm not sure. But maybe the problem is that they can't serve the data from the other websites like this, if they use it commercially. Right now they have non-commercial use, from what I understand.
- mrweasel 18d ago> Shouldn't the solution be to gate bulk access for automated services for a price? The problem is that many of the people who are scraping this data doesn't want to pay. These are organisations who would rather not clone your git repo, and instead scrape every single page on your Forgejo installation. These are NOT nice people.
- josefritzishere 18d ago[dead]
- tgtweak 18d agoCan't wayback machine just offer direct access to the archive for a premium and in doing so, pay for the service?
- edelbitter 18d agoNot while the new dukes of the internet wielding massive armies of hijacked smart TVs have a better time browsing the web than I have; as a mere peasant with just a few IP addresses. There would be no reason to sign up and pay up for bulk access, unless open access is shut down.
- robotmay 18d agoUnrelated, but this week I've been on a memory binge with the Wayback Machine, trying to find old content of mine from the early 2000s. Took me a while but I've finally put together a good bit of info about myself at the time that I'd completely forgotten, and it's all thanks to the Internet Archive storing my little gaming review website from when I was 16. I could barely remember any of the other stuff, it's been genuinely surprising figuring out what I'd forgotten. I couldn't even remember most domains I owned aside from one, which I used as the starting point. Still can't remember what my Tripod site address was, but that might be lost to time. Thank you, Archive.org.
- halfblood_walks 18d ago[flagged]
- msephton 18d agoWhy can't they capture OS, Browser, and IP address at the time of error? All that information is available at the point of failure, the user should not need to email it in.
- ericpauley 18d agoPresumedly they collect that, but the vast majority of blocks are correct and not errors. This info allows them to look up the user’s request to label it as legitimate.
- sicktriple 18d agoAnyone else feel like making a new internet and starting over
- righthand 18d agoPeople will just bring their bots over and you'll be back to square one. Bots scraping existed before LLM companies decided to go nutso on the internet.
- mitxela 18d agoYou can try a higher level of identity verification on the new internet. It shouldn't be fully ID verified, but more like how it used to be - users on a network were anonymous to other networks, but you could email the admin of a network to track down bad behavior with their cooperation if they agreed it was bad. You could even build this as an overlay on the current internet. DN42 is like this.
- righthand 18d agoSo then identity theft and fake identities will sky-rocket.
- petterroea 18d agoI'd be happy to pay a 5$/month donation to get a higher rate limit/more lenient filter put on me
- throwaway456754 18d agoIf you get something, it isn't a donation.
- 1vuio0pswjnm7 18d ago"Here's what's going on." Thank you https://news.ycombinator.com/item?id=49571448 https://news.ycombinator.com/item?id=49571448 I had a feeling it was due to "AI" companies and developers using "agents" Not surprised
- potato-peeler 18d agoWayback can’t be accessed through vpn, atleast on proton. Heck, most sites simply block you for using vpn.
- sehw 18d ago[dead]
- delis-thumbs-7e 18d agoI recently remembered a wonderful comic blog from 2010’s that is not online anymore. It was a sonderful Finnish LGTG-thened comic blog that I use to read, then forgot completely until few weeks ago. WM had it stored of course, so I could read through this amazing piece of internet art again. I really so through some money their way, they do wonderful work.
- userbinator 18d agoThe Internet Archive’s Wayback Machine has been hit by waves of high-volume automated traffic Thank you for not immediately blaming it on "AI bots". I suspect there's some entity manufacturing consent for strong identity/age verification/sanctioned-browser-OS "walled garden" Internet, and these random DDoSes are part of that. I knew something was up when a few alternative YouTube front-ends I use suddenly put up the 'nubis and complained about the high volumes of traffic they were getting flooded with; of course someone actually going after that data would be aiming their "AI bots" at YouTube directly instead of trying to suck it through a tiny little-known proxy-site, so it really strained the credibility of the argument.
- bothers 18d ago[dead]
- roughly 18d agoBonus points for anyone who’d like to guess how the tragedy of the commons was resolved in the times before the enclosure movement.
- RobotToaster 18d agoTorches and pitchforks?
- mitxela 18d agoby the enclosure "movement"? i.e. greedy powerful people walling off everything they could and declaring it was theirs and you'd have to pay a tithe to use it? (I should really start calling rent "tithes" more often)
- deleted 18d ago[deleted]
- mrhcon 18d ago[flagged]
- Roark66 18d agoI think it's a matter of time before archive.org gets "bought" and dissappears. There should be government sponsored mirrors in many places of the world. The amount of data in archive.org is about 100PB. We're talking 10 racks of disks. I think archive.org should sell "archive as a service" for let's say $15mln. Half of that would be hardware cost and the deliverable could be 12 DC racks containing entire archive.org.
- tim333 18d agoThe archive is a non profit funded by various foundations and a a congressionally designated depository for U.S. Government documents. They may have a job selling it off without objections.
- hedora 18d agoThey should sell copies. I’m sure some foundation model company would happily hand them more than enough to establish a self sustaining foundation. Also, then there would be multiple copies. They could even give torrent access to libraries.
- knd775 18d agoIf they did this, many websites would immediately opt out or block them.
- GaryBluto 18d agoI've been doing a bit of (very careful to be polite) scraping of Wayback to get archives of now-offline sites, so I hope this won't cause any significant issues for me. I did contact them beforehand to request a direct copy of the sites/networks required (which I believe they offered at some point) but unfortunately received no reply.
- cranberryjoe 18d agoWeird that a library is restricting free access to other libraries wanting to preserve history. I guess it’s not a library after all.
- Mr_Minderbinder 17d agoIt feels like the Web is breaking. More and more sites are retreating behind Cloudflare, captchas and PoW defences. When reading Wikipedia on my phone, I have noticed for a while that sometimes half the thumbnails on an article fail to load. Non-corporate browsers are even more unusable than before.
- tallow_ghost 13d ago[flagged]
- MollyRealized 13d agoI know I'm coming to this post rather late, but there's actually a legal consequence to this as well. So many attorneys and law firms rely upon reference to the Wayback Machine as legal proof of a website's state as of a certain date. If more and more companies begin to opt out, or if that becomes less and less reliable, there may literally be changes in lawsuit outcomes based on the sudden unavailability of evidence. IANAL, but I am a litigation legal assistant. (Out of work and Chicago-based, if anyone knows of someone looking for a tech-aware assistant.)