9 ms·
Blocking Internet Archive Won't Stop AI, but Will Erase Web's Historical Record
- xnx 7mo agoDoes Internet Archive have a distributed residential IP crawler program? I would enthusiastically contribute to that. There must be some mechanism to prevent tampering in such a setup.
- progval 7mo agoThe Internet Archive does not, but Archive Team does: https://wiki.archiveteam.org/index.php/ArchiveTeam_Warrior https://wiki.archiveteam.org/index.php/ArchiveTeam_Warrior
- xnx 7mo agoYes! I'm running an instance right now.
- Retr0id 7mo ago> There must be some mechanism to prevent tampering in such a setup. Trivial as long as they terminate the TLS on their end, not yours. So you'd just be a residential proxy.
- gzread 7mo agoNo, IA does everything above board and even honors invalid DMCA takedowns.
- SlinkyOnStairs 7mo agoDevil's advocate: Anyone seeking to limit AI scraping doesn't have much of a choice in also blocking archivists. And it's genuinely not that weird for news organisations to want to stop AI scraping. This is just a repeat of their fight with social media embedding. Sure. The back catalogue should be as close to public domain as possible, libraries keeping those records is incredibly important for research. But with current news, that becomes complicated as taking the articles and not paying the subscription (or viewing their ads) directly takes away the revenue streams that newsrooms rely on to produce the news. Hence the "Newspaper trying to ban linking" mess, which was never about the links themselves but about social media sites embedding the headline and a snippet, which in turn made all the users stop clicking through and "paying" for the article. Social media relies on those newsrooms (same with really, most other kinds of websites) to provide a lot of their content. And AI relies on them for all of the training data (remember: "Synthetic data" does not appear ex nihilo) & to provide the news that the AI users request. We can't just let the newsrooms die. The newsroom hasn't been replaced itself, it's revenue has been destroyed. --- And so, the question of archives pops up. Because yes, you can with some difficulty block out the AI bots, even the social media bots. A paywall suffices. But this kills archiving. Yet if you whitelist the archives in some way, the AI scrapers will just pull their data out of the archive instead and the newsrooms still die. (Which also makes the archiving moot) A compromise solution might be for archives to accept/publish things on a delay, keep the AI companies from taking the current news without paying up, but still granting everyone access to stuff from decades ago. There's just major disagreement about what a reasonable delay is. Most major news orgs and other such IP-holders are pretty upset about AI firm's "steal first, ask permission later" approach. Several AI firms setting the standard that training data is to be paid for doesn't help here either. In paying for training data they've created a significant market for archives, and significant incentive to not make them publicly freely accessible. Why would The Times ever hand over their catalogue to the Internet Archive if Amazon will pay them a significant sum of money for it? The greater good of all humanity? Good luck getting that from a dying industry. --- Tangent: Another annoying wrinkle in the financial incentives here is that not all archiving organisations are engaging in fair play, which yet further pushes people to obstruct their work. To cite a HN-relevant example: Source code archivist "Software Heritage" has long engaged in holding a copy of all the sourcecode they can get their hands on, regardless of it's license. If it's ever been on github, odds are they're distributing it. Even when licenses explicitly forbid that. (This is, of course, perfectly legal in the case of actual research and other fair use. But:) They were notable involved in HuggingFace's "The Stack" project by sharing a their archives ... and received money from HuggingFace. While the latter is nominally a donation, this is in effect a sale. --- I find it quite displeasing that the EFF fails to identify the incentives at play here. Simply trying to nag everyone into "doing the thing for the greater good!" is loathsome and doesn't work. Unless we change this incentive structure, the outcome won't change.
- onetokeoverthe 7mo ago[dead]
- Obscurity4340 7mo agoIt would be better if there was some arrangement the papers could reach with Archive where they just delay the release or wait a week then its part of the archive. That way, news stuff gets paid for when its hot and fresh but then it gets archived and the record is preserved
- maltyxxx 7mo ago[flagged]
- user_7832 7mo ago> But in recent months The New York Times began blocking the Archive from crawling its website, using technical measures that go beyond the web’s traditional robots.txt rules. That risks cutting off a record that historians and journalists have relied on for decades. Other newspapers, including The Guardian, seem to be following suit. I'm a bit surprised I never read about this till now, though while disappointing it is unfortunately not surprising. > The Times says the move is driven by concerns about AI companies scraping news content. Publishers seek control over how their work is used, and several—including the Times—are now suing AI companies over whether training models on copyrighted material violates the law. There’s a strong case that such training is fair use. I suspect part of it might be these corps not wanting people to skip a paywall (whether or not someone would pay even if they had no access is a different story). But this argument makes no sense for the Guardian.
- user_7832 7mo agoI went to Guardian's website to cross check their motto (getting confused with WaPo's motto) and got served this (hilarious? sad?) banner. As if blocking cross website tracking is somehow bad. > Rejection hurts … You’ve chosen to reject third-party cookies while browsing our site. Not being able to use third party cookies means we make less from selling adverts to fund our journalism. We believe that access to trustworthy, factual information is in the public good, which is why we keep our website open to all, without a paywall. If you don’t want to receive personalised ads but would still like to help the Guardian produce great journalism 24/7, please support us today. It only takes a minute. Thank you.
- mocd 7mo agoThe Guardian’s ads asking for contributions have got progressively more desperate. I find their commitment to keeping their site paywall free admirable, but the current almost-begging (and selling off their Sunday paper) has got so intense that it feels like it’s only a matter of time until they introduce some kind of paid content.
- ryandrake 7mo ago
- gzread 7mo agoThis is why archive.is was created. Should we stop trying to hunt down and punish its creator and support it as the extremely useful project that it is?
- philistine 7mo agoThe creator can maintain anonymity. The creator does not deserve to continue being celebrated when they embarked on a DDOS campaign using the traffic of archive.is against a journalist trying to uncover their identity. By these actions, they have shown to be capricious, vindictive, and willing to ensnare their users in their DDOS of others. Whoever they are, they’re terrible.
- MSFT_Edging 7mo agoIf there's ever something a journalist would never ever do, it's destroy someone's life for a headline. Never ever. Totally impossible.
- gzread 7mo agoTheir life is in danger and one particular journalist is making it so
- choo-t 7mo agoWell, if they deserve anonymity, they also deserve to be able to protect it, and they have really few tools against a doxxing, the DDOS was one of them, corrupting the archived article was another, albeit dangerous for their own reputation as an archiver. The crux of the problem was the doxxing, not the defense against it.
- ajam1507 7mo agoYou don’t think leveraging your site to DDOS someone is a problem? Do people not also deserve to be protected from being DDOSed? Do people also not deserve to not have their internet traffic be used to DDOS someone?
- tossandthrow 7mo agoI think media outlets think way too highly of their contribution to AI. Had they never existed, it had likely not made a dent to the AI development - completely like believing that had they been twice as productive, it had likely neither made a dent to the quality of LLMs.
- Freak_NL 7mo agoHow do you think those models get trained? You can only get so far with Wikipedia, Reddit, and non-fiction works like books and academic papers.
- RugnirViking 7mo agoHow does the entire textual corpus of say, new York times compare to all novels? Each article is a page of text, maybe two at most? There certainly are an awful lot of articles. But it's hard to imagine it is much more than a couple hundred novels. There must be thousands of novels released each year
- Freak_NL 7mo agoLike apples to oranges. LLMs are (apparently) massively used to get information about topics in the real world. Novels aren't going to be much help there. Journalism, particularly in written form, provides a fount of facts presented from different angles, as well as opinions, and it was all there free for the taking… Wikipedia provides the scantest summary of that, fora and social media give you banter, fake news, summaries of news, and a whole lot of shaky opinions, at best. Novels give you the foundations of language, but in terms of knowledge nothing much beyond what the novel is about.
- olalonde 7mo agoLLMs can get up to date information from primary sources - no journalists required.
- 7mo ago
- daliliu 7mo ago[dead]
- Havoc 7mo agoAs someone perpetually online it’s also making me rethink that a bit Unless you love walled gardens, doomscrolling and endless AI slop that seems like the fun is over
- stuaxo 7mo agoThe New York Times is awful I want it to be archived so people can see that in the future.
- Archonical 7mo agoI don't read it. Why is it awful?
- lyu07282 7mo agoFrom Manufacturing Consent: > by selection of topics, by distribution of concerns, by emphasis and framing of issues, by filtering of information, by bounding of debate within certain limits. They determine, they select, they shape, they control, they restrict — in order to serve the interests of dominant, elite groups in the society." > "history is what appears in The New York Times archives; the place where people will go to find out what happened is The New York Times. Therefore it's extremely important if history is going to be shaped in an appropriate way, that certain things appear, certain things not appear, certain questions be asked, other questions be ignored, and that issues be framed in a particular fashion." The propaganda in the New York times is especially precious because of how highly respected it is, there never was a war or other elite interest they didn't push along.
- mikkupikku 7mo agoThey have a very long track record of pretending to be independent but actually toeing the government's line at key pivotal moments in history when an independent newspaper is needed the most. Everybody here knows how they helped start the second Iraq war I hope, but that wasn't a one-off fluke. Go back through the major wars in American history and you can find the New York Times championing the cause of war before each of these. World Was 2, they uncritically accepted Walter Durranty letting Stalin ghostwrite for him, specifically w.r.t. Stalin's man-made famine in Ukraine, because America was allied with Stalin. WWI, frequent editorializing of Germans being wild Asiatic savages while the Anglos were good and noble people that Americans owed something to for some reason nobody could explain. Vietnam, they uncritically accepted government reports on the second Gulf of Tonkin incident which never happened and broadly accepted the governments own reports about how the war was going, at least in the early years when it still might have been possible to avoid further engagement. Korean war, they supported the government narrative of communist containment. First Iraq War, they uncritically reported very dubious atrocity propaganda, like the fraudulent "Nayirah testimony" given by the teenage daughter of a diplomat pretending to be a politically uninvolved hospital worker. The pattern here is deference to official narratives at precisely the times when criticism is needed the most.
- VladVladikoff 7mo agoAs a site operator who has been battling with the influx of extremely aggressive AI crawlers, I’m now wondering if my tactics have accidentally blocked internet archive. I am totally ok with them scraping my site, they would likely obey robots.txt, but these days even Facebook ignores it, and exceeds my stipulated crawl delay by distributing their traffic across many IPs. (I even have a special nginx rule just for Facebook.) Blocking certain JA3 hashes has so far been the most effective counter measures. However I wish there was an nginx wrapper around hugin-net that could help me do TCP fingerprinting as well. As I do not know rust and feel terrified of asking an LLM to make it. There is also a race condition issue with that approach, as it is passive fingerprinting even the JA4 hashes won’t be available for the first connection, and the AI crawlers I’ve seen do one request per IP so you don’t get a chance to block the second request (never happens).
- mycall 7mo agoEvasion techniques like JA3 randomization or impersonation can bypass detection.
- noads2000 7mo ago[dead]
- VladVladikoff 7mo agoI am aware, fortunately I haven't seen much of this... yet. Also JA4 is supposed to be a bit less vulnerable to this. Also this is why I really want TCP and HTTP fingerprinting. But the best i've found so far is https://github.com/biandratti/huginn-net https://github.com/biandratti/huginn-net and is only available as rust library, I really need it as an nginx module. I've been tempted to try to vibe code an nginx module that wraps this library.
- andrepd 7mo agoI wonder if it would be practical to have bot-blocking measures that can be bypassed with a signature from a set of whitelisted keys... In this case the server would be happy to allow Internet Archive crawlers.
- b1n 7mo agoArchive now, make public after X amount of time. So, maybe both publisher and archiver are happy (or less sad).
- Zopieux 7mo agoI see the attraction, however this will immediately get abused and this delay increased to unreasonable within months/years of it being introduced. Some kind of slippery slope if you will.
- catapart 7mo agoI'm seeing a lot of comments about how we maintain the status quo, but I'm very interested in hearing from anyone who has conceded that there is no way to stop AI scrapers at this point and what that means for how we maintain public information on the internet in the future. I don't necessarily believe that we won't find some half-successful solution that will allow server hosting to be done as it currently is, but I'm not very sure that I'll want to participate in whatever schemes come about from it, so I'm thinking more about how I can avoid those schemes rather than insisting that they won't exist/work. The prevailing thought is that if it's not possible now, it won't be long before a human browser will be indistinguishable from an LLM agent. They can start a GUI session, open a browser, navigate to your page, snapshot from the OS level and backwork your content from the snapshot, or use the browser dev tools or whatever to scrape your page that way. And yes, that would be much slower and more inefficient than what they currently do, but they would only need to do that for those that keep on the bleeding edge of security from AI. For everyone else, you're in a security race against highly-paid interests. So the idea of having something on the public internet that you can stop people from archiving (for whatever purpose they want) seems like it's soon to be an old-fashioned one. So, taking it as a given that you can't stop what these people are currently trying to stop (without a legislative solution and an enforcement mechanism): how can we make scraping less of a burden on individual hosts? Is this thing going to coalesce into centralizing "archiving" authorities that people trust to archive things, and serve as a much more structured and friendly way for LLMs to scrape? Or is it more likely someone will come up with a way to punish LLMs or their hosts for "bad" behavior? Or am I completely off base? Is anyone actually discussing this? And, if so, what's on the table?
- titzer 7mo agoYou're going to hate this, but one answer might be blockchain. A crytographically strong, attestable public record of appending information to a shared repository. Combined with cryptographic signatures for humans, it's basically a secure, open git repository for human knowledge.
- catapart 7mo agoSounds interesting, but I guess I'm a little unsure of how to connect the dots? Are you suggesting that websites would be hosted on a blockchain and browsed by human-signed browsers? Or more like there would be a blockchain authority, which server hosts could query to determine if a signature, provided by their browser, is human? Would you mind painting the picture in a little more detail?
- ryguz 7mo ago[flagged]
- rkwtr1299 7mo agoThe EFF has a lukewarm stance on AI, but criticizes everyone else. AI is clearly ruining the Internet and the job market. How about thinking about your mission and take an anti-AI hardliner stance? But I see multiple corporate sponsors that would not be pleased: https://www.eff.org/thanks https://www.eff.org/thanks All these so called freedom organizations like the OSI and the EFF have been bought and are entirely irrelevant if not harmful.
- rdiddly 7mo agoWhen you disappear from the historical record, that's called you becoming irrelevant. The world moves on, and pays attention to someone else. Not sure why the Times doesn't seem to see this angle.
- paseante 7mo ago[dead]
- lich_king 7mo agoI am really tired of this kind of moralizing. The reality is that every time geeks come up with some utopian ideal, such as that we should publish all our software under free licenses or make all human knowledge freely accessible to anyone, the same geeks later show up and build extractive industries on top of this. Be a part of the open source revolution... so that you do unpaid labor for Facebook. Make a quirky homepage... so that we can bootstrap global-scale face recognition tech. Help us build the modern-day library of Alexandria... so that OpenAI and Anthropic can sell it back to you in a convenient squeezable tube. Maybe it's time to admit that the techie community has a pretty bad moral compass and that we're not good stewards of the world's knowledge. We turn lofty ideals into amoral money-making schemes whenever we can. I'm not sure that the EFF's role in this is all that positive. They come from a good place, but they ultimately aid a morally bankrupt industry. I don't want archive.org to retain a copy of everyone's online footprint because I know it be used the same way it always is: to make money off other people's labor and to and erode privacy.
- Planktonne 7mo agoAgreed; again and again, we see that the utopian ideals of the tech world are only the ones that let them extract value without consideration.
- netsharc 7mo agoWhere your argument falls apart: > the same geeks Proof?
- alexpotato 7mo agoAs someone who did a lot of work on early spam fighting only to see it replaced by things like DKIM, I wonder if we are going to start having the "taxi medallion" style approach but for people connecting to your site. e.g. IA will publish out signed https requests with their key so you, as the site owner, can confirm that it is indeed from them and not from AI. Feels like that would be very anti open internet but not sure how else you would prove who is a good actor vs not (from your perspective that is).
- m3047 7mo agoI'll tell you what I expect to see from crawlers, agents and which I'm enforcing on everybody who doesn't look distinctly human: * Reverse DNS which points to a web site which has a discoverable / well-known page which clearly describes their behavior. * Some sort of reverse IP based, RBL and SPF -inspired TXT records which describe who, what, when, why, how, how often so that I can make automated decisions based on it. Yah, I don't have a lot of crawlers that I welcome... but I'm building a pretty good database of the worst offenders. At scale... there are advantages to scale which work in my favor, actually. I documented this at the end of a blog post when I made blocking Amazon incoming requests a default policy several years ago.
- ashwinnair99 7mo agoWe're essentially burning the library to punish the arsonist. The arsonist already left.
- tremon 7mo agoWhat do you mean, "the arsonist already left"? Isn't it more accurate to say that 90% of the library's visitors are arsonists?
- pamcake 7mo agoIt is not accurate. A very small number of actors pose as many and make up the majority of traffic. For example, your User-Agent block may cut traffic by 10%, 99% of which malicious - but you blocked 1000 individuals, only 1 of which malicious.
- deleted 7mo ago[deleted]
- paseante 7mo ago[dead]
- phendrenad2 7mo agoDoes IA use a known set of IPs? Should be trivial to let them through. But yeah, news companies aren't technically capable of this kind of finesse, they probably have by-the-hour contractors doing any coding/config changes, and closing the ticket is the goal there.
- neilv 7mo agoI'm now an AI bro, and a long-time fan of the EFF (though they occasionally make a mistake). I think this EFF piece could be more forthright (rather than political persuasion), since the matter involves balancing multiple public interest goals that are currently in opposition. > Organizations like the Internet Archive are not building commercial AI systems. This NiemanLab article lists evidence that Internet Archive explicitly encouraged crawling of their data, which was used for training major commercial AI models: | News publishers limit Internet Archive access due to AI scraping concerns (niemanlab.org) | 569 points by ninjagoo 34 days ago | 366 comments | https://news.ycombinator.com/item?id=47017138 https://news.ycombinator.com/item?id=47017138 > [...] over a fight that libraries like the Archive didn't start, and didn't ask for. They started or stumbled into this fight through their actions. And (ideology?) they also started and asked for a related fight, about disregard of copyright and exploitation of creators: | Internet Archive forced to remove 500k books after publishers' court win (arstechnica.com) | 530 points by cratermoon on June 21, 2024 | 564 comments | https://news.ycombinator.com/item?id=40754229 https://news.ycombinator.com/item?id=40754229
- charcircuit 7mo agoThe EFF is being obtuse. Using archives sites is a known bypass for reading news articles for free. Every time a paywalled site someone posts an archive link so others can read for free. >Archiving and Search Are Legal But giving full articles away for free to everyone is not. Archive.org has the power to make archives private.
- m3047 7mo agoIf you're selling ammonium nitrate and diesel, it's a reasonable presumption that you're in the agricultural supply business. It's also reasonable to expect you not to sell a truckload of both to someone who you don't know to be a farmer.
- EchoReflection 7mo ago[dead]
- lzhgusapp 7mo ago[dead]
- pugchat 7mo ago[dead]
- alyandon 7mo agoIt seems that AI-scraping hysteria/paranoia has reached a point that more and more resource sites completely block me without any recourse (e.g. having an account) if my traffic is coming from a datacenter IP. I really wish sites blocked agents/IPs based on actual bad behavior instead of reflexively blocking by IP != residential/mobile. I do understand the frustration though. I've had to deal with bad bots scraping every crevice of a few web based services I host and while most went away after I tweaked robots.txt there were a few I've had to blanket ban IPs assigned to entire ASNs because the bots on their networks refused to play nice.