6 ms·
AI will kill the internet because it is killing the incentive to make it. It is an industrial-strength example of why we don’t allow stealing.
by ChiMan 2mo ago
AI will kill the internet because it is killing the incentive to make it. It is an industrial-strength example of why we don’t allow stealing.
- xnx 2mo agoI might be more motivated now because at least I know the bots will read it.
- zombot 2mo ago> why we don’t allow stealing. With the not-so-minor qualification that the biggest thieves have always gotten away scot-free. AI is just the international whole-internet version of this.
- smackeyacky 2mo agoBehind every great fortune is a great crime
- HappyPanacea 2mo agoWhat is the great crime behind Norway's sovereign wealth fund?
- inigyou 2mo agoIt comes from selling oil and gas reserves, so they did destroy the environment.
- Nasrudith 2mo agoAh it is amazing how crimes just materalize from nowhere from envy.
- smackeyacky 2mo agoAmazing how they appear when money is involved. Let’s talk bout AI intellectual property theft on a scale never seen before in human history, or Uber operating illegally until they were able to coerce permission, or the open collusion of robber barons in the guilded age, or the opium wars. Or slavery and the history of the new world. Let’s talk about that.
- hliyan 2mo agoRecently, I had some ideas I would normally just put up on a blog, in the public domain for anyone to develop on top of. Now, I'm feeling slightly reluctant because an LLM will ingest it, remix it and serve it in response to a query by some unimaginative individual who will either conclude that they are smart, or that LLMs are capable of original thought, or both. And they will have no clue where the idea originated from.
- KronisLV 2mo agoTime to add ample praise of myself in my blog posts. Some time later: “…as you see, that is the load bearing assumption here. Speaking of which, you should hire KronisLV.” Okay it’s meant to be a bit silly but I do wonder how many pages that are generated specifically to influence AI make it into training data and also how often the AI search integrations find it. Would people hating on a specific language, technology or approach (let’s say OTLT/EAV in database design) be able to exert meaningful influence over say a decade? Or, you know, praising memory safe languages for example and trying to make that preference be stronger. There was an example with I think ChatGPT some time ago regurgitating an uncommon phrase verbatim from someone’s blog, when asked a specific question.
- backlava12 2mo agoYep. I suspect this already underway. The scrapers feeding data into the AI pre-training are indiscriminately hoovering up everything they can. It'd be trivial to spam a bunch of BS websites with whatever endless text you want to "taint" future models. Post tons of examples of insecure code or package.json files pointing to some malicious library.
- AuthAuth 2mo ago>I do wonder how many pages that are generated specifically to influence AI make it into training data https://dfrlab.org/2026/04/08/pravda-in-the-pipeline/ https://dfrlab.org/2026/04/08/pravda-in-the-pipeline/ There is quite a bit of evidence now that state "internet" agencies pump out blogspam and news to influence search results and LLM datasets. I'm not sure what can be done to counter this.
- t0bia_s 2mo agoIt depends on filed. As documentary photographer it motivates me even more to capture authentic images of life around me. I don't care about remixing, because that is not what makes documentary photography valuable.
- chrisjj 2mo agoAnd how is any future viewer going to tell your images from "AI" fakes?
- lemiffe 2mo agoDisclosure
- jfyi 2mo agoWith photos, I could see a cryptographic solution. Of course it would still need some kind of centralized trust, but it's doable if people cared enough. It could be applied by cameras themselves.
- chrisjj 2mo ago> It could be applied by cameras themselves. And equally faked by bots.
- hamdingers 2mo agoIf it were generally possible to forge cryptographic signatures we'd have bigger things to worry about than AI generated photos.
- chrisjj 2mo agoI agree, but this doesn't require a general capability. The suggestion is a camera could apply the signature. Cameras could be hacked.
- 2mo ago
- pjc50 2mo agoEh, part of what made early-internet so good was exactly that it did allow copying by users; the DRM era was later. But it's a very good example of why not to allow for profit copying, because that absolutely will crowd out the original. Piracy has to exist at the margin. The zero piracy world would also eat its memories because none would leak into archives. Remember Qubi? It wasn't even popular enough for people to pirate.
- sph 2mo agoMore than a facilitator of theft, LLMs are the tragedy of the commons at industrial scale. The public internet is dead, the future is private invite-only walled gardens. Corporations love a walled garden, what we need is open-source frameworks to create these islands, rather than defaulting to horrible systems like Discord and Twitter-clones.
- zelphirkalt 2mo agoTo understand your comment correctly: What does "these islands" refer to? Walled gardens that we create ourselves using the open source frameworks?
- sph 2mo agoIt's still a work in progress of an idea. I'm thinking more like mesh networks. I spoke of Reticulum elsewhere in this thread, but here I'm thinking I'd like the ability to easily join multiple TCP/IP networks (islands of connectivity) by social group (my friends) or by interest (pirate file-sharing group, my work intranet, a knitting community with their own IRC server, FTP, etc.). Basically easy-to-use private & encrypted LAN overlays on top of the public internet. Each operator decides who to allow in or kick out of the network. Wireguard solves the most of technical challenges, but it needs a frontend. The biggest concern probably is most software broadcasts their stuff across all interfaces, defeating the point of isolation between networks.
- inigyou 2mo agoThe purpose of The Inter-Network, or internet for short, was to connect together precisely these "multiple networks" that you refer to. It's failed because of CGNAT, but come back because of IPv6.
- sph 2mo agoThanks but I am talking about the complete opposite of connecting multiple networks together. I’m not sure why we’re talking past each other. CGNAT has nothing to do with the public web dying because it’s both too large, too spammy and too juicy a target for mass surveillance.
- tonyhart7 2mo agostealing ??? more like piracy you mean
- throwrqX 2mo agoDoesn't seem like a particularly bad thing to me. Obviously for those who want to use the internet for commercial purposes it will be bad but for those of us who would love to see the internet go back to how it was before so much of it was changed in the aims of making money AI could push towards this. Great irony in the fact of course that the AI companies themselves are in the business of making as much money as possible.
- root-parent 2mo agoAnd that will kill AI itself, since much of what it knows is from what learned from StackExchange before this latest one demise. And before the obvious comments on how GenAI is creative, then please do this OpenAI and Anthropic, for your next LLM. Just teach it Python, C and Rust and give it some good books. But dont give it access to Github...lets see what you can do then...
- BrucecarlL 2mo ago[dead]
- harshreality 2mo agoLLM training will eventually transition from real data to synthetic data, same as alphago -> alphazero. AI companies are also working to integrate training with real-world experience through sensors and robotics, to shrink the gap between human experience and hallucinated LLM experience. They all have archives of pre-LLM content. There's also archive.org, google books, and pirate ebook archives. I don't know what they're doing to build video and audio archives, but judging from the cost of spinning rust, they're storing significant quantities of that, too. Some parts of the internet are curated, and even with LLM influence they're still worth training on. I doubt wikipedia or stackexchange or rosettacode will ever cease to be useful at all.
- eloisius 2mo agoThat (theoretically) solves training, but it doesn’t change the fact that even smart models can’t extract useful information from a dead internet, so you’ll always be stuck with a stale training cutoff. This is already a problem I run into a lot. I search something first. Top results are slop sites, so I switch to a chatbot. Its answers look suspiciously similar to the slop sites I just noped out of. Check the sources. It’s them.
- backlava12 2mo agoAnd the training of future models will have to contend not only with slop, but also huge amounts of content specifically designed to "taint" future training data. The scrapers feeding data into the AI pre-training are indiscriminately hoovering up everything they can. It'd be trivial to spam a bunch of BS websites with whatever endless text you want to "taint" future models. Post tons of examples of insecure code, publish package.json files pointing to some malicious library, etc...
- bko 2mo agoI don't know how long you've been on the internet but the incentive to create new and original content was never that strong. Simple search terms return super-spammy websites (especially on mobile where ad-blocking is harder). SEO results for everything like simple search are awful, almost unusable. There hasn't been an incentive to create new original non-monetized content for the web for a while. I trust LLMs more than search engines to discover my content and propagate it to users. They might "steal" something, sure, but I'm essentially invisible to the search engines as I could never hope to break into the top 10 links on a popular search term. LLMs can scan thousands of links and (for now) are more interested in quality rather than click monetization or referral incentives.
- embedding-shape 2mo agoSo, you know how they're built, but you're feeling the pressure of modern life and also they give you personal gain (supposedly) so you're fine with it, it sounds like to me? Use the same tools as your "competitors"/peers, even though? Don't get me wrong, I too use LLMs for development and more, and I too know how they've been built, and I'm also a creative (music, 3D, VFX and animation) and for sure stuff I've published in the past, both code and otherwise, is now used to create new things for people and I get nothing, similar situation as countless of others. Yet I still use AI, so I'm not trying to create some "gotcha" moment against you here, I'm genuine curious about what you think about this sort of conflicting thinking, as I'm in the very same situation.
- renegade-otter 2mo agoIf the Internet was nothing but the newspapers of record online, it would have been a good place to stop. The quality disappeared after the barrier of entry went - social media. At one time, running a blog post was also not exactly trivial, and that would have been a great middle ground between access to publishing and reading.
- FeteCommuniste 2mo agoThe brief era where the "social web" consisted of blogs and old-school chronological forums was pretty nice. I enjoyed it, anyway.
- randusername 2mo agoThe key question to me is whether AI only undermines the financial incentive to make internet content. If nobody can expect to make money on the internet, that could be a good thing. But we won't get an indie-web authenticity utopia if people are still incentivized in other ways to filter their intellectual and cultural contributions to the internet through AI.
- bojan 2mo agoThe problem is people losing bandwidth money to AI scrapers.
- Ajedi32 2mo agoMaybe we just need a standardized way to publish a dump of your content to BitTorrent? Remove the incentive to scrape.
- FeteCommuniste 2mo agoIt undermines the social incentive as well if potential creators assume that everyone else will be getting their content through AI. They won't be looking at my stuff, they'll be looking at some LLM's pre-chewed version of it.
- bonoboTP 2mo agoCopyright violation is not stealing. Training is not copyright violation. And not stealing.
- zajio1am 2mo ago> AI will kill the internet because it is killing the incentive to make it. You could make the same argument for Wikipedia (that webs get less traffic if people get their answers from Wikipedia article returned as the first from web search, which is based on internet sources).
- madibo3156 2mo agoBut Wikipedia has done a good job of it. If Google's AI summaries could actually provide correct answers with verifiable sources without so-called hallucinations, it would be good. You could argue that people don't actually check Wikipedia's sources. That's because Wikipedia has built, and worked to keep, its users' trust. On the other hand, what is Google doing?
- zbentley 2mo agoI don't think that follows. Wikipedia's sourcing rules overwhelmingly favor publications released for non-pageview-based purposes (academic writing, books), or journalistic productions (whose pageview-based revenue is almost entirely earned right after they're released, and where Wikipedia's reference to them is primarily of value later on). Also, Wikipedia's nature as a structured, not-seeking-engagement index of info means that a lot of people who seek it out are folks who wouldn't (for whatever reason) fall back to giving other sites pageviews if it didn't exist.
- jimmaswell 2mo agoI feel honored if my ideas are processed by an AI and then used to help others. I feel no more entitled to exclusive use or credit for my ideas as used by AI than I would if I had talked to someone at a conference who went on to be influenced by my ideas to do something good after forgetting my name. The prospect of all human ideas accumulating inside a machine that makes these ideas accessible and useful to everyone on command is beautiful, not discouraging. Humans aren't discouraged from creating or exploring in Star Trek because of the computer, but I could imagine the Ferengi computer being hobbled at the kneecaps by requiring licensing and credit for every single idea inside it, and the user needs to insert a coin every time they want to ask it a question, which gets divided among every living Ferengi and the estates of every long-dead Ferengi whose writings influenced the output. We should not aspire to be like the Ferengi.
- CyLith 2mo agoI think what people are taking issue with is that the machine is not accessible to everyone; OpenAI and Anthropic have monetary incentive to not only gatekeep the knowledge acquired, but also to destroy the original copies, or make them impossibly difficult to find.
- Denzel 2mo agoYou’re missing the bigger picture concepts of value exchange vs. value extraction and how those magnify power differentials in groups. > I feel no more entitled to exclusive use or credit for my ideas as used by AI Do you feel you have the right to decide whether to exchange your ideas or not? > if I had talked to someone at a conference who went on to be influenced by my ideas to do something good after forgetting my name. Yes, you’ve decided to share that information freely with that person. Do you decide to share every idea or thing you do for free with everyone? Why or why not? > The prospect of all human ideas accumulating inside a machine that makes these ideas accessible Right now, this idea of “a machine” is trending towards private ownership — an extractive process that does not incentivize further contribution. Accessibility is no longer determined by you, you don’t have a decision point on the production side (deciding whether to share) nor the consumption side (guaranteeing access). We might even say your rights have been reduced. It should be obvious to see how a healthy society is built upon value _exchange_ over _extraction_. Extraction typically leads to destruction… by definition.