18 ms·
Microsoft exec called AI scraping 'the largest theft of labor in human history'
- cmiles8 14d agoThe evidence here is quite damning for OpenAI and Microsoft is clearly trying to distance themselves from OpenAI’s behavior here.
- Weryj 14d agoI think it’s more like ‘The absolute maximum possible degree of theft’ there can’t be larger, it’s everything current and past.
- redsocksfan45 14d ago[dead]
- deleted 14d ago[deleted]
- jacquesm 14d agoIt's the robbery of all of our culture to sell it back to us at a mark-up. Crimes this large are crimes against humanity. So many people whose life's work got appropriated without consideration, compensation or consent it is baffling. It is said that at the heart of every great fortune there is a great crime, so it should be no surprise that the most valuable companies on the planet will most likely result from this crime. And given that justice can be bought by those with the most money you can forget about anything coming of this.
- cyber_kinetist 14d agoAt least the Chinese AI companies are doing good service open-sourcing their models back to the public.
- mitxela 14d agoThere are no open-source LLMs. There are only downloadable LLMs, with no source for the model being provided. The source for the training and inference programs is not the source for the model itself, which would be the training set, training program, and random seeds.
- okeuro49 14d agoOLMo
- pluc 14d agoHas it affected DD as deeply as it as affected software engineering? Guessing clients feel a lot more "empowered" or "independent" and knowledgeable these days? I liked doing DD, just as much as I enjoyed developing, but it must be dying a slow death too. What's changed in how DD reports are produced?
- dgellow 14d agoWhat is DD supposed to mean? Nit: Please don’t use obscure acronyms when writing things to an international audience without defining them first… DD can mean so many different things
- pingou 14d ago>It's the robbery of all of our culture to sell it back to us at a mark-up Would regulation help with that? Right now you can download free models that have been trained on that "stolen" data. With regulation and compensation, only rich companies would be able to do that, and they would definitely not give it back for free. I put "stolen" in quotation marks because it's still unclear if we can call that stealing. Nobody would say a human reading a book and learning from it is stealing. I'm not saying that a machine doing the same is equivalent, but the only thing I am sure of is that I am not sure we can call it "stealing".
- embedding-shape 14d ago> With regulation and compensation, only rich companies would be able to do that Well, with some imagination, you can have regulation that forces companies to open up, not just close down. Imagine a law that stipulates that if you want to offer "LLM-inference-as-a-service", you need to also publish exact details about how it was trained, what datasets were used and also offer those exact weights for download. Sure, this would never happen, but just offering another perspective on how laws and regulation can be used if it was wanted, locking stuff down and pulling up the ladder behind you isn't the only way to use laws, although that is a very popular reason and approach.
- ben_w 14d agoI am unclear how this would help anyone? Any argument that writers and artists lose from these existing, would remain unchanged.
- embedding-shape 14d agoThe argument was "It's the robbery of all of our culture to sell it back to us at a mark-up". Remove the "selling" part, and force them to give the weights away for free, and at least it's no longer robbery that few rich people benefit from. Kind of like how public and free torrent piracy is easier to morally and ethically defend than piracy where they sell access to pirated content. I think we're past the point were we can feasible pay for "IP-protected bytes" digitally, better to just move past the concept. It's been slowly disappearing for a long time now already, most of us make most of our money on live events and other AFK activities rather than actually selling our art, maybe time for the rest to get onboard with this too.
- CrimsonRain 14d agoEvery time you're writing software or building machines/factories (which is automating things), you are committing a crime. Every time you learn from your superiors or colleagues, get better than them, get promotion or they get fired, you are committing a crime. Provide justice there first.
- oblio 14d agoScale matters. The average human doesn't do much on his own, and definitely doesn't uproot society or risk siphoning/leeching wealth from every person on this planet.
- CrimsonRain 14d agoYes, the whole IT sector is built on it. Robbing jobs, money, power, opportunities from billions of people and delegating them in to shitty jobs. Average human on his own...why draw the line there? It doesn't matter much what one human does...but what many/collective/society do and society has been "ripping off", "uprooting", "leeching (read: creating)" wealth since dawn of time. It is called PROGRESS.
- oblio 14d ago> It is called PROGRESS. There is no such thing as PROGRESS for progress' sake. And FYI, agriculture is a wonderful invention. Yet for about 5000 years after its introduction the average human had worse nutrition than the average hunter gatherer, which led to such things as height decreases for those 5000 years. Industrial agriculture is another wonderful invention. Yet 150 years later we're not sure it's sustainable and it's likely many of its aspects aren't, which will raise some sticky issues soon ("which billion people do we decide to let starve since we can't make enough food for everyone after most of our soil eroded?"). Repeat this for industrial textile production, mining, etc. I won't even go into climate change. And again, scale matters. Most individuals can only control what they do, and what they do generally doesn't impact much. But companies can impact a whole lot. > "ripping off", "uprooting", "leeching (read: creating)" wealth Let's not be 100% cynical here. A lot of what humanity has achieved has been genuine wealth creation and distribution/re-distribution. I would say more wealth has been created than leeched off. * * * And before you think I'm some starry eyed teen, I'll play the game. At the end of the day, me and mine have to outrun you in the face of PROGRESS. May the odds be ever in your favor.
- Razengan 14d ago> sell it back to us at a mark-up. What if it was for free, like Wikipedia? > Crimes this large are crimes against humanity. jfc no, sit down. Try doing something about the actual evil shit like arms manufacturers and the politicians ordering the deaths and misery of millions from the comfort of their couch. At this point in our civilization, all human knowledge NEEDS to be collated in one place and easily queryable. Otherwise it's just too damn difficult to make any further progress at the edge of our understanding; there's just too much shit for one person to learn "manually" (wait I'm not advocating for low-effort slop, chill) It's helping common folk who wanted to do something but didn't know where to start, while legacy search engines increasingly lead to spam, shallow knowledge or outright predatory shit (ofc AI could go this way too) Example: Not long ago I had the misfortune of becoming interested in some WarHammer 40K lore. Most of the links led to Fandom (the enshittification of Wikia) and that place is a cesspool of obnoxious ads. That content was written by unpaid volunteers. Should Fandom keep profiting from their work for perpetuity? Should I not be able to get the gist of what the heck a Qoiazrjirnowerx@# is without wasting my mortal lifespan on a horrible website? Or, if I need to ask something peculiar, should I post on Reddit or StackOverflow or HN and wait for someone to see it and deem to give a sufficient answer, only to have a pricky mod decide that the question doesn't "fit" the community? God hell no, if you don't know how much bullshit AI could eliminate for the silent majority then you were probably part of that bullshit. (that's a general "you" for whomever was fine with the status quo and not a personal insult @ anybody) If you see something you dislike increasing in popularity but can't figure out why, it's probably because a lot of people were sick of the way things used to work but their complaints were ignored by the people who now find themselves disrupted.
- steveBK123 14d agoDefending the same entities getting large DoD contracts to use AI for killing?
- Razengan 14d agoThey've been using computers for killing for decades, who's taking up pitchforks against computers? Who's taking up pitchforks against THEM? How do we keep getting deflected into hating the TECH instead of the people who abuse it??
- TacticalCoder 14d ago[flagged]
- api 14d agoIMO if they didn’t have proper licensing to train on the data the model should not be copyrightable. In the long term though I think models have no moat, so the cost will fall to the cost of compute and storage. Which is why they’re pushing AI safety panics: regulatory capture to outlaw open models and outlaw competition. And yeah, EA is neither effective nor altruistic. It’s a cult, part of the “Rationalist” and adjacent cluster of tech cults. They’re to tech what Scientology is to Hollywood I guess.
- sneak 14d agoCopying data isn't a crime.
- nullbio 14d agoAgreed, but now they're trying to stop other people from copying data so that they can be the sole gatekeepers of humanities collective knowledge.
- CrimsonRain 14d agoAnd they should get effed. Distilling should be a right.
- lostmsu 14d agoGood luck with that.
- djierardi 14d ago"copyright". its right there in the name of the rights.
- noosphr 14d ago>It's the robbery of all of our culture to sell it back to us at a mark-up. Crimes this large are crimes against humanity. Yeah the introduction of copyright was truly criminal. > So many people whose life's work got appropriated without consideration, compensation or consent it is baffling. Oh wait ...
- erulastiel 14d ago“It is said that at the heart of every great fortune there is a great crime” lol at this edgy 5th grade statement. So ridiculous.
- dataviz1000 14d agoOwning ideas with copyright and patents is what separates the United States from communism. The first time a saw a documentary about Tetris it really hit me what communism is -- nobody owned anything they invented or created. [0] It was a long time ago and I remember feeling sad watching the story. In the Soviet Union, a group of ~15 people, Politburo, controlled everything including any thought written to paper. It is this one line, Article 1 Section 8 Clause 8, that separates the United States from the disaster that was the Soviet Union: > To promote the Progress of Science and useful Arts, by securing for limited Times to Authors and Inventors the exclusive Right to their respective Writings and Discoveries; I don't think it is far fetched to call ignoring and disregarding the Copyright Clause a communist revolution, violent or not. That is the one thing the communists -- there have been many over the years inside the United States -- would change to make the United States a communist country. [0] https://en.wikipedia.org/wiki/Tetris#Spread_beyond_the_Soviet_Union_(1985%E2%80%931988) https://en.wikipedia.org/wiki/Tetris#Spread_beyond_the_Sovie...
- vitorfblima 14d agoIn a communist society there is no profit (or incentive for), thus no need for copyright laws. It's not abolishing copyrights that would turn the US into a commie country, communism is about abolishing private ownership to the means of production.
- deleted 14d ago[deleted]
- card_zero 14d agoWell it's worth reading the linked wiki article section, which includes a link to another article "Copyright law of the Soviet Union", flatly contradicting you unless you maintain that the USSR wasn't truly communist.
- echoangle 14d ago> unless you maintain that the USSR wasn't truly communist. Didn’t they themselves say they were a socialist society on the path to communism? Does anyone think the USSR was communist?
- logicchains 14d ago>It's the robbery of all of our culture to sell it back to us at a mark-up. Crimes this large are crimes against humanity. So many people whose life's work got appropriated without consideration, compensation or consent it is baffling. This is such a brain-dead take. By that logic there could never be any kind of AI, because unlike a human it'd be completely forbidden from learning from the sources of knowledge from which humans learn. It's stupid to suggest silicon brains should not legally be able to read copyrighted material just because you hate bigcos and capitalism. Learning isn't stealing, regardless of whether it's done by a human or a machine. By your logic someone reading and memorizing all somebody's life work is appropriation; completely inane.
- kg 14d agoYou're ignoring that there are options other than theft. Heard of buying things?
- px43 14d agoDo you think OpenAI has a New York Times subscription? Do you think that's actually relevant to what the New York Times is arguing here? The vast a majority of text that is being claimed to have been "stolen" was never for sale. Reddit posts, deviant art images, personal websites, etc.
- chrisjj 14d ago> This is such a brain-dead take. By that logic there could never be any kind of AI, because unlike a human it'd be completely forbidden from learning from the sources of knowledge from which humans learn. Incorrect. You overlooked consideration.
- layer8 14d agoLearning isn’t stealing, but if you’d set up a forum or hotline where paid staff would answer questions and write essay directly based on NYT content without a license for that content, that would probably be deemed illegal.
- 14d ago
- user43928 14d agoIt is the largest democratization of knowledge that ever happened. The 'sell it back to us' argument falls short in my view. Free versions are abundant, and in some time useful models will ship preinstalled on all mobile phones. The comment here seems incredibly pessimistic and quite dramatical.
- oblio 14d ago> Free versions are abundant, Awesome, can I make my own competitive LLM, just like I can make my own open source software? > and in some time useful models will ship preinstalled on all mobile phones. Considering hardware prices, that "some time" is doing super heavy lifting. It could be 10-15+ years before that happens and the local LLM is actually useful. Most people don't see hardware prices declining from current prices until at least 2030, likely much longer.
- user43928 14d agoNo. I also cannot make a competitive computer, yet can use one for my work to remain competitive. Yes, it could take some time to arrive on phones. It is questionable if it will ever make sense compared to using a paid hosted provider. But what are 10 years in the grand scheme of things? Should we have scrapped it all, called it a crime against humanity, and never developed AI, because it will take a decade to disseminate the benefits to everyone?
- oblio 14d ago> Should we have scrapped it all, called it a crime against humanity, and never developed AI, because it will take a decade to disseminate the benefits to everyone? Nah, we should have: 1. invested less, in a more targeted way 2. ideally like ARPANET, with the benefits given to all humanity 3. with fair royalties paid to all (where relevant) 4. and by creating an ever growing shared curated and high quality data set that would allow anyone to create their own competitive LLM ARPANET & co were all taxpayer funded and the internet has created more wealth than most human inventions. The base should be part of the commons, everyone should knock themselves out by building on top. Exactly like the internet.
- cluckindan 14d agoWorks produced by AI are not copyrightable. Are they in the public domain? Can an entity sell unique works which are in the public domain?
- card_zero 14d agoYes? I own a lot of books that were public domain when published (as reprints). I could read them on Gutenberg, but I'm paying for the nice paper formatting.
- godwinson__4-8 14d agoOur "culture" has long been the province of corporations. In prior epochs it was still the product of patronage and power. At least with LLMs we can glimpse an escape route to that which generations of humans have strived for - a world in which the labor required of each human to lead a flourishing life approaches zero. Instead of fixating on a remedy that seeks to criminalize AI, maybe focus on the relatively rather achievable goal of redistributing LLM gains. Would that not be the most desirable justice? What is your alternative, and would you foreclose the future in the name of a past that never really existed in the first place?
- jaybeavers 14d agoFlourishing life? Where is your data leading to the conclusion that AI models are leading to the world populace leading a ‘flourishing life’? Are taxes on revenue of companies like OpenAI somehow being collected and turned into a UBI and I just didn’t hear about it?
- godwinson__4-8 14d agoI said a glimpse. I didn't say it's here now or it would be easy. It is however, achievable. Certainly more so than engaging in the fantasy that we can criminalize LLMs out of existence. And it is likely more desirable than such an effort anyway.
- bayindirh 14d agoLike trickle down economics? We all experience the substance which is trickling down.
- jaybeavers 14d agoOr, perhaps, the companies forming the for profit LLMs could be made to pay a license fee for the copyrighted data they lifted from behind the paywall and then used to form a for profit entity with. Nothing about banning technology. How about enforcing DMCA and then applying a penalty for the knowing theft rather than negotiating a license?
- philipallstar 14d agoIf piracy isn't stealing training definitely isn't stealing.
- inquirerGeneral 14d ago[dead]
- c1sc0 14d agoAt the very least we should foribly confiscate the models & make them available free-for-all as open-weights downloads. Failing that, bring back the guillotine.
- kurtis_reed 14d agoRobbery implies taking it away from you so you can't have it anymore. How is that the case?
- sicher 14d agoIntellectual property works differently.
- 2OEH8eoCRo0 14d ago*Rent it back to us
- themgt 14d agoSo many people whose life's work got appropriated without consideration, compensation or consent it is baffling. Let's say hypothetically a solution was legislated globally, wherein each living individual whose work was scraped for LLM training is compensated with royalties relative to the work's value. Would that resolve the injury caused by the intellectual osmosis? Of course many of the original thinkers are now dead, and this system would mostly benefit those writing before the LLM age rather than help people going forward. The more fundamental objection seems to just to the concept of a machine that "learns" by ingesting public information, which is maybe ultimately a feeling that reality itself constitutes a crime against humanity.
- patrickmay 14d ago> Would that resolve the injury caused by the intellectual osmosis? Monetary compensation doesn't address the lack of consent. This type of usage was not anticipated when people made their intellectual product available for other humans to use. Scale does matter. What would resolve the injury would be to ask people if they are willing to have their content used in this way and to not train on material without consent. This includes open source software with particular licenses requiring attribution. Obviously there is too much money involved for this approach to work, but it strikes me as the most moral.
- brookst 14d agoMusicians and writers build new works on top off millennia of literary and musical history, then copyright their works and sell it back to us. Is this so bad? If Taylor Swift, consciously or subconsciously, gets an harmonic idea from a 1970's song and a fragment of a melody from some 1990's song... is that theft? I'm ambivalent about AI and, like all gold rushes, many of the players are terrible, dishonest, egomaniacal jerks. But I am deeply skeptical of the idea that aggregation of knowledge and culture is itself wrong. That's literally how culture has worked since the dawn of time, and our modern era obsession with credit and perpetual copyright is unhealthy. \ It's only in the past 100 years or so that this idea of "if you create it, it's yours alone and nobody can build on it without paying you" became current, and it was largely driven by the megacorps that AI haters used to hate (remember the despite for RIAA? I do). It's bizarre to think that someone's life work is entirely their property, as if they grew up in a box and did not build on hundreds of generations of other peoples' lives work. I don't object to disliking these companies; I object to the idea that you, me, anyone remixing culture is committing a crime. What the hell happened to the hacker ethos?
- bayindirh 14d agoThe problem is not Taylor Swift is being "inspired" from other artists. We have tons of examples this throughout history. The problem is, replacing Taylor Swift with its AI counterpart, and to use Taylor's own material to do that without getting her permission or compensating her. This is not about Taylor even. It's about everyone, you and me, and Taylor and Haggard and Blind Guardian and Sia, etc... We hated RIAA because they prevented us from listening to the music while trying to get it was hard and expensive. In short, we were not angry because they wanted compensation, but because they have cut the supply without giving us a solution. Now we have iTunes Store and Bandcamp for DRM free music, and nobody is against musicians getting their fair share. As a side note, I used to make music, I know what it entails. Hacker ethos has ethics. It has do experiment but don't cause harm embedded all over it. It's about experiment and discovery. Not about ripping people off for their own profit (unless you're a black hat of course), and getting things were free was part of sending a message, not monetary gain.
- 14d ago
- JimDabell 14d ago> It's the robbery of all of our culture to sell it back to us at a mark-up Learning isn’t stealing. They didn’t take our culture away from us and nobody is “buying our culture back” from them. We never lost it; it never went anywhere.
- 2ahsg1 14d agoYes, you are getting a laundered version for a monthly payment to the rent seekers, who are 1000x worse than the RIAA. And the internet is destroyed by slop spam in the process, so our culture is diluted and crowded out.
- bkaae 14d agoThat's true. However it does sometimes feel like our culture is being buried in slop.
- franktankbank 14d agoImagine someone asks you how to do something at work, you tell them. Then they create a huge packet filled with bullshit about how it got done and now they are your boss.
- TeMPOraL 14d agoIt's a shitty move, but ultimately, in between the bullshit narrative, they also did the thing - not you - so the promotion rightfully belongs to them. Execution trumps ideas, impact trumps raw effort, and such. Isn't this the entrepreneurial narrative?
- franktankbank 14d agoSure thats obviously how it works today. Now what do I do next time someone comes and asks how to do something?
- TeMPOraL 14d agoOf course you tell them, because you're nice person, and not jealous of someone else succeeding in a thing you aren't even pursuing? You wouldn't want to act like "the dog in the manger" from childhood stories.
- HumblyTossed 14d ago> It's the robbery of all of our culture to sell it back to us at a mark-up. Considering the audience here (aspiring tech billionaires), it'll be interesting the responses to this.
- TeMPOraL 14d ago> It's the robbery of all of our culture to sell it back to us at a mark-up. Except, of course, no one has actually been robbed, the culture has not been stolen - it's still there - nor are the people involved selling it back in any form. This rhetoric sounds impressive, but really looks more like "piracy is theft" line from early 2000s, similarly flawed in basic premise. Whether the end result threatens the form in which culture is created, at least beyond just threatening the business models of the gatekeepers, is a separate discussion, but you can't draw the heart-string-pulling "life's work got appropriated" arguments there so easily. And let's not forget what we got back for this: reified intelligence on a chip almost too cheap to meter, available to everyone across the world - not just rich West, inference is so dirt cheap that whole world uses it. It exploded in popularity organically, because of how many real problems of real people, including individuals and non-profits, it addresses. There's plenty to hate about how AI is transforming the world, but one thing it's not, is "robbery of all of our culture to sell it back to us at a mark-up".
- wjnc 14d agoCreators are unwillingly and contra to economic systems that have evolved over a few millennia entered in the Borg or the Matrix. So their achievements are reused for private and public benefits without their permission. Piracy is theft. And this piracy is a bigger theft than piracy on an individual download basis. Was there any doubt? (Linguistically I agree that you can’t “rob culture”. You can rape or reap culture though, and that is the point at hand.)
- mrngld 14d agoHN crowd: Check out my Plex server and 132TB media collection! Also HN crowd: Reading publicly posted information on the open internet is morally outrageous theft of the highest order
- TeMPOraL 14d agoAlso, The crowd: "Knowledge should be free! Culture belongs to all of us! Information must be democratized!" AI companies *proceed to copy literally everything and put it through virtual blender, until an universal general-purpose problem-solver tool comes out, then give it out for free or serve for peanuts to literally everyone on the planet with Internet connection * Crowd: bbb..buuut not like that!
- intrasight 14d ago> sell it back to us at a mark-up sell it back to us as markdown
- boringg 14d ago“It is said that at the heart of every great fortune there is a great crime“. That quote is from a fiction author - you’re essentially quoting Spiderman “with great power comes great responsibility”.
- cryptonym 14d agoAre you seriously putting Balzac and Spiderman on the same level to try to ridicule thoughts? Ad hominem.
- boringg 14d agoThey are both quotes that people say things like "they say" to imply someone of import said it to give it credence and truth to it but were in fact from fictional sources. Exact same pattern - take it as an ad hominem all you want.
- Vaslo 14d agoOther than the rare books than have been ruined, all the same knowledge is still out there though. So it’s not robbed in the sense of a bank heist. Maybe in the sense of pirating a movie. The markup is the millions in training they committed and the connecting the knowledge. Seems like a reasonable trade off to me. You can choose not to use it though.
- teh_klev 14d ago> all the same knowledge is still out there though > The markup is the millions in training they committed and the connecting the knowledge Sure, but gated behind a hallucinating idiot.
- ycsucks2 14d ago[dead]
- farseer 14d agoNobody has made much money from AI yet unless you count the shovel sellers (Nvidia). Not sure closed models would make any money ever and eventually the benefits should flow to everyone.
- bsoqk 14d agoIf LLMs scraping the web is robbery then piracy is stealing.
- ls-a 14d ago[dead]
- dhx 14d ago"Sweat of the brow" doctrine has been rejected in most countries.[1] Even Europe's Database Directive, probably the closest thing to an implementation of this doctrine, largely doesn't do much in practice. An example of "sweat of the brow" doctrine would be the series of "Beaches of ..." books by Andrew D. Short of the University of Sydney where significant sweat has been expended to visit and document every beach of Australia, particularly from a swimming safety perspective. That's a lot of very remote beaches, and many with crocodiles. Across the Northern extent of mainland Australia from Broome to Cooktown (this being one of the books in the series), 3500 beaches were visited and documented along 12000km of coastline.[2] AI could train on these books and gain an understanding of whether some small and unknown beach that receives <100 visitors a year has fine sand composition, pebbles, etc. Without "sweat of the brow", this use of AI is completely fine to regurgitate the facts learned from the book (regardless of the accuracy of the book). If "sweat of the brow" did exist, there would be some very significant (probably insurmountable) challenges to overcome, including: 1. You're a different expert in beaches and also want to visit all 3500 beaches across Northern Australia to provide a more up-to-date database, just in case beaches have changed in the last 10 years (e.g. sand washed away). In your database/book series, can you write "Andrew D. Short observed ACME Beach in 2006 to have fine sand. We observe 10 years later in 2026 the beach is now entirely pebbles of 15-20mm diameter", or is this infringing? 2. You're a researcher studying drowning deaths at Australian beaches and wish to extend the data published by Andrew D. Short's series of books with additional fields--dates of drownings at a beach, weather conditions on the day of drownings, etc, and then make some novel observations from the expanded dataset. Is this infringing? 3. You visit ACME Beach and observe and document it--what type of surface, dimensions, presence of reefs/rips/etc. You then put this information on your blog or social media account and it becomes a social media phenomenon as people are attracted to what has been revealed to be the best "secret" beach in the world. A few days later your website or social media account is blocked/deleted without warning--apparently there has been a complaint that you might have copied some facts out of a book you've never heard of. "Sweat of the brow" doctrine would almost certainly result in a tragedy of the anticommons[3] situation which would be worse for humanity as a whole. [1] https://en.wikipedia.org/wiki/Sweat_of_the_brow https://en.wikipedia.org/wiki/Sweat_of_the_brow [2] https://sydneyuniversitypress.com/products/9781920898168 https://sydneyuniversitypress.com/products/9781920898168 [3] https://en.wikipedia.org/wiki/Tragedy_of_the_anticommons https://en.wikipedia.org/wiki/Tragedy_of_the_anticommons
- Betelbuddy 14d agoWe know what is going inside black holes in other galaxies, we know the details of Israel nuclear program..., the Windows source code got leaked, we got the NSA tools and full details and locations of the Echelon architecture... Phds in Maths warns us daily about the terrible secrets of the evil AI inside their labs. How, their numeric matrices and gradient descent Python scripts, are about to kill 10% of us all, I guess the sick or genetically less interesting ones...and use the rest, as some meat/metal hive drones part of some Borg collective... There are ONLY TWO Stories and their details, that we collectively will never see. 1) One could come from the these brave souls that warns about an impending death...but their courage falters on another subject.... From Jacob Coxon to Evan Hubinger or Julie Steele, Samuel Marks, Josh Angels, Mrinank Sharma, Dario Amodei, Demis Hassabis, Geoffrey Hinton, Yoshua Bengio, Stuart Russell....The story of the full datasets they used to train the models, the data they stole, how many PB was, the amounts of data, the nights setting up torrents from unsuspicions IPs, where is it currently stored and how many exabytes is now... the massive data cleansing and data quality program to conform all the different formats, the internal discussions on the ethics of the stolen files, how large was the team, the CSAM content they sucked with their automated scripts and who was handling it internally, the porn, the massive amount of porn that is after all 80% of the internet, the leaks their data sucked with their automated scripts... And the other... 2) The Epstein Files.
- ThrowawayTestr 14d agoI haven't paid a dime for any of the LLMs I've used (electricity and internet aside)
- deleted 14d ago[deleted]
- Cthulhu_ 14d agoSo why is there so much outrage now when Google and other search engines did it 30 years ago? They literally said their objective was to collect and index all the world's knowledge, and when they scraped the internet clean a thousand times over they invested in scanning and digitizing everything that wasn't on the internet yet. AI is better at repackaging it back to the end user but ultimately I'm arguing it's the same thing. (caveat: yes I know there were plenty of people that objected to Google et al indexing everything; famously, Gmail was scary to a lot of people because they read your email to give you ads)
- seanw444 14d agoIt's not the same at all though. Google indexes the original content, hosted at the original location. It just gives you a way to find it. LLMs ingest the original data, throw it away, and spit it back out in whatever form they like back to the user. And since it's gotten so popular, many peoples' interactions with the content are no longer in the original form at all; their site traffic / book sales diminished.
- sillyfluke 14d agoGoogle didn't immediately create $20 and $200 per month paid tiers. The free thing lasted a long time (and you didn't need to sign in), and the initial monetization with ads, when it eventually came, was not immediately shoved in the user's face. the enshittification and bean counting came way later compared to the monetization blitz that occured with the AI companies.
- afavour 14d agoBecause they aggregated links to that information rather than copy it outright. It’s like the difference between an encyclopedia and its table of contents.
- giaour 14d agoGoogle (and other search engines) gave you an easy and reliable way to opt-out of having your content indexed and linked to. People did in fact object pretty loudly every time Google attempted to surface the information on google.com rather than sending traffic to the source.
- cineticdaffodil 14d ago[dead]
- Aurornis 14d ago> It's the robbery of all of our culture to sell it back to us at a mark-up. At a mark-up would mean it’s more expensive. The outrage is that they’re taking knowledge that was expensive to access because you had to hire experts or otherwise pay a lot of money for it and making it accessible to anyone who signs up for the ChatGPT free tier. Calling it “robbery” is also specious as no knowledge was taken away from anyone. The content in the training sets was out there in the world one way or another. It still is! I’m really perplexed by this sudden swing toward the idea that knowledge is something that we should encourage or incentivize to keep locked away or that other people should be forced to pay for use of knowledge. Roll back the clock a few years and tech sites would be almost unanimous about knowledge being free and unrestricted for the benefit of humanity. I’m keeping knowledge separate from actual direct rote duplication of content. Now we have this amazing era where I can download models to my computer, run them locally, and have enormous amounts of derived knowledge at my fingertips for the cost of some compute cycles. Except now it’s a “crime against humanity”?
- d3rockk 14d ago[dead]
- streetfighter64 14d agoI don't know about "crimes against humanity", seems to diminish a bunch of actual crimes against humanity. And Microsoft is one to complain! Talk about a glass house. Just checking the annual report for 2025, the median employee compensation was 200k per year, but the total dividends divided by number of employees was 100k, meaning that each employee gets only about 66% of the value they've generated. Isn't that theft? In any case, Microsoft has stolen 25 billion from its employees in 2025, and OpenAI has got 13 billion in revenue from "stolen" content in the same period, so that'd make them about equally bad villains, except OpenAI has mostly stolen from other companies.
- gkoz 14d agoSince they didn't improve the culture, it makes no sense to buy anything back at a mark-up. We can keep using the one we still have.
- antics9 14d agoIt’s a crime against humanity if the only way to retrieve the digital version of the works is through an LLM. The web getting flooded with slop and drowning all original work is one way to go about that. Another is paywalling and gatekeeping en masse. A third is shutting down shadow libraries.
- underlipton 14d agoPeople are getting lost in the weeds and ignoring the simple, fundamental fact that these companies are making fortunes off of labor that they didn't attempt to compensate the laborers for. "Property" and "IP" discussions are distractions; no amount of it can rationally get us around the utter unfairness of what occurred and the way it will warp our economy at a basic level if not addressed. It's not even enough to make the weights and models free; access should be free, and everyone who hitched their horse to this wagon should be on the hook for keeping the systems running, on their dollar. They took ownership of a venture that is short one (1) "Humanity's entire cultural corpus", and the only question is if we're going to issue a margin call.
- larodi 14d ago> you can forget about anything coming of this. first of all - a lot of people are definitely not forgetting it. perhaps many more are waking up to the fact. when so many people wake up to the fact that a massive theft of intellectual property IS what enabled present day AI, they will inevitably refuse to a) publish that much openly; b) respect any kind of copyright claims imposed by those who perpetuated, facilitated, enabled the theft. so, really, a lot will be coming out of it, we like it or not.
- aaron695 14d ago[dead]
- gyosko 14d agoAnd here we are, just watching and doing nothing..
- deleted 14d ago[deleted]
- Neil44 14d agoI understand the sentiment and partly agree. But also, the original has not gone anywhere. You're free to accumulate knowledge in the old way just as before. So maybe it's not theft of knowledge that we should be angry about, it's something else harder to define.
- pluc 14d agoLots of the original content is no longer available. Bots kill sites, AI kills monetization - both results in the original material disappearing.
- bcjdjsndon 14d agoCopying means we can both share in the knowledge, surely everyone on HN wants that right? Share the open source code for the good of everyone? Hackers used to say "information yearns to be free" now they're saying "that's my information and I don't want you using it" Probably indicative of America's wider downfall that they've all become so self interested
- sneak 14d agoHackers are irrationally anti-corporation. This is where the nonsensical AGPL came from, too.
- proc0 14d agoIf corporations weren't already owning the consumer, with AI it does this by many orders of magnitude. If something isn't done to prevent AI from being used to farm the masses for data, we will be living in a sci-fi dystopia without a doubt.
- sajithdilshan 14d agoIf someone asked what is 'the largest theft of labor in human history' I would have thought slavery.
- mitxela 14d agoNever ended, just changed in form.
- not_a_bot_4sho 14d agoYou're right that it never ended. But it didn't change form much. Still around 50 million people enslaved nowadays.
- midtake 14d agoAre you comparing modern workplace aches and gripes to literal 1800s slavery?
- someguynamedq 13d agoNo need to balk. Different forms of coerced labor can all be bad even if some are worse than others.
- cindyllm 13d ago[dead]
- y-curious 14d agoYeah gulags and other forced work camps also come to mind. But I guess this is a larger scale in terms of man hours
- bcjdjsndon 14d agoBut it's copying...how is it theft? Your labour WASNT stolen was it?
- deleted 14d ago[deleted]
- American87 14d agoI remember techchrunch.com making the argument that IP Infringment != Theft in the music piracy era.. how quickly the tide turns :)
- mitxela 14d agothey did say theft of labor, not theft of the things being trained on
- leonidasrup 14d agoIn case of programming. How much do the current LLMs invent solutions for user tasks, how much they just copy and adopt existing open-source solutions from from Github and other code repositories? This not a problem for open-source code under permissive software license, but works derived from open-source code with copyleft software license should be also under copyleft license. Could the biggest commercial benefit of LLMs be just working around limitations of copyleft licenses? What is the monetary value of human work put into copyleft software and later used to train LLMs? It's hard to estimate, but the study "Estimating the Total Development Cost of a Linux Distribution", estimated that it would cost $1.4 billion to develop the Linux kernel alone. https://consortiuminfo.org/metalibrary/estimating-the-total-development-cost-of-a-linux-distribution/ https://consortiuminfo.org/metalibrary/estimating-the-total-...
- menaerus 14d agoThey do invent code solution for the problem that exists in your codebase. Latter implies that the code solution LLM synthesizes is usually unique of a kind, so, it's not a copy-paste neither it is a simple extract from "another codebase" and adopted. IMO they operate pretty similarly to humans - we synthesize our solutions, and therefore build-up our knowledge, by collecting knowledge from multiple other sources, including technical books and blogs, open-source code repositories, and our past experiences.
- leonidasrup 14d agoLLMs can output near-exact segments of copyrighted code used for training. https://arxiv.org/html/2408.02487v3 https://arxiv.org/html/2408.02487v3 I wonder how would Microsoft react if someone would synthesize a code solution based on Windows source code. https://en.wikipedia.org/wiki/Shared_Source_Initiative https://en.wikipedia.org/wiki/Shared_Source_Initiative
- menaerus 14d agoI guess you're not writing code much or haven't done much so as your professional career?
- TutleCpt 14d agoThe most shocking point is that they have a Microsoft exec who knows what he's talking about.
- vintagedave 14d agoIn my experience many execs know what they're talking about. Where I feel you may see real variance is ethical and capability standards: willingness to stick to a line, and competence in analysis and execution based on what is known. Sometimes, hidden agendas can be misread as lack of competence, ie ethical lapses cause actions that are misread as capability lapses. Knowledge alone is less often a factor. Of course this varies widely across companies. I've been fortunate to work with some excellent folk at executive and C-level. Here, an exec clearly (a) understands or can make a clear, direct assessment and (b) was willing to do so in writing. Kudos on both grounds.
- ekunazanu 14d agoI think a different variant/opposite of Hanlon's razor applies when it comes to corporate or political decisions: Don't attribute to stupidity when it can be adequately explained by malice or greed. This sounds rather obvious, but I feel people forget it far too often.
- TeMPOraL 14d agoI call this the Hanlon's Handgun: "Never attribute to stupidity that which can be adequately explained by systemic incentives promoting malice." Previously: https://hn.algolia.com/?dateRange=all&page=0&prefix=false&query=hanlon%27s%20handgun&sort=byDate&type=comment https://hn.algolia.com/?dateRange=all&page=0&prefix=false&qu...
- soraminazuki 14d agoYou can see the same dynamic playing out in this very thread too.
- Sohcahtoa82 14d agoOh, I think MS execs (and all other execs, top-level politicians, and pundits for that matter) know what they're talking about, but they'll tell whatever lies they need to enrich and empower themselves.
- wj 14d agoHow is this different from Microsoft scraping to build Bing? Honest question. There is a line in the sand somewhere apparently.
- oreoftw 14d agoHuge difference between building AI and a search index.
- sneak 14d agoWhy? In both cases the SaaS downloaded the whole web and derives 100% of revenue from content they didn't make.
- haxiomic 14d agoLinking to work, where ownership and attribution is clear and the owner has the ability to commercialise is a very different thing to “laundering” content through the model, quoting the midjourney developers here > "We just need to launder it through a fine-tuned codex." [0] [0] https://cybernews.com/news/midjourney-ai-images-art-lawsuit-copyright/ https://cybernews.com/news/midjourney-ai-images-art-lawsuit-...
- sethops1 14d ago
- bcjdjsndon 14d agoCopying isn't stealing you babies
- drstewart 14d agoI remember when this was the prevailing thinking online until about 2024. But that's when everyone was trying to justify their own piracy of GTA or whatever. Suddenly, they don't like other people pirating.
- winstonwinston 14d agoImo copying (like carbon copy) makes it fair compared to taking it (without attribution) and claiming it yours. For example, for software I usually use MIT license, which is a permissive license with attribution.
- noosphr 14d agoI'd call the introduction of copyright the largest theft of human labor in history. No one was compensated for all the free labor they did before the introduction of copyright which copyright holders then privatized. For example the Disney corporation would have had to pay the Brother's Grimm estate for the use of Snow white under the copyright regime they instilled in 1998 with the Mickey Mouse Protection Act. That we are finally having a sane pendulum swing towards no copyright is a breath of fresh air. The only way the AI bubble could improve the world more is if we end up becoming a Type I Kardashev civilization to feed the data centers. Then when the bubble pops we suck up all the extra CO2 with all the now idle nuclear power plants we can't shut down. At the same time it's truly baffling going on a site called _hacker_ news and seeing corpo talking points from the 90s/00s regurgitated wholesale. Information wants to be free.
- catdog 14d ago> That we are finally having a sane pendulum swing towards no copyright is a breath of fresh air. Might be a short one though if all goes to plan. Just another form of gatekeeping the worlds information and with new gatekeepers replacing the old ones. > At the same time it's truly baffling going on a site called _hacker_ news and seeing corpo talking points from the 90s/00s regurgitated wholesale. Information wants to be free. Look at who owns that site, no surprise here.
- aeon_ai 14d agoVectoralist News doesn’t have the same ring to it
- 47282847 14d ago“Information wants to be free“. It’s not “theft of labor”; the work was already done. If anything it is theft of “intellectual property” (aka “copyright infringement”), if you believe that is a thing, but not of the “labor” that went into it. My personal take: anyone producing content, everyone’s creativity, is fed by something that others did before. We’re all standing on the shoulders of giants composed of previous generations and their “content’s” distribution and dissemination. I have an immense gratitude for all the labor before me that I was and am allowed to partake; without that, I would be nothing. Sharing information is an act of love; gatekeeping it is short-sighted greed. New technologies have always “killed” previous “labor”, out of which new opportunity grows. I just wished the collected data was public. I hope we all get a mega-leak at some point.
- bcjdjsndon 14d agoUnpopular opinion on here
- ambicapter 14d agoProbably because it claims there’s no problem, and then makes a tiny little mention of the BIG problem at the very end.
- newsclues 14d agoThe unpopular part is the hypocrisy where their work has already or must be rewarded while other people’s work is not.
- proc0 14d ago"I just wished the collected data was public. " That's the entire contention here. It's a double standard. Companies will sue the living hell out of anyone taking their IP, whether it's code or art, yet they have no qualms taking all the data they need from anyone and everyone. It was already a problem before, i.e. artists getting paid very little for work that companies profit a lot from like musicians or digital artists, but now with AI it's on steroids.
- nullbio 14d agoIt's humanities collective knowledge and work. That's why nobody should ever buy the narrative of distillation being a crime or theft. It should be a human right to distill these models. Distillation should be being provided as a service.
- rich_sasha 14d agoYeah, distilled, hosted by OpenAI and charged for. And don’t you try reverse engineer what they did! If this was all open, I’d maybe half agree.
- riskable 14d agoThat's what this boils down to: What are our rights? Everyone has a right to scrape the Internet. That includes corporations who scrape the Internet to train AI models. If we take away that right, how would the Internet even work? It wouldn't. Example: I could tell curl right now to download this techcrunch article and all the comments about it on HN and I'd be violating no law. I'd be infringing on no one's rights. If I then distributed these downloaded files without permission then I'd be violating copyright law. The thing it certainly would not be is theft! People claim AI companies are "stealing" human labor but that's not true. They're saving (in their databases) the fruits of human labor and other bots/software. Then they're using that data to train AI models. The only conclusion I can make whenever someone says "AI is theft!" is that they have no idea what they're talking about. My assumption is that what they really mean is, "AI is bad for labor!" and possibly, "cheap AI is incompatible with capitalism." Which very well could be true. But if AI really undermines the value of labor that much, the problem isn't the AI, it's capitalism.
- anon7000 14d ago> People claim AI companies are "stealing" human labor but that's not true. They're saving (in their databases) the fruits of human labor and other bots/software. Then they're using that data to train AI models. And profiting on it on a scale that’s hard to fathom. Someone who spent effort creating a great resource or doing some research and maybe got some income via donations, ads, whatever. Now that information from their resource is distilled into a big model. The original author is screwed, the model provider makes money through the effort of everyone else. It worked well for everyone before, because there was recognition, prestige, a sense of doing good for people, even a chance for some income. That’s completely eliminated with AI.
- rich_sasha 14d agoIt’s not that different to the US helping itself to indigenous peoples’ lands in North America, decimating them with smallpox and alcohol, then generously offering reservations. At least it’s consistent, is what I’m saying.
- ks2048 14d agoGenocide vs non-destructive copying of digital information - I'd say that's pretty different.
- m4rtink 14d agoAsk Anthropic how non-destructive their copying is, when the clandestinely buy rare books for cheap & shred them after scanning and not sharing the result.
- KingMob 14d agoFirst, most of these books are old and plentiful, not rare at all. Think "Master Windows Exxcel '95!", not first editions of "Master and Margarita". Libraries routinely destroy plentiful old books that nobody is reading any more. Anything actually rare and valuable is too pricey to hand over. Second, they HAVE to destroy the books because US copyright law REQUIRES it. They're only allowed to digitize the books if it's considered "transformation" and not "copying". That's only allowed if they destroy the book afterwards.
- ks2048 13d agoThat's true, but the article was about "scraping".
- sedan_baklazhan 14d agoAI overall is the ultimate piracy crime. I wonder what a token cost would be if AI companies were to pay royalties to every author who made their business even possible.
- logicchains 14d agoThen let's make humans pay royalties to every author from whom they ever learned something, even if it was offered freely to them, only fair?
- sedan_baklazhan 14d agoIndeed. If I take somebody's work, transform it somewhat and sell it, I should pay royalties (unless the author explicitly allowed me to do so). That is exactly the case.
- KingMob 14d agoSadly, they're just following in the footsteps of every large media/publishing/music conglomerate that already screwed over the vast majority of artists/musicians/writers. For every Taylor Swift striking it rich, there's 999,999 who can't even pay their bills with what their copyright gets them. I don't always agree with Doctorow, but he's written a lot of good stuff on how stronger copyright won't help broke artists. Even just today, it turns out: https://pluralistic.net/2026/08/18/enron-corpus/#sign-here https://pluralistic.net/2026/08/18/enron-corpus/#sign-here
- sedan_baklazhan 14d agoIt is not just about artists. It is about every kind of intellectual work: scientific research, essay, fiction, painting, software, etc - the list goes on..
- alansaber 14d agoSpiderman pointing
- jappgar 14d agoAll the "LOL you wouldn't steal a car???" posts in this thread miss the point entirely. AI is cannibalizing information. It is literally destroying information and impoverishing those who would produce more of it. At a long time scale, AI dominance is apocalyptic even if it never intentionally hurts anyone.
- sebastiangrill 14d agoI think so too. The only way to redeem this theft would be to force all AI companies to open source their models if they cannot prove that copyrighted material was not used to train them.
- simonw_simonw_ 14d ago[dead]
- totetsu 14d agoAre those factory workers we saw photos of now, wearing cameras to capture the movement of their hands stitching getting compensated for a generations worth of wages? Do they even have any choice but to give away the copy-right to their labor?
- pbasista 14d agoI do not understand what "theft" they are talking about. Those AI bots were scraping publicly accessible internet. Publicly. Accessible. Of course there are some parts of the publicly accessible internet which host content that may be considered illegal or has been obtained illegally. If those AI bots used such content as well, it is fair to call it out as wrong, in my opinion. But that is a separate topic. Blindly calling scraping of publicly accessible internet a "theft" is, in my opinion, disingenuous. Especially when coming from a company operating a web search engine. Which itself has its own bots scraping the same parts of the internet 24/7.
- cluckindan 14d agoA lot of sites have terms and conditions which explicitly disallow the use of site content as a part of another service. If I have a bike and you start renting it out without my permission, surely you are committing theft of some sort. If I build a complex custom bike and you start copying individual features from it on your custom bikes, surely you are committing theft of some sort, but whether it’s punishable depends on whether I’ve decided to go full corporate and protect my designs with patents and trademarks. You’ll be hard pressed to patent or trademark anything if I have published and documented prior art.
- pbasista 14d ago> A lot of sites have terms and conditions which explicitly disallow the use of site content as a part of another service. Having terms and conditions in itself is irrelevant. Because in order for them to have any legal meaning, it is necessary for the other party to agree to them. An agreement can be implicitly enforced by law. Or explicitly enforced by the website itself before giving access to the data. If neither of those are present, there is no enforced agreement. And agreeing to it becomes optional. Such sites should be considered, in my opinion, publicly accessible. > If I have a bike and you start renting it out without my permission, surely you are committing theft of some sort. Of course. But that is a bad analogy. No one is "renting" or "taking" anything from those websites. The bots are just reading it. Therefore, a better analogy would be that you have a bike, parked out in the public, and people are looking at it. By looking at it they steal nothing from you. And the bike and all of its parts remain yours at all times. That is a suitable analogy, in my opinion, to what those bots are doing.
- tom2026hn 14d agoThe problem isn’t just “stealing the fruits of human labor”, it’s also driving down the value of human skills and even taking away human jobs.
- ambicapter 14d agoOne leads to the other so its simpler to point the root issue.
- Trasmatta 14d agoAnd the cruel irony is that they stole our work to train the AI that devalues our work going forward, and will cause many of us to lose our jobs. I regret every line of open source code I ever wrote.
- tom2026hn 14d agoThis process can even feel a bit like parasitism, it empties out the host, like in <Alien>.
- bluefirebrand 14d ago> I regret every line of open source code I ever wrote And every stack overflow post, every reddit post, everything. I regret participating in the open Internet. Here I am anyways, I guess. It's just in my genes or something.
- Trasmatta 14d agoSame. I've scrubbed my presence as much as possible from the majority of the internet (except HN for whatever reason) to prevent future models from being trained on my output, but the damage is already done. Part of me exists in pretty much all AI models now, without my consent. And those models are actively stealing my career and passions.
- keeda 14d agoI think you just described "technology" -- would you outlaw all technology then? ;-)
- fwlr 14d agoThe largest theft of labor in human history … and it’s to do away with the laborers by making a device that produces labor substitute, with full awareness that the substitute produced is not fit for the purpose of making more such devices. It’s like burning all the crops for heat, which you use to boil the oceans for salt, which you use to salt the earth so no more crops can grow. If AI wants to destroy humanity it better get its boots on, or else AI companies might get there first.
- GardenLetter27 14d agoI think it's okay to advance humanity, but they can GTFO when they then try to ban distilling and open models.
- tremon 14d agoIs AI advancing humanity, though? Or is it only advancing technology while divorcing it from the human?
- riskable 14d agoDefine, "advancing humanity." Are we talking about turning everyone into philosophers and somehow ascending to a higher plane of existence? Yeah, AI isn't going to help with that (probably). Or are we talking about useful, positive benefits to every day people like better speech recognition, tools for the visually impaired, disease research, physics research, science in general, and loads of other areas where AI is improving things?
- aitoolcrux 14d ago[flagged]
- kunley 14d agoBut what about M$ owning Github and doing the same with its content? Github even did not deny scanning private repositories. (Gitlab denied the same when asked). So...
- jaybeavers 14d agoYou have to admit there is now some lovely schadenfreude to be had from the whole ‘Chinese free LLM companies be stealing our theft! Stop them!’ whining.
- Joel_Mckay 14d agoThere is nothing funny about $9Tn in FOSS getting misappropriated, and sold as isomorphic plagiarism tokens... or the estimated $4.6Tn in loan debts these 7 companies incinerated when the bubble pops. Anyone that lived through the dot-com or housing market bubble know what a collapsing Ponzi scheme does to real businesses, and peoples retirement funds. Popcorn ready =3
- juvvel 14d agoI wouldn't have a problem with working off the fruits of other people's labor because most of us are essentially doing that everyday anyway, the issue is that big tech companies (want to) reap all the benefit and create profit from something that should be accessible to everyone. Everything is getting privatized -- housing, water, electricity, and now, thinking and knowledge. We are heading towards a world where you have to pay even more excessive fees just for existing and for completing any basic task.
- Draiken 14d agoIt's the age old privatize the profits and socialize the losses. People lose their jobs, the environment is destroyed, our bills skyrocket and all of the gains go to the people who own all the shit... I honestly cannot believe some people still believe that we'll ever get to a society where nobody has to work and we can live our lives happily ever after. Maybe too many Disney stories?
- lofaszvanitt 14d agoPeople are like sheep. Plus those who work are preoccupied and are too tired to react to these changes.
- sleight42 14d agoFor the life of me, I can't understand why your comment is being downvoted or flagged or whatever makes it go gray on HN. I'm guessing it's people knee-jerking that you're being political? I can't understand the people who don't see it. The data centers strain the power grids then electricity costs go up for everyone else. This is de facto a regressive tax because everyone needs electricity and the poor pay proportionately more of their income for the increased cost. Live in San Francisco? Probably not now unless you're rich because of the skyrocketing cost of living due to Tech and AI money. Another de facto regressive tax, driving away other people. Environmental damage? The poor are the most impacted and the least able to absorb the costs. Do they have the property or renter's insurance to protect them from these disasters? Another de facto regressive tax. Need a new phone? Or a computer? Same problem. Want to dabble in AI? You're not going to get too far on $20/month. It's mostly a wealthy person's game. Or there are the statistics that the vast majority of successful founders from up upper middle class families or wealthier. Wealth centralization is what our economic system does. The purpose of a system is what it does. If it wasn't, the system would have been changed.
- meerita 14d agoThere will be a point where companies will not need to scrape any content. Agents will create endless streams of probes, and they will end up solving all kinds of knowledge problems.
- chrisjj 14d agoAbsolutely. The parrots will get so clever that they'll extract knowledge from pure vaccuum.
- JohnFen 14d agoAs someone who thinks that genAI is harmful, I deeply resent that any of my work has been used to help train it. I will never forgive these companies for forcing me to contribute.
- dev1ycan 14d agoBecause it was, it completely defaced all copyright and similar laws, like there is ZERO ground to stand against China now regarding theft... it's so weird how this is being allowed.
- iamflimflam1 14d agoI don’t mind these companies scraping my content. But for love of god, my blog changes at most every couple months. You don’t need to scrape it every few minutes.
- UltraSane 14d ago"The Net interprets censorship as damage and routes around it."
- rafaelmn 14d agoCopyright is artificial scarcity rationalized by arguing that producing novel intellectual work is valuable, but requires substantial effort that can't be recouped, so we have to incentivize it somehow. LLMs and AI are changing that proposition substantially - human effort involved in producing copyrightable content is getting reduced constantly to the point that if we abolish copyright entirely we'll still have more content than we could ever hope for. AI/robotics eliminating scarcity of physical goods sounds very far fetched but in the intellectual space it looks very very plausible in the near future - so it could be time to abolish IP laws soon, especially if AI manages to advance enough in R&D and research space.
- DarkNova6 14d agoIf there is no legal guarantee that human creativity can pay off we are starving art and humanity from its inception. If ordinary people cannot participate in the act of creation, you get exactly what hollywood has become. Sorry, but this reads like a mouthpiece exactly from those companies that benefit the most from having no copyright and I doubt your have thought this actually through.
- rafaelmn 14d agoLack of way to capture value from intellectual labor is considered a market failure that leads to suboptimal market results for consumers, but with AI that argument becomes very weak. Sorry but the point isn't to create artificial scarcity just so intellectual labor is well off, that's a negative side for the consumer that was considered necessary tradeoff. Market economy should be about providing the most value to the consumer. Disclaimer - I was never a fan of IP laws despite them working in my favor, with AI I can see them finally being abolished.
- fwlr 14d agoWe could imagine, as an extreme case, a technologically highly advanced society, containing many complex structures, some of them far more intricate and intelligent than anything that exists on the planet today – a society which nevertheless lacks any type of being that is conscious or whose welfare has moral significance. In a sense, this would be an uninhabited society. It would be a society of economic miracles and technological awesomeness, with nobody there to benefit. A Disneyland with no children.
- rdsubhas 14d agoThey are not selling the information. They are selling a service for easy access to that information. These are two different things. Note: am not an AI fanatic.
- AdamN 14d agoCorrect solution here is to make sure royalties are embedded in the AI responses (and work output). These should be appropriately priced and go back to the owners of the IP. If the IP is no longer owned then it can be free use. There should be a carveout for non-profit or government AI.
- pier25 14d agoAI content cannot be copyrighted.
- AdamN 14d agoFirst of all it can. Second I'm talking about attribution to copywritten material (and royalties being paid accordingly). This would only be for for-profit AI implementations though.
- pier25 14d agoNot in the US. https://www.reuters.com/world/us/us-appeals-court-rejects-copyrights-ai-generated-art-lacking-human-creator-2025-03-18/ https://www.reuters.com/world/us/us-appeals-court-rejects-co... As for your second point I doubt it will ever be economically feasible. AI companies cannot generate profits even while stealing their training data.
- juiceland 14d agoCopyright infringement, if this even were that, is not theft. Chattel slavery is the largest theft of labor in human history.
- Joel_Mckay 14d agoTrademarks are intellectual property, and every LLM knows what Mickey Mouse looks like. =3
- jMyles 14d ago[dead]
- alex1138 14d agoInformation wants to be free and all that but there's a sense in which AI really is real intellectual property theft in an ethical sense compared to others and of _course_ it was Facebook who steals from everyone where Zuckerberg personally approved it Their bots are also apparently the worst. Google does not put huge strain on your public-facing website (I think). Facebook does, they're incredibly malicious about it
- Ylpertnodi 14d ago> Information wants to be free Does it?
- alex1138 14d ago"And all that" is a linguistic trick to imply "yes, this is being discussed -"
- haritha-j 14d agoI just don't understand people saying "but a human learning from a book isn't illegal". How do people not understand that some laws only make sense at a certain scale? One human learning from resources and being added to the labour pool is not the same as an infinitely copyable entity doing the same thing. One has negligible impact on the demand for the original, and the other replaces 99% of the demand." And creating a rule that says you cannot train on any material unless the rights holder authorises it via license is not complicated. That will creat a amrketplace where creators can decide the price for their content. It's just inconvenient.
- DownGoat 14d agoI agree that there is not an orange to orange comparison, but I have still not seen a law that could scale as well. The example that comes to mind with a proposed law like this is how would you license work that build on another work? What if I decide to publish a blog post after taking some course, that distills whatever I learned in the course for free?
- californical 14d agoWe already have laws that cover this. If you know enough to write your own course that completes and steals significant share from the original, you likely have so much background knowledge that you didn’t need to take the course in the first place. If you only ever learned about the topic from this course, you likely have an uninteresting shallow understanding that won’t take share from the original. And if you substantively copy the course and publish your own version which is heavily taken from the original, then you may be violating their intellectual property. Seems like it’s still fine to keep that as-is. We can still charge a license fees to use somebody’s works to integrate into their algorithm, since algorithms aren’t humans. I do think copyrights should be shortened to 20 years but that’s another discussion
- coffeefirst 14d agoBecause it’s a bad faith argument that presupposes integrating someone else’s work into your algorithm is equivalent to me reading a book. Your rule would be the right way to do all this. You could even have a mechanical royalty that applies by default where you can train on anything that hasn’t set rules and a preset rate.
- b3lvedere 14d ago"The question of whether AI firms can legally use copyrighted material to train AI has no clear answer, but judges have been largely favorable to AI companies’ arguments that training constitutes “fair use.” This legal rule lets people use copyrighted work without permission in certain cases, like parody, news reporting, or criticism. Earlier this month, the Trump administration contributed a brief in defense of OpenAI’s unlicensed use of copyrighted material to train its LLMs. " So training can make it legal as well. Interesting...
- vegnus 14d agoTheyre just jealous that theyre being surpassed on their market capture of computing
- ohrus 14d agoYet we still tell students to buy textbooks. The individual must always pay. The corporation can do whatever the hell it wants. The hypocrisy of this new world is already catching up to us.
- Waterluvian 14d agoI have this weird vision of an alternate reality where governments (say, National Archives) are the ones creating the models as a public service and then the rest of the industry is just commoditized pricing of hosting them, competing with value add bits. And we’re on here reading articles about how the latest release of the EU model does a better job generating maps now and the new Canadian model seems to apologize less and whatnot.
- Kuyawa 14d agoGoogle has been scraping everything from us since day one. Meta, Microsoft, Github, Slack, Reddit, StackOverflow, big and small, every single app that interacts with people uses our own data to make money and create walled gardens. I haven't seen a single one opening their silos to the world. That's our data, we produced it, you captured it and now you think it's yours So no, your cries for regulating others because you are losing the race won't work this time.
- WarmWash 14d agoHow much money have people paid to use these services over the years? None? Ok, now you understand the business model.
- Kuyawa 14d agoI understand the business model, we are data providers, they are aggregators. That's fine. What's not right is that they want to limit the use of such data when it's not theirs in first place. They just store it but that doesn't give them a license to prohibit the use by a third party since we all are owners of that data
- WarmWash 14d agoNo, they are data sellers. You pay them in data. Just like money, it becomes theirs.
- cute_boi 14d agoIt works if you can bribe politicians.
- hereme888 14d agoMicrosoft is one to talk.... Remember when MS trained copilot on all your github code?
- seydor 14d agoImagine the parthenon marbles. When they were looted it was even a celebrated act, but they are still stolen in the british museum centuries later.
- deleted 14d ago[deleted]
- OrvalWintermute 14d agoThroughout history we’ve been able to retell stories, to copy content, to create shallow clones or synthesis It is only now in human history that we are able to create nearly perfect copies, and we’ve been taxed incredibly for this with overpriced everything.
- rietta 14d agoI remember when open source software, and Linux in particular, was the threat to the world according to Microsoft execs.
- joduplessis 14d agoMicrosoft executives levelling "tone-deaf" up in realtime.
- dzink 14d agoAre we considering what is the shelf life of information? If you build a building, the expense on materials determines longevity. If you build a city. The robustness of government and the economy in it determines the property taxes and value of property over time. If you make or cook food. The majority of the nutritional value of it goes to the initial consumption. Once the food has stayed out without refrigeration it is taken over by bacteria and fungi. Refrigeration seems to be paywalls. Once the information is out it accumulates at exponential rates - the amount of text on the internet does not diminish but increases. Some people may “prune” old content away, but that is rare. Human attention is somewhat a fixed number. Thus text left out is not consumed, but sits idle and decays in accuracy and value over time. The fresh content of valuable should be in a fridge. If not valuable it is released - thus scavengers and those hungry and motivated to dig can consume it. If spammy and sales-y / propaganda-y which a lot of content farms are doing, the goal is for it to be consumed by the masses and push the zeitgeist to buy its premise. That’s Sugar or addictive shelf-stable junk foods. AI model companies are the bacteria / fungus/cockroaches/rats of the information dumpster. They sneak out any remaining energy from content that would otherwise be buried by other content and try to give it a second shelf life - one reachable and accessible and consumable by humans. They make alcohol. Alcohol is addictive. Ir may mess with your brain - it may make you lazy. It will sneak in bad decisions because it lowers your judgement. It is repurposed food, not the one you are used to injesting. It may even have its own agenda - depending on how the information is reprocessed. And it also has a shelf life since humanity continues to have new insights and people keep getting new alcohol brands to try.
- shevy-java 14d agoYet he also helps destroy all those jobs. The thief is calling "Catch the thief!". He does not see the moral dilemma here?
- levischoen 14d agoTransatlantic slave trade calling - we’ll hold. I know there’s a memory shortage but history books are cheap.
- MaxHoppersGhost 14d agoSlavery was widespread and universally practiced across the world throughout history. Everyone posting on this board is a descendent of someone who was a slave at some point. Transatlantic slave trade is a drop in the bucket.
- anon48293 14d agoFunny. The end of copyright and patents is by far the best thing about AI to me. All of it is nonsense. Great that you drew a picture of a mouse once, I really fail to see why I couldn’t draw it and sell it either. It was moronic from the get go.
- Ylpertnodi 14d agoDraw a mouse then. But, I do agree in principle that disney's mouse is theirs. Forever. Whenever I see the little rodent, I think 'disney', and if you drew a similar mouse and you aren't disney, then you've cheated me.
- danesparza 14d agoI mean ... I know a few African nations that might disagree.
- MaxHoppersGhost 14d agoThe black folks still in Africa were the ones selling the slaves so maybe not.
- danesparza 12d agoSo ... that makes it somehow "not theft"? I'm confused at the point you are trying to make.
- WarmWash 14d agoThe righteousness of the internet, the same internet that desperately called to end IP laws, championed piracy, ad-block everything, and always use proxy services to backdoor paywalls/login walls This same group of people, now being on the other side table, are screaming an crying that it's not fair. Grow up and reap what you sow.
- 1p09gj20g8h 14d agoIt literally is. Anyone saying otherwise is deluding themselves
- 1234letshaveatw 14d agoWhat if all this slurping and training empowers humanity to cure cancer? feed the starving? travel the stars? We are not only making the accumulated knowledge of the world accessible, we are making it actionable. Sure I'm ignoring all the possible bad outcomes lol, but if an independence day (movie) type scenario was playing out nobody would be batting an eye. I guess cancer is not as sexy though
- FLeXMurphy 14d agoApart from jacquesm's wonderful milquetoast comment, if anyone has any practical solutions to this problem that do not involve suspension of disbelief that voting (with or without wallet) and calling "your congresscritter" or whatever other nonsense people spout, now would be a great time to voice it.
- gaigalas 14d agoFraming it as property is the wrong angle though. It's an ecosystem, not a cache of good writings that was stolen. That ecosystem was hurt severely and its recovery is uncertain. I think we'll not have people writing good content for a long time (there's no reason or incentive to), and the effects of this will splash back heavily on AI companies themselves. You can see AI as a battery for intelligence that took a long time to charge and it's being used right now. For years, it was charged with all sorts of novel content that went undiscovered and AI is making available. That charge is the production of novel content, new insights, cross-pollination between areas, slowly driven by humans. My view also draws a conclusion about recursive self-improvement: it is impossible for a battery to re-charge itself. I don't particularly think it can be done with this technology (LLMs). I could be wrong though, but I don't think I am, and we'll know within our lifetimes. If things stall, it's likely because it has ran out of seeds/charge/substrate and not a technical limitation. It is in the long-term interest of AI companies to make incentives for people to generate novel public insights, they just don't know that yet.
- Fnoord 14d agoCopyright infringement, not theft.
- j3th9n 14d agoEveryone benefits.
- jolt42 14d agoYes. The resulting LLM is much easier to find information than ever before. To me, that's an improvement on the source data, much like the hitchhiker's guide to the universe. If it results in breakthroughs especially in health, steal away baby.
- kiicia 14d agoinformation was easy to find since wikipedia and google (when both were unpoisoned), LLM-s were already good as "books you can talk with", now when LLM-s start cosplay experts they are not (and try make decisions they are not fit to take responsibility for) it gets literally insane, like looking at car wreck in slow motion
- heaney-555 14d agoLLMs are not compression algorithms. From an information theory perspective, that's impossible given their size. Thus, a distinction needs to be made between viewing material to _learn_ and viewing material to _verbatim repeat_. It's not illegal to read the New York Times and then start giving paid advice based on what you learned, as long as you don't repeat the text verbatim.
- cush 14d agoIt’s irrelevant if it’s legal today or not. This is new technology and may be new precedent.
- bustadjustme 14d ago... but you have to pay to read the NYT. You paid for the information. Guess who didn't.
- Timon3 13d agoI can see how LLMs can't be lossless compression algorithms, but why not lossy?
- someguynamedq 13d agoCitation needed. Of course they are compression algorithms
- cineticdaffodil 14d ago[dead]
- nunez 14d agoYou or I scrape a website for casual use? Straight to jail, right away. Big tech scrapes ALL OF THE WEBSITES CONSTANTLY to resell to you as knowledge? Shut up and take all of my money. The 2020s is the wildest timeline indeed.
- jgalt212 14d ago> Several of the new admissions, however, run counter to OpenAI’s fair use defense, particularly the rule’s requirement that use doesn’t substitute or harm the market for the original work. Sounds about right for the people and orgs involved.
- pianoben 14d agoAs if the transatlantic slave trade wasn't a thing. What a perfect illustration of our industry's self-absorption and self-regard.
- markhahn 14d agoHas anyone found a meaningful discussion of how scraping is theft? Obviously, reproducing works in whole is infringement. That's not what AI is doing, so the question becomes: how is scraping different from ordinary reading? Is it just that site owners want to play back history and retroactively create high-cost licenses for scraping?
- fwlr 14d agoA key consideration in most legal definitions of theft is “intent to permanently deprive the owner”. While this has historically meant scraping is not theft (because copying doesn’t erase the original, nobody is deprived), in this specific case the AI companies’ business plan (copy a person’s content and train on it to make their model more capable of replacing that person) could very well meet the bar of intent to deprive.
- Madmallard 14d agoIt absolutely meets the bar of intent to deprive. Anything otherwise is willful ignorance or astroturfing.
- IX-103 14d agoI'm not sure how to be upset over this. For decades copyright has been extended and extended. Meanwhile, the ease and speed of spreading published works across the globe have increased massively. Don't forget that when copyright was first created, it could take multiple years for first editions to make it across the globe. In that environment, multiple decades of copyright makes sense. Whereas today that would actually stifle innovation and creativity rather than incentivizing it. I mean, because of the length of copyright we've gotten all of these live action remakes of Disney films or superhero movies rehashing the same stories. I find it a little hard to be upset about AI. Supposedly stealing copyrighted works when the vast majority of those works. Probably should have been in the public domain to begin with. I have a faint hope that this scuffle between the AI companies and the publishing industry will result in more reasonable copyright laws, but I think it's more likely that exceptions will be made and AI will be treated as a special case.
- underlipton 14d agoI'm upset because I have to pay to use it or else be hit with onerous rate limiting for stuff that "should have been in the public domain to begin with."
- Arcuru 14d agoIt's the scale that's the problem. I've been saying for a while, the value of any individual piece of work is not terribly valuable for an LLM, but the aggregate value of all human work is obviously very valuable. I think this line of thinking can be the argument for why we should heavily tax these AI companies above and beyond how we tax other industries. Also that if you're going to use these things to write software, you should make it as virally copyleft as possible https://jackson.dev/post/moral-ai-licensing/ https://jackson.dev/post/moral-ai-licensing/
- fhn 14d agoBut Windows telemetry, github code training, scanning all OneDrive documents/outlook emails is not theft
- bentt 14d agoThis is a great basis for a dividend from AI revenue to be paid back to society in a more inclusive form than stock. As more money flows into AI companies, data centers, and other related infrastructure, an amount should be extracted and redistributed in the name of balancing this equation.
- montjoy 14d agoWhat’s the Microsoft angle here? A few years ago they were ready to hire anyone from OpenAI that was willing to leave. Is it just catering to the anti-AI sentiment going around?
- 123176 14d agoAtlassian, powered by this spy tool, is truly insane: “We've always believed the best way to move work forward is to capture context once and let it flow everywhere. With Grok powering Loom's speech-to-text and Cursor turning that into code, we're closing the loop from context to code: record what you mean, and the work gets done. It's a glimpse of where AI-assisted development is headed.” All these failing companies are trying to bullshit their way out of the decline. Atlassian could have, you know, come up with a usable GitHub competitor. Instead they dream about coding by yapping.
- wrs 14d ago>[Nadella said] if he “had been made aware that OpenAI had scraped and trained on information that was behind a paywall,” he would have “invoked [Microsoft’s right to] require OpenAI to retrain its models.” Whew, good thing it’s too late to be accountable for that now, huh? Water under the bridge. Mistakes were made. Eggs, omelets.
- nathias 14d agodistillation is a human right
- mike_bob 14d agoGo cry me a river of lies Mr. Microsoft Exec. Corporate culture is a blight on humanity.
- thunkshift1 14d agoThis will lead to a massive settlement between the big boys and most people who put stuff out in good faith will be left out of it. And that will be the end of it. We will never hear anything about this ever again and the ‘theft’ will continue like normal.
- 1vuio0pswjnm7 14d agoSource: https://storage.courtlistener.com/recap/gov.uscourts.nysd.612697/gov.uscourts.nysd.612697.1587.1.pdf https://storage.courtlistener.com/recap/gov.uscourts.nysd.61... p. 1 "This case is about, as Microsoft's Director of Applied Science put it, an astonishing theft of unprecedented proportions; SF1437, perhaps the largest theft of labor in human history. SF1652" p.11 "As Microsoft recognized: millions of people around the world will soon consider large models hoovering up all their work to be an astonishing theft of unprecedented proportions and admitted that almost no one intended for content they created to be used in this fashion, nor are they compensated for its use. SF1437." p. 74 "As Microsoft's Dr. Glen Weyl put it, compensating creators is in the best interests of my employer, of my country, and of many other groups I belong to. SF1657." Hyperbolic quotes from Microsoft employees are, IMO, the least interesting elements of this brief Here is Microsoft's brief. Note how MSFT responds to the "web grounding" claims https://storage.courtlistener.com/recap/gov.uscourts.nysd.612697/gov.uscourts.nysd.612697.1585.0_2.pdf https://storage.courtlistener.com/recap/gov.uscourts.nysd.61... It seems OpenAI does not want the public to know about (a) OpenAI's data collection and retention practices and (b) the number ChatGPT users have requested deletion of conversations https://storage.courtlistener.com/recap/gov.uscourts.nysd.612697/gov.uscourts.nysd.612697.1608.0.pdf https://storage.courtlistener.com/recap/gov.uscourts.nysd.61... "OpenAI seeks to redact specific information about [(a)] the number of users who requested deletion of ChatGPT conversations and [(b)] OpenAI's related data collection and retention practices." "Disclosure would give OpenAI's competitors insight into OpenAI's confidential business practices and customers and cause competitive harm to OpenAI. Yeats-Rowe Decl. 4." Perhaps it would causes competitive harm because, upon learning about OpenAI's privacy practices, ChatGPT users might reduce their usage of ChatGPT Declaration is sealed so we can only guess
- 1vuio0pswjnm7 12d ago*cause
- sharts 14d agoSo Microsoft exec stating what most people have been saying already is newsworthy.
- keeda 14d agoWe may not like this, but let us contemplate what laws would be in place to prevent a thing like this; I suspect we would like those laws even less. The laws at play here are related to Intellectual Property, specifically Copyright. Yes, it is terribly flawed, but it is the product of centuries of case law dealing with very hairy issues, and I believe it is fundamentally sound, and here's why. As the name implies, it deals with only verbatim copies of works or subsantial portions thereof. It very expressly does not cover abstract things like concepts, ideas, themes, facts, or patterns, and rightfully so, because we really do not want anyone owning something that broad. But these abstract things are precisely what have been extracted, at unimaginable scale, to build these models! Each pattern in the tokens derived from these works contributed imperceptibly tiny perturbations to randomly initialized weights, interacting in incomprehensible ways into vectors representing concepts and ideas and facts, the cumulative aggregate of which has somehow created a form of intelligence. There is no copying, only gleaning, and so Copyright Law falls short. But what is the alternative, and do we want it? To prevent something like this would require some sort of legal protection on the more abstract things. We do have a legal framework for those: Patents! But as is very clear on HN and in many Tech circles, those are an extremely contentious topic (even though they actually protect much narrower ideas than most presume.) I don't think anybody anywhere really wants any protection on broader abstractions, and rightfully so. So: we as a society expressly decided these abstract things belong to the commons, and those are the exact things these labs harvested. This is probably the only logical culmination of our technological journey, and is within the very reasonable legal frameworks we have evolved over centuries. As such, it is not productive to dwell on fighting this or bemoaning this. Instead we should focus on ensuring that this technology -- with its immense potential and opportunities and dangers -- benefits everybody as much as possible. That is a better way to compensate everybody's labor, and that is a much richer and fruitful discussion to be had.
- deleted 14d ago[deleted]
- sinan-faizal 14d agowell it can be used for good as well, how we stopping it?
- __bjoernd 13d agoBut US labs have been doing it right, while Chinese labs are only distilling from their superior products, right?
- aitoolcrux 13d ago[flagged]
- nullpoint420 13d agoOne could call it the largest distillation attack in history
- neop1x 12d agoNo wonder LLMs are "better" every year. No juest scraping, they train on more and more data people send to them. So basically, it steals other people's data to sell back to you.