34 ms·
Meta torrented & seeded 81.7 TB dataset containing copyrighted data
- jokethrowaway 2y agoGreat, can we get the full Kim Dotcom treatment for Zuckenberg now? I'm also ok with abolishing copyright all together if he's too untouchable
- TZubiri 2y agoI love it. This plotline feels out of cryptonomicon or silicon valley series.
- tremarley 2y agoebooks are a 1-2 mb each max. 81.7 TB are a lot of books, like 42-85 million books.
- thunkingdeep 2y agoI’ve got 70-80mb pirated books, I think because of the illustrations. Guess it depends on the book.
- mateus1 2y agoI don’t think they’re using picture heavy book for LLM training, no?
- moralestapia 2y agoYes they do, there's multimodal models.
- mnsu 2y agoFor multi-modal models, why not? They would be probably some of the best data.
- michaelt 2y agoSometimes the PDF of a book is big because the book's packed with important illustrations and charts - like a textbook or journal paper. Other times a PDF of a book is big because someone scanned it and didn't have trustworthy OCR, so they figured distributing images of text at 1.5 MB per page was better than risking OCR errors.
- WithinReason 2y agoPresumably they didn't create the torrent
- littlestymaar 2y agoEven if they didn't use the illustration(which isn't clear given multimodal models), they'd still make use the text in the books.
- deleted 2y ago[deleted]
- rbanffy 2y agoI don't think they need to be selective. It's not like Meta can run out of storage.
- RIMR 2y agoJust because the LLMs are trained on text doesn't mean that images we're a part of what they downloaded. You clean up the data after you acquire it, not before.
- hulitu 2y agoWhy not ? Do you think that AI doesn't enjoy porn ? /s
- squigz 2y agoIt could be anywhere from a few million to a hundred million https://annas-archive.org/datasets https://annas-archive.org/datasets
- weberer 2y agoThe article says they got datasets from Anna's Archive. It was most likely the scihub/libgen torrent which is 96.0 TB right now and contains 92,872,581 files. That's about 1 megabyte per file. https://annas-archive.org/datasets https://annas-archive.org/datasets
- southernplaces7 2y agoWhere does one find these torrent datasets? Did they download the books in bits and pieces or as a single huge multi-TB file?
- panki27 2y agoThey could have at the very least seeded some more, to give something back to the, uh, community.
- Havoc 2y agoReally curious what the judges are going to do here. Horse has functionally bolted on this already I’m guessing slap on wrist despite courts going after individual for a couple of movies torrented pretty hard
- WhereIsTheTruth 2y ago[flagged]
- deleted 2y ago[deleted]
- aprilthird2021 2y agoIs there any other possible outcome than a fine? That too one which will not really affect Meta's overall earnings
- Havoc 2y agoIdeally we have a conversation about how we as society have ended up in a situations where we have a two tier justice system. At a minimum the starting point of discussion here should be that if life ruining $80,000 per item is an acceptable fine for individuals then why is it not the same for corporations. Which would probably get you a number in the trillions at which point we could have a discussion about reforming this entire system. But yes realistically slap on wrist is what is going to happen here.
- hnfong 2y ago> Is there any other possible outcome than a fine? Yes, of course. It's quite possible that judges realize that if they restrict training data to licensed materials, LLMs will become stupid and China will overtake the US to become the leader in AI, and because that can't happen, they'll make up some reason to make training on unlicensed data legal. It's definitely fair use! I'm not even joking. Last time the US Supreme Court basically said "Android is too important, we have to declare its use of Java API fair use."
- 2y ago
- mnsu 2y agoSo according to some AI, the damages awarded per infringed work is ~$750 minimum in the US. 80TB of books, each let's say 10MB on average, would be 8 million works. So Meta should pay 6 billion USD for their copyright infringement?
- oersted 2y agoNice calculation, that’s actually quite doable for them, they have already been paying similar fines for a while.
- gorbachev 2y agoMinimum doesn't cover willful copyright infractions, for which maximum penalty is $150K per work. That comes out to quite a different number.
- timeon 2y agoProsecutors filed for Swartz 50 years of imprisonment and $1 million in fines. Can you calculate how many years that would be for Mark and his people?
- qup 2y agoI ran it, it came out to zero
- bmsleight_ 2y agoSo if I torrented and seeded, I would be doing it for my own entertainment, not commercially. I expect big copy-write holders to come after myself. If Meta does it - I guess they have better lawyers ? Could make interesting case law.
- unification_fan 2y ago> Could make interesting case law. Yeah, to perpetuate this system where only those who can afford lawyers get to benefit
- echoangle 2y agoSince it’s case law, everyone would benefit from the precedent
- timeon 2y agoThere already is precedent with cases like Aaron Swartz.
- echoangle 2y agoNo there’s not, he killed himself before there was a decision. That doesn’t create precedent.
- hnfong 2y agoThe last time the US Supreme Court decided on copyright law, they basically said "we like what Google is doing with Android, so what they did is fine". So, no, not necessarily.
- deleted 2y ago[deleted]
- passwordoops 2y agoEye for an eye. Meta losses rights to 81.7 TB of IP. Transcribed into a text file
- cma 2y agoMeta already does that to themselves every year or so, deleting all internal communications. They've thrown away a huge amount of communication to source code commit reinforcement training data as a result. They do it to avoid emails making it into trials like this.
- yodsanklai 2y ago> Meta already does that to themselves every year or so, deleting all internal communications. Aren't they obligated by law to keep all internal communication?
- stingraycharles 2y agoYes, they are. But I can imagine the fine/impact for this being much, much lower than the consequences of all their nefarious communication being used in trials.
- cma 2y agoWhen there is a specific order after proceeding starts, but not before. There can sometimes be other orders as part of govt settlements like Google was recently accused of violating. You may be thinking of certain financial institutions where it is a hard requirement, and maybe there are some other regulated industries too that have it.
- immibis 2y agoCompanies often don't do what they're obligated to. As long as they can keep plausible deniability.
- zaik 2y agoNo large company will ever consider training a public LLM on all their internal communications.
- openplatypus 2y agoSomething tells me uncle Donald will exonerate his new favourite lapdog from any criminal or civil liability.
- Terr_ 2y agoIANAL but the pardon power (A) only extends to criminal punishments, not civil liabilities and (B) copyright lawsuits can be launched by anybody, not just the Department of Justice. So, barring further Might Makes Right shit--which I'm not willing to fully rule out--Trump can't fully shield Zuckerberg et al.
- impossiblefork 2y agoThey can also sue in France, or Spain, or Japan.
- openplatypus 2y agoHow unlikely it is for Trump to declare AI national security and simply make it lawless playground fro Zuck & Co.
- ksynwa 2y agoA good chance for federal prosectutors to "send a message" as they did with Aaron Swartz but I don't see things going that way.
- courseofaction 2y agoEven after JSTOR declined to press charges in that case. Despicable. The US has dug the hole it's going down.
- acomjean 2y agoIf you were wondering why meta was making a lot of donations to the new government (including settling a lawsuit for 25 million with the New president, 1 million to the inauguration)…. I suspect there will be no federeal charges. The rules have always seemed different for corporations regardless. https://www.businessinsider.com/trump-settles-lawsuit-meta-million-presidential-library-2025-1 https://www.businessinsider.com/trump-settles-lawsuit-meta-m...
- Nasrudith 2y agoWell of course, bullies always prefer targets that can't fight back. That itself is unfortunately a basis of the legal system from it being run on flawed monkey brains. Why else is hitting vulnerable children okay but getting into a consensual bar fight illegal?
- smgit 2y ago[flagged]
- deleted 2y ago[deleted]
- 9dev 2y agoNot that I have any particular sympathy for the guy, but could we keep the tone a bit more civil around here? HN is one of the few bastions of grounded discussion on the internet, and I’d prefer to keep it that way.
- zfg 2y ago> HN is one of the few bastions of grounded discussion on the internet Hacker News has its fair share of irrationality. I submitted an article about efforts to undermine Wikipedia: https://news.ycombinator.com/item?id=42962971 https://news.ycombinator.com/item?id=42962971 But it was flagged and locked to commenting. Wikipedia is one of the internet's greatest projects. Hacker News apparently doesn't have the stomach to discuss the threats to Wikipedia.
- RobotToaster 2y agoBefore I decided my opinion on this I need to know their ratio.
- adamsocrat 2y agoArticle states: Meta also allegedly modified settings "so that the smallest amount of seeding possible could occur"
- RobotToaster 2y agoIn that case, throw the book at them.
- MaKey 2y agoDamn leechers!
- moffkalast 2y agoThe jury of their peers finds them guilty!
- malfist 2y agoBig tech taking and not giving back, where have I heard this before?
- nyoomboom 2y agoRemembering Aaron Swartz in this moment
- stingraycharles 2y agoWhich was arguably more innocent — scientific papers.
- piyuv 2y agoMeta is not “innocent”, and comparing this instance with Swartz is a huge offense to his legacy.
- Philpax 2y agoI don't think you've read the parent comment correctly?
- piyuv 2y agoParent comment implies Swartz was guilty of some degree. I vehemently disagree with that.
- aruametello 2y ago> Parent comment implies Swartz was guilty of some degree as a constructive criticism, you might want to reconsider your interpretation of >"Remembering Aaron Swartz in this moment" -> Which was arguably more innocent — scientific papers. As in, both hold some degree of illegality (objectively), so when pointed that "he is guilty of some degree" is due to the jurisdiction laws (broken or not) regardless of societal/moral values that the context may apply. perhaps a better answer would be to point that he shouldn't be punished for those actions.
- HeatrayEnjoyer 2y agoYou're quoting a different user
- gameshot911 2y agoBeyond illegal downloading and distribution of copyrighted content, the article also describes how Meta staff seemingly lied about it in depositions (including, potentially, Mark Zuckerberg himself).
- malfist 2y agoHuh, a big tech CEO lied to us? Flippant response I know, but too many people worship at the alter of the job creater and believe these folks are moral upstanding citizens
- Ekaros 2y agoConsidering prices for single work, this must be multi-billion dollar compensation. Take for example 675k paid for 31 songs. So 20k a song. If we estimate book to be say 10MB that would 8 million works. So I think reasonable compensation is something along 163 billion. Not even 10 years of net income. Which I think is entirely fair punishment.
- pinoy420 2y agoFor creating a backup of library genesis. No. They should be awarded a philanthropic prize.
- striking 2y agoThere's evidence of them seeding back as little as possible. I'm not sure how that's "creating a backup".
- ralusek 2y agoThey're talking about creating and releasing Llama...not seeding the torrent
- pseudalopex 2y agoA model is not a backup.
- qup 2y agoThen why are we mad about the copyright stuff?
- striking 2y agoBoth things can be true: * an AI model is not a backup of the contents of all of the books in the sense that it would preserve their contents or similar such it might e.g. be useful for future generations * Meta has (allegedly) been unfairly benefiting / profiting off of the copyrighted work of others by illegally reproducing copies of their work. Not just in the AI model sense[1], but actually (allegedly) downloading them directly from pirate repositories in a way that isn't straightforwardly fair use and even uploading some amount of this pirate data in return. I feel like the parent commenter may have been making the typical argument for preservation of copyrighted materials, and I'm amenable to it... when it's regular people or non-profits doing that work, in a way that doesn't allow them to benefit unfairly or profit off of the hard work of others (or would be connected to such a process in some way). Plaintiffs allege that Meta didn't just do all this, but also talked about how wrong it was and how to mitigate the seeding so they might upload as little as possible. So no matter how you slice it they allegedly 1) knew they were doing something at least a little bit wrong and 2) took steps to prevent the process that might otherwise have preserved the copied materials for the public interest. And I feel like you probably knew all this, but maybe I'm missing something. 1: the typical argument wherein the model wouldn't exist without the ingested data, a lot of it is still in there, it is of course a derivative work and the question is really how derivative is it and what part of the work can they claim is their own contribution
- gizmo 2y agoBased on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the early days. The contracts with the music labels came later. GPL violations by commercial products fits the theme also. Companies aggressively protect their own intellectual property but have no qualms about violating the IP rights of others. Companies. Individuals have no such privilege. If you plug a laptop into a closet at MIT to download some scientific papers you forfeit your life.
- newsclues 2y agoComprehensive intellectual property needs to happen for the modern (digital) era. Basically the entire legal system needs to be retooled and rethought for computers.
- actionfromafar 2y agoLooks like the entire legal system is being retooled at the moment.
- threeseed 2y agoNo we just need to enforce the existing laws. And the legal system is for humans not computers.
- newsclues 2y agoThe existing laws are a problem, and are not enforced in a fair and just manner. Yes, the legal system is for humans, but we can use technology to improve the system for humans, so it's faster, better and more fair, because humans aren't perfect, and now we have technology to be better than the system create a long time ago. You don't think the legal system should run on pens and paper right? Adapting to typewriters, was a benifit to the system? Well, video on demand, live streaming, and things like LLMs can also make the system better for humans.
- fimdomeio 2y agoIt really makes you think about those crazy internet folks from back in the day who thought copyright law was too strict and that restricting humanity to knowledge in such a way was holding us all back for the benefit of a tiny few.
- stefan_ 2y agoThe more concerning thing is that the best thing these overpaid people could come up with was.. download the torrent, like everyone else. Here you are, billions of resources, and no one is willing to spend a part of it to at least digitize some new data? Like even Google did?
- dietr1ch 2y agoI think they are morally required to improve the current state. - Seed the torrent and publicly promote piracy pushing lawmakers. - Contribute with digitisation and open access like Google did in the past. - Make the part of their dataset that was pirated publicly accessible. - Fight stupid copyright laws. I can't believe that copyright lasts more than 20 years. No field moves that slowly, and there should be tighter limits on faster moving fields.
- malfist 2y agoCopyright and patent aren't the same thing. "Fast moving field" doesn't make sense in terms of copyrights. There's no reason the copywriter should last some minimum duration after the life of the creator. If I write a really popular book, I don't want Hollywood to make it into a movie without compensating me just because they waited a few years
- everforward 2y agoFast moving field does make sense in terms of copyright because the knowledge is recorded in documents which are then copyrighted. E.g. research papers. > If I write a really popular book, I don't want Hollywood to make it into a movie without compensating me just because they waited a few years I genuinely don't understand this. Even at a decade copyright, pretty much anybody who was going to buy the book and read it has already done so. It costs you virtually nothing in sales, and society benefits from the resulting movie. Your goal is to deprive everyone of having a movie, because someone who isn't you is going to make some money that was never going to you anyways? Your goals for copyright appear to be a net negative to the system that enforces copyright, which begs the question why should the system offer protection at all?
- postepowanieadm 2y agoThat's horrible! Magnet anyone?
- pinoy420 2y agoLibrary genesis
- ykonstant 2y agoWeird shenanigans are happening in libgen at the moment; better go through Anna's Archive to look for the items you want, it will link you to the corresponding mirrors more reliably. At least this has been the recent experience of a friend who used libgen and anna's archive to download legal, public domain works!
- bmacho 2y agoNo, AA is rate limited to being unusable, while libgen is fast enough.
- addandsubtract 2y agoAnna's Archive: https://annas-archive.org https://annas-archive.org
- immibis 2y agospecifically https://annas-archive.se/torrents https://annas-archive.se/torrents - this is a meta-project which aggregates illegal copyrighted material from other illegal projects. You absolutely should not download any material this page links to, although you can use it for the purpose of researching about shadow libraries.
- gorbachev 2y agoPrevious: https://news.ycombinator.com/item?id=42673628 https://news.ycombinator.com/item?id=42673628
- uncomplexity_ 2y agodid they not seed enough, is that the crime? lol
- palata 2y agoGood, we know it. Nothing will happen, because nothing happens to billionaires and their companies. Musk is proving it every day now.
- jokethrowaway 2y agoThis is why we need to abolish the government. If the government doesn't have any power, they can't do preferential treatment to their cronies. Enough with laws for thee but not for me!
- nprateem 2y agoIf you're an author with a book likely to have be hoovered up, I wonder what you'd get from the fb models if you asked "complete this in the style of [author] in [book]: [quite a long excerpt]" If you get a direct quote then you're good with your claim, surely.
- aprilthird2021 2y agoI believe that is part of this lawsuit pretty much
- Nemo_bis 2y agoThat's the NYT's case. Not necessarily very strong. https://www.techdirt.com/2024/03/05/openais-motion-to-dismiss-highlights-just-how-weak-nyts-copyright-case-truly-is/ https://www.techdirt.com/2024/03/05/openais-motion-to-dismis...
- unraveller 2y agoThe way it works counts if you bring prompting into it. It could easily have learned enough style chops of [author] from other sources to mimic/predict those stanzas from raw data points. Whatever the ruling one thing is for sure, plagiarism is no longer the sincerest form of flattery. The human authors are out for AI blood on this.
- yoavm 2y agoWe all like hating big corporations, especially Meta, and people seem to use this as an opportunity to advocate for punishing them. I think it's wiser to advocate for changing our IP laws.
- Ekaros 2y agoFirst punish them. Then change the laws.
- DaSHacka 2y agoI bet you and my "first build the product, then worry about security" manager would get along.
- Ekaros 2y agoMy approach is same. First fire that manager. Then define security.
- ngneer 2y agoThat one is tough, because they are blind to the risk. I try to only work with people who have been burned before or have been around long enough to have seen the aftermath. Let me guess, they are probably telling you "show me the vulnerability", but refuse to delay shipping or fund the PoC. Best advice is to communicate in writing the most likely risk and threat scenarios, with as much data or extrapolated data as possible. When the security flaws are later discovered, that is data you can refer to. From what I read, this is what Zoom was like early on. They had amateur hour security and then when s*t hit the fan they beefed it up and retained a security team. I guess you could say it worked for them?
- anticensor 2y agoIn many countries, that'd trigger an automatic release/repayment of unjustly sentenced fines.
- aprilthird2021 2y agoI think most of the public is probably in favor of stronger IP laws now that big corps are threatening to make them jobless with IP-disrespecting AIs
- bjourne 2y ago[flagged]
- kevingadd 2y agoPretty confident most of us haven't pirated 82tb of media in order to train an AI but maybe I'm naive?
- slig 2y agoI'm doing exactly that over the weekend with a RPi that I found in the junk drawer.
- miltonlost 2y agoLmao. Using A Jesus quote to defend a megacorp’s illegal piracy. Wow
- perihelions 2y agoBest way to "punish" Meta is to slash the Gordian knot and abolish copyright. Level the playing field, incrementally, for everyone else who isn't a trillion-dollar corporation. The alternative is a futile legalistic attack against a monopoly entity too powerful to be meaningfully punished. That won't accomplish anything useful. It would, rather, help cement this status quo, where copyright infringement is selectively legal or illegal, for different entities at the same time; and companies like Meta thrive arbitraging that difference. You can't defeat Meta—but you can help dig them a moat.
- miltonlost 2y agoRidding copyright would level the playing field for individuals and companies????!!!! Getting rid of laws that protect the individual only will help the larger empowered businesses.
- Workaccount2 2y ago>only will help the larger empowered businesses. I'm pretty sure I could list ten megacorps that would collapse overnight if copyright was abolished. The music groups, movie studios, streaming platforms...
- nkrisc 2y agoWhat's the alternative to copyright then? Anything I create will be instantly reproduced and sold for less than I can afford to by some entity far larger and more efficient than me. > Level the playing field, incrementally, for everyone else who isn't a trillion-dollar corporation. There is no level playing field when you have individuals and trillion-dollar companies in the same market.
- clueless 2y agoRight, all this talk about getting rid of copyright and no one is talking about what should replace it? how would we we incentives people to write good books? to pour 1000s of hours of their time to produce new knowledge?
- lrvick 2y agoThis should be legal. Copyright law does more harm than good. The only ethical problem here is that only Meta sized companies can afford to pay the "damages" for such blatant law violations at worst, or the fees of their lawyers at best.
- pleeb 2y agoIf an individual was the one tormenting almost 82 TB of copyrighted books, the damages they would have to pay would be in the trillions (mostly because of how broken the copyright law system is)
- maronato 2y agoCopyright law does more harm than good to individuals who just want to learn and enjoy content without profiting from it. Companies like Meta and OpenAI, however, should definitely have to pay to use the hard work of humans to train their AI.
- moffkalast 2y agoIf only these corporations with vested interests in permissive copyright would put their money where their mouth is with lobbying for a change. Or is that only allowed when they're trying to do something scummy? I forget.
- woadwarrior01 2y agoI wonder what happened to the related OpenAI training GPT3 on the books3 dataset story[1] from ~2 years ago? [1]: https://www.wired.com/story/battle-over-books3/ https://www.wired.com/story/battle-over-books3/
- gundmc 2y agoI think this one is different because the legality of training on copyrighted material is an open legal question while distributing/seeding copyrighted material is decidedly illegal.
- iimaginary 2y agoWe need better laws that would create a better way to do this legally whilst compensating rights holders.
- miltonlost 2y agoWe need better justice system that enforces the laws we have in the books that would help compensate right owners when big companies in emails pirate terabytes of data.
- SketchySeaBeast 2y agoI really don't think that Meta did this because the alternative would have been too onerous; they are a huge org, they could work through whatever loopholes required. They did it because it would have cost money and there will be no penalty for not paying.
- impossiblefork 2y agoSo, if they're sued in Japan, or France, do you think that the courts will take any special measures because it's a valuable American corporation? I suspect that if the case is reasonable they will just convict, and quickly-- appeal denied and all simply because the laws are so straightforward.
- SketchySeaBeast 2y agoI must have failed to clearly express myself - I don't think Meta should be doing what they are doing, I hope they do end up being punished. But the only way that Meta is going to change its behaviour is by being held accountable in a way that's much more difficult and costly than if they'd simply followed the law in the first place.
- impossiblefork 2y agoAh, I'm not sure exactly what I believe here, but this kind of torrenting is obviously illegal-- I'm personally split on how I feel about it morally, because some of these people really are trying to preserve knowledge, and I think that's commendable, at the same time, commercial piracy is something which really does screw over authors with it being some kind of theft-of-service type thing where people exploit other people's work-- and if they felt that the work had no value they could have written another text themselves. I only really wanted to convey that I believed that it probably isn't obviously easy for Meta to get away with anything in this, even if the US government decides to be lenient for the sake of a high market-cap US company simply because other countries are a viable place to sue as well. I think I misinterpreted your comment as that you thought that Meta thought that costs would be low because they imagined a US court system that simply ignored the illegality because it's they who committed it, when nothing like that is actually implied in your comment.
- belter 2y ago"Supposedly, Meta tried to conceal the seeding by not using Facebook servers while downloading the dataset to "avoid" the "risk" of anyone "tracing back the seeder/downloader" from Facebook servers, an internal message from Meta researcher Frank Zhang said, while describing the work as in "stealth mode." Meta also allegedly modified settings "so that the smallest amount of seeding possible could occur," a Meta executive in charge of project management, Michael Clark, said in a deposition..." They will be getting a lot of Frommer Legal letters...
- mik1998 2y agoLibgen is a civilizational project that should be endorsed, not prosecuted. I hope one day people will look at it and think how stupid we were today to shun the largest collection of literary works in human history.
- rafram 2y agoI think you’re overstating its importance. The internet already makes it possible to order almost any book in existence and have it arrive at your doorstep within a week or so, or often on your ebook reader instantly. And your local library probably participates in an interlibrary loan system that lets you request any book held by any library in the country for free. LibGen gives you access to a much smaller body of works than either of those. It’s a little more convenient. But the big difference is that it doesn’t compensate the author at all. Just go to a real library.
- mik1998 2y agoNo one sells scans of older books, which are often sparsely available in obscure (often private) libraries.
- rafram 2y agoSure, but I have a strong feeling that scans of out-of-print books only constitute a small portion of LibGen’s traffic. It’s like the idea that most BitTorrent users are just using it to share free software and Creative Commons media. (See the screenshots on every BitTorrent client’s website.) It would definitely be helpful if it were true, but everyone knows it’s just wishful thinking.
- crazygringo 2y agoWhy does the proportion matter? Academics are huge users of LibGen for academic books from the entire past century and beyond. It's infinitely more convenient to instantly get a PDF you can highlight, than wait weeks for some interlibrary loan from an institution three states away. Just because the majority of people might be downloading Harry Potter is irrelevant.
- Refusing23 2y agotheir whole business is stealing data.. so its quite funny to see they freely share it too.
- seydor 2y agoWe have at least 4 types of ill-defined concepts of property in the 21st century , largely due to our laziness, intellectual inertia and lack of motivation to make forward-thinking definitions for the coming age of AI and ubiquitous access to all information and all communication. 1) the concept of copyright is as old as the word suggests (copies are the least of our worries going forward - it should be possible to define processes for exploitation of ideas in a fair way) 2) we allow humans to learn from other people's ideas and transform them to commercial products and the same should happen for AIs in the future 3) we have an ill-defined concept of "personally identifying information" which gives people ownership to information that others have created via their own means - there should be better ways to ensure a level of privacy (but not absolute privacy) without overly-broad, nonsensical definitions of what is personally protected information 4) We allow social media and other telecommunications media to arbitrarily censor people's speech without recourse. This turns people's speech to property of the social media companies and imposes absolute power on it. This makes zero sense and is abusive towards the public at large. We need legal protections of speech in all media, not just state-owned media.
- thfuran 2y ago>we have an ill-defined concept of "personally identifying information" which gives people ownership to information that others have created via their own means - there should be better ways to ensure a level of privacy (but not absolute privacy) without overly-broad, nonsensical definitions of what is personally protected information What information about me could a corporation create via its own means that would be legally protected but shouldn't be? PII is generally information that a corporation collects. Unless you mean that my cellphone provider creates the association between my name and phone number and should therefore be able to do with it as they please?
- seydor 2y agoIt's not just about corporations. Banking and government services e.g. are required to keep your personal information stored for years and years even against your will
- wnevets 2y agoMy ISP will shut off my internet if it catches me torrenting copyrighted material but if you're a massive corporation that steals TBs of data its barely a blip in the news.
- hackerbeat 2y agoOne of the many reasons why Zuck’s been sucking up to Trump. He’s in desperate need of some Get-Out-Of-Jail-Free cards. Same for all the other sleazy tech bros.
- aucisson_masque 2y agoYou wouldn't download a car.
- rvz 2y agoMaybe you should go after the worst offender (OpenAI) first before going after Meta, since the latter already gave back their model away for free for everyone and the architecture. We will know why OpenAI isn't getting investigated.
- unraveller 2y agoCould be why OpenAI paid them so much, to go after their open-source competition hardest of all.
- deleted 2y ago[deleted]
- hruzgar 2y agoSo true. It seems like there is a controlled operation to shut open models down starting with Meta. Obviously they can't go after deepseek atm
- peterclary 2y agoI strongly urge people to read Thomas Babington Macaulay's speeches on copyright, its aims, terms, and hazards. Very well reasoned and explained. In particular, people often cited the case of authors who had died leaving a family in destitution, and claimed that copyright extension would be a fair way of preventing this, but in most cases the remaining family had never held the copyright; the author had initally sold the reproduction rights to a publisher who had then sat on the work without publishing it. The author, driven into penury, was then induced to sell the copyright to the publisher outright for a pittance. So in such cases a copyright extension only benefited the publisher, and indeed increased their incentive to extort the copyright.
- bbor 2y agoI’m a huge IP hater and am sure that happens, but to be fair, letting copyright extend past death also increases the amount the author can sell it for in the first place.
- ttyprintk 2y agoThe current workaround is to attribute footnotes to your beneficiaries, or quote them in the dedication. Those become derivative works subject to the lifetime of your beneficiary.
- kshri24 2y ago> Thomas Babington Macaulay The one who got Hindu Sanskrit books translated in a horrible manner and then claimed: "I have no knowledge of either Sanskrit or Arabic. But I have done what I could to form a correct estimate of their value. I have read translations of the most celebrated Arabic and Sanskrit works. I have conversed both here and at home with men distinguished by their proficiency in the Eastern tongues. I am quite ready to take the Oriental learning at the valuation of the Orientalists themselves. I have never found one among them who could deny that a single shelf of a good European library was worth the whole native literature of India and Arabia." This chap will educate us on copyright? No thanks!
- 2y ago
- jfbaro 2y agoThey are getting shittier and shittier
- snapcaster 2y agoThe powerful do what they can, the weak suffer what they must
- api 2y agoOne of the largest businesses of the Internet to date has been piracy. Individual informal piracy has been the smallest component of this. By far the largest has been corporate mass-scale piracy, and LLMs are probably the largest heist to date. They've literally downloaded the sum total of all human thought and knowledge, compressed it into queryable lossy compression models (which is what LLMs are), and are selling it back to us. Meta, with its "open weights" models, is one of the least guilty parties, since at least they've made the resulting blobs of mass piracy available to us. Same with Mistral, Deepseek, etc. ClosedAI, Google, and others have all probably done this and more and refuse to make even the model available. I think the way to deal with this is very simple: If you have trained your model on works to which you do not have rights or permission, the resulting model is not copyrightable and cannot be sold. It must either be kept for research purposes only or released free of charge and in the public domain. All these models that have been trained on pirated works should become public domain. Of course now that we have full capture of the US Federal Government I'm sure any suggestion like that would be neutralized with one bribe to Trump.
- HPsquared 2y agoIf you owe the bank $1,000 it's your problem; if you owe the bank $1,000,000,000 it's the bank's problem.
- fsflover 2y agoSupport EFF if you think that the copyright laws should be changed and also applied equally to all: https://www.eff.org/issues/innovation https://www.eff.org/issues/innovation
- maxwell 2y agoI'm sure they'll throw the book at them.
- caterwhal 2y agoReally strange how much torrenting is demonized by all of these companies and ISPs when individuals want to use it but when a company like Meta uses it there is so little scrutiny.
- toss1 2y ago>>"vastly smaller acts of data piracy—just .008 percent of the amount of copyrighted works Meta pirated—have resulted in Judges referring the conduct to the US Attorneys’ office for criminal investigation.".....While Meta may be confident in its legal strategy despite the new torrenting wrinkle... Zuckerberg has paid the vig several times [0,1,2], which is evidently the best legal strategy under this administration. OFC, considering there are already multiple payments, there is no assurance the vig payments won't substantially increase as the Capo sees more opportunity for profit. [0] https://en.wikipedia.org/wiki/Vigorish https://en.wikipedia.org/wiki/Vigorish [1] https://www.politico.com/news/2025/01/29/meta-settles-trump-facebook-ban-lawsuit-007810 https://www.politico.com/news/2025/01/29/meta-settles-trump-... [2] https://www.bbc.com/news/articles/c8j9e1x9z2xo https://www.bbc.com/news/articles/c8j9e1x9z2xo
- breppp 2y agoYes it smells bad but facebook did the right thing (at least for facebook) After OpenAI trained their models on the famed books2 dataset, and seeing the technological implications of ChatGPT, there was a good chance they would let them get away with it. Would the USA really surrender its AI technological advantage for trivial matters like copyright? They would make some royalty arrangement and get it over with
- z7 2y agoHow about a consequentialist argument? In some fields, AI has already surpassed physicians in diagnosing illnesses. If breaking copyright laws allows AI to access and learn from a broader range of data, it could lead to earlier and more accurate diagnoses, saving lives. In this case, the ethical imperative to preserve human life outweighs the rigid enforcement of copyright laws.
- KolmogorovComp 2y agoThere’s nothing particular to AI about your comment, it’s a general downside of IP.
- z7 2y agoNo, the development of an artificial general intelligence does seem like a special case compared to usual IP debates, particularly in the potential multiplicative positive-sum effects on society overall.
- imgabe 2y agoBoo hoo. We are trying to advance civilization here. To accumulate and make available all human knowledge to date. And you stand there with your hand out to stop this? You are a villain. There is no sympathy for you.
- zelphirkalt 2y agoCome on publishers! This is your chance! Now you can really show, how you will treat all copyright infringements equally and not only go after easy target. Show us, how you spend all that money in a lawsuit against Meta!
- ej1 2y ago[flagged]
- JW_00000 2y agoI don't understand why it's even a question that Meta trained their LLM on copyrighted material. They say so in their paper! Quoting from their LLaMMa paper [Touvron et al., 2023]: > We include two book corpora in our training dataset: the Gutenberg Project, [...], and the Books3 section of ThePile (Gao et al., 2020), a publicly available dataset for training large language models. Following that reference: > Books3 is a dataset of books derived from a copy of the contents of the Bibliotik private tracker made available by Shawn Presser (Presser, 2020). (Presser, 2020) refers to https://twitter.com/theshawwn/status/1320282149329784833 https://twitter.com/theshawwn/status/1320282149329784833. (Which funnily refers to this DMCA policy: https://the-eye.eu/dmca.mp4 https://the-eye.eu/dmca.mp4) Furthermore, they state they trained on GitHub, web pages, and ArXiv, which are all contain copyrighted content. Surely the question is: is it legal to train and/or use and/or distribute an AI model (or its weights, or its outputs) that is trained using copyrighted material. That it was trained on copyrighted material is certain. [Touvron et al., 2023] https://arxiv.org/pdf/2302.13971 https://arxiv.org/pdf/2302.13971 [Gao et al., 2020] https://arxiv.org/pdf/2101.00027 https://arxiv.org/pdf/2101.00027
- gameshot911 2y agoCritically, by torrenting they also directly distributed the copywritten material itself. That is a standalone infringement separate from any argument about trained LLMs.
- qup 2y agoAnd punishing them in the normal manner will be an incredibly small slap on the wrist, and do absolutely nothing to help us find out what will play out in court regarding a fair-use defense on training AI with copyrighted material.
- lucianbr 2y agoIsn't there a "fruit of the poisoned tree" kind of thing? Sounds to me quite similar to the situation where you would murder your parent and get to keep the inheritance, even if you are convicted of murder. Inheriting stuff isn't illegal, yet, I think most jurisdictions would not allow you to keep it in this case. There should be a problem with stuff obtained through illegal means, even if having that stuff is in principle legal. In this case, copyrighted material. Obviously they would argue that having the data is only a consequence of the download part, and that part is legal. What I see is that these situations are always complicated, and if you're rich enough, you get to litigate the complications and come out with a slap on the wrist or maybe even clean hands, while if you are an ordinary citizen, you can't afford to delve into the complexities and get punished. These days I'm starting to give up on the whole concept of the legal system being fair. They're not even pretending anymore.
- antirez 2y agoCopy-right is not learn/train-right. That said Meta full its mouth with open source while they release models that are not SOTA nor usable for commercial purposes.
- deleted 2y ago[deleted]
- StefanBatory 2y agoI as a individual would be liable to pay ~1000$ of damages if I'd downloaded a movie in Germany or Poland and the publisher would get to me. I'm going to assume as it's a corporation, then the laws no longer apply.
- Anamon 2y agoThat's okay, they should just charge The Zuck with it personally; I'd be fine with that.
- 1970-01-01 2y agoAnd they're going to get away with it simply because if you or I openly did this the DMCA fines would be for a million trillion dollars. Since Meta shareholders can't stomach a million trillion dollars in fines, their lawyers will wave their magic wands and poof! No laws were broken!
- sva_ 2y ago> By September 2023, Bashlykov had seemingly dropped the emojis, consulting the legal team directly and emphasizing in an email that "using torrents would entail ‘seeding’ the files—i.e., sharing the content outside, this could be legally not OK." I'm pretty sure you can theoretically download torrents without seeding, although this is frowned upon. If they really seeded (with full bandwidth?) that's indeed pretty brazen. It is sort of strange that Meta is being singled out here though, and sort of sad considering they at least release the model weights. What's the signal? Do illegal shit to be competitive, but make sure there is no evidence?
- voidUpdate 2y agoYou can, in transmission for example you can just set the seed percentage to 0%. I recognise that this makes me a bad torrenter, but I've been told in the past that my ISP wont be too happy about me seeding, and they already do something screwy to torrents I access through the surface web, so I'm just playing it safe
- sva_ 2y agoI think your client may still be sharing IP addresses, not sure about the legality of that
- reverendsteveii 2y agoSo they're gonna go through every book that was stolen and apply the appropriate penalty, right? Each copyrighted work has a minimum penalty of $750 under the DMCA. That will be applied fairly in order to ensure that the rights holder is made whole by the infringer, right? It's so funny to see the law blatantly ignored by the overlords. Like, there isn't even a pretext anymore. They just steal what they want and budget for the fines and campaign donations to make the consequences go away.
- nickpsecurity 2y agoThat they’d focus on file sharing over transformation or outputs is exactly the risk I warned the companies about in my AI report. Most datasets, like RefinedWeb and The Pile, also require sharing copyrighted workers between people who are not licensed to do that. Many works also prohibit commercial use or have patents on them. They need to make datasets which don’t have this problem or have entities in Singapore train the foundation models within their rules. The latter has a TDM exemption that would let AI’s use much of the Internet, maybe GPL code, licensed/purchased works they digitize, etc. Very flexible.
- ofou 2y agoWho would have known that BitTorrent, shadow libraries, and seeders will help to train the best AI models out there, that adds a whole new meaning to a "seed".
- 65 2y agoI'm more interested in piracy not being highly prosecuted than I am in Meta getting punished for this. I'm not trying to spend 20 years in jail for pirating a TV show.
- ngneer 2y agoSounds just like how Facebook got started, harvesting photos without permission. From the Wikipedia article, the Facebook precursor was known as Facemash. On Zuckerberg, "He hacked into the online intranets of Harvard Houses to obtain photos, developing algorithms and codes along the way. He referred to his hacking as "child's play."" If I were younger, I would be livid.
- lazycog512 2y agoabolish knowledge rentiers
- ezekiel68 2y agoUnless Meta 'fessed up to this (which seems unlikely), the headline here is missing the word "allegedly".
- papercrane 2y agoMeta admitted to the torrenting more than a month ago. The reason this is in the news is because some of the emails discussing it have been unsealed.
- ofslidingfeet 2y agoYeah well, OpenAI compressed the whole internet into proprietary weights and is now providing access via paid subscription while the original internet gets deleted from our culture.
- 999900000999 2y ago"Say they hood robin, ain't that a b*, take from the poor and give to the rich." - Ice Cube. Meta will face no consequences. Say your a small publisher and you'd like a bit of compensation. If you dare sue Meta can just blacklist your books on its platforms. Even if they don't, you probably don't have the money to sue one of the biggest companies on earth. I think copyrights should be limited to 25 years after first publication. This would fix plenty of issues and give the AIs of the world plenty to learn from. Who am I kidding, Meta will take what they will. For that author making 20k a year, be honored to be of use to Meta.
- bwfan123 2y agocan people vote with their feet, and leave the platform ? but the masses are addicted to the slop that meta feeds them.
- losvedir 2y agoHooray! Or wait, are we not doing that anymore?
- ej1 2y ago[flagged]
- dansitu 2y agoI'm fine with them using my books to train an open source model, but it would have been nice to be asked.
- deleted 2y ago[deleted]
- Der_Einzige 2y agoThe only bad thing about this is that small time players who do it are treated poorly (Aaron Swartz). IP de-facto not existing for AI companies is a feature, not a bug. The fact that most of the world embraced hardcore copyright troll ludditism when the means of their (badly paying creative) jobs economic production was democratized implies that most people do not believe in any "egalitarianism" and especially not the left-wing form many profess to believe in. Certainly not "information wants to be free" or any of the other idealist shit that I or Aaron Swartz believed in. What meta did was software communism - full stop. They literally released their models to the public! I support all of this 10000%. The only issue is that they're not open enough (fully open source the dataset) So, unironically, good! Thank you, please pirate more! Please destroy the US IP system while you're at it. Copyright abolitionism is good and thank you Zuckerberg!
- zackmorris 2y agoIs there a concept in the legal system of first-come-first-served that could be used as precedent? What I mean is: when someone is prosecuted for copyright infringement, but Meta isn't, then could the case be put on hold until Meta is found guilty and pays a fine? Also maybe the fine on the later case would have to be proportional to the prior case. So if Meta pays $1 per infringement, the penalty might be $1 for torrenting something else (which is immaterial and not worth the justice system's time) so pretty much all copyright infringement cases would get thrown out. It reminds me of how mainstream drug addicts get convicted and spend years in prison, while celebrities get off with a warning or monetary fine.
- hnfong 2y agoLawyers (and hence, judges) are really good at arguing why the earlier case does not apply in a present case, even if most reasonable people would think the two cases are essentially the same. It's a fundamental part of lawyer training, and if they want to let BigCorp go and bring the hammer down on the little guy, they can make up a hundred reasons for it.
- kpgraham 2y agoDamn! One of my old books can be found in the Anna's Archive search. The book has been out of print for years. I pity the Meta users who get results based on something that I wrote. (Check Anna's for 'Keith P. Graham', and the first book listed is mine.)
- yalogin 2y agoLLMs are worse than search for figuring out what value a specific asset provides to the LLM. Atleast with search your work or page is not lost and still gets a click/user interaction, and may be give you a chance to monetize the interaction. However, LLMs just don’t have any such option. Gemini adds links but the links they add are completely editorialized by the LLM and need not reflect the original at all. So how does anyone ask for compensation even if they sue?
- bigmattystyles 2y agoThe question is, if they could and would have paid for each book, would it be ok to train the LLM on them? I'm talking about prior books, I'm sure new books have language forbidding their use to train LLMs at the point of sale. But legally, how does using a book to train a LLM differ from a teacher learning from a book and teaching its contents to their pupils. Obviously, the LLM can do so at scale, but is there a legal difference?
- CryptoBanker 2y agoA LLM is not a person. That is the legal difference...until we have Citizens United v2
- dragonwriter 2y ago> The question is, if they could and would have paid for each book, would it be ok to train the LLM on them? Whether training on AI model on an array of diffentent works, many of which are copyright protected, is itself a copyright violation, in addition to or distinct from any copyright violation that goes on gathering the dataset for training (and separate from any copyright violation in the actual or intended use of the LLM), remains to be resolved as a legal question, and may or may not have a simple yes or no answer (or the same answer under every system of copyright laws globally). My inclination is that it is probably generally not a violation in US law, but that's not something I am very confident in; how the definitions of copy and derivative work apply to determine if it would be without fair use, and how fair use analysis applies, are not clear from the available precedent. > But legally, how does using a book to train a LLM differ from a teacher learning from a book and teaching its contents to their pupils. It is very clear, by looking at how US copyright law is written and even more clear in its history of application, that information stored in brains of people are without exception neither copies nor new works that can be derivative works under US law, and so cannot be infringing, no matter how you gain them. It’s also very clear in the statute itself and the case law that data in media used by artificial digital computers, on the other hand, can constitute copies or derivative works that can be infringing. Even if the process is arguably similar in legally relevant manners, copyright law is critically focussed on the result and whether it is a particular kind of thing which can be infringing, not just the process.
- djyaz1200 2y ago“Behind every great fortune lies a great crime” -Honoré de Balzac
- liendolucas 2y agoFor some misterious reason I can't see Zuckerberg in front of a judge facing 50 years imprisonment. Anyone can? I truly hope that whoever takes the case goes after Meta with 1000 times the pressure that was put on Swartz, but honestly I don't expect much just as the top comment precisly expressed. And if we are going to be fair please also let's not forget about the other usual suspects, or anyone thinks they are falling behind?
- impossiblefork 2y agoThere are other countries than the US though and if rightsholders wish to sue, lawsuits can happen there too. Several EU countries, Switzerland, South Korea, Japan, etc. are viable countries to sue from. Even in Japan which has a law specifically permitting training on copyrighted material you must still obtain it legally-- i.e. you must license it.
- hnfong 2y agoThat's irrelevant. Switzerland (for example) isn't going to arrest Zuckerberg and put him in jail for this either. Nobody will. But if you're operating a site called Pirate Bay or something like that and it's not earning billions of dollars, expect countries to chase you across the globe trying to arrest you.
- impossiblefork 2y agoIf there were a criminal prosecution for willful copyright infringement in some non-US country an extradition request for the relevant people is not an impossibility though, and there would be no legal reason to deny it.
- lvl155 2y agoI’d think people can get together to put this on a public space strictly for training purposes and have the consortium of some sort get paid per use. But we live in this stupid society where you have to move mountains to change things an inch.
- scotty79 2y agoSeeding it was probably most societally useful thing Meta ever did.
- mrinterweb 2y agoRemember people getting sued insane amounts of money per-song they torrented. If we applied that precedent to Meta, Meta would need to declare bankruptcy. https://www.cbsnews.com/news/file-sharing-mom-fined-19-million/ https://www.cbsnews.com/news/file-sharing-mom-fined-19-milli...
- Pxtl 2y agoLaws are for poor people.
- srameshc 2y agoAt OpenAI we have seen some employees expressed their concern publicy about the moral grounds on which company was acting. We never heard about it from anyone at Meta but there were some jokes ofcourse. I guess everything is fair in AI and Corporates.
- josefritzishere 2y agoZuckerberg did more copyright infringement? Shocking!
- peterbonney 2y agoThe more I learn about how AI companies trained their models, the more obvious it is that the rest of us are just suckers. We're out here assuming that laws matter, that we should never misrepresent or hide what we're doing for our work, that we should honor our own terms of use and the terms of use of other sites/products, that if we register for a website or piece of content we should always use our work email address so that the person or company on the other side of that exchange can make a reasonable decision about whether we can or should have access to it. What we should have been doing all along is YOLO-ing everything. It's only illegal if you get caught. And if you get big enough before you get caught then the rules never have to apply to you anyway. Suckers. All of us.
- clueless 2y agoyep, pretty much.
- wrs 2y agoAnd if you were in any doubt before, this lesson is now exemplified by the holder of the highest office in the land and approved by popular vote. The rewards of acting ethically are, unfortunately, sometimes only personal. This must be a hard environment to raise children in, given the examples they see around them.
- formerphotoj 2y agoTHIS.
- hamburga 2y agoParent here: it takes a lot of discussion, but it's a great time to talk about the reality of evil and villains. My kids are on the good side, or at least I like to think so ...
- a123b456c 2y agoThis argument may focus too much on the category of external rewards. I might well be kidding myself or self-justifying, but I believe internal rewards are at least as important. Some materially successful people are deeply unhappy.
- deleted 2y ago[deleted]
- swozey 2y agoI deleted my facebook account about 10 years ago. Downloaded data, deleted. Not deactivated. Nothing in my life made me ever want to go back except for when I got back into playing hockey, and all the hockey leagues use facebook to communicate a few months ago. I made a new account, had to literally upload a picture of my face to pass verification.. and then a few days later I was immediately banned and couldn't use my account. I assume because they searched previous data and compared my face to find out I have a "deleted" (lol) account and matched me. I've assumed they'll only let me log in if i use my original 10 years ago deleted account. Fuck meta. Fuck zuck.
- deleted 2y ago[deleted]
- black_puppydog 2y agoWouldn't it be a real shame if the entirety of US constitution, laws, and legal precedence went out the window these days, and the only thing left unscathed was the rotten mess that is copyright law? Just saying, this might be the moment to burn it to the ground. Not that it makes up for any of the other stuff going on, but why waste a perfectly good crisis?
- abigail95 2y agoThis reminds me of Peter Sunde's "komimashin" https://www.engadget.com/2015-12-21-peter-sunde-kopimashin.html https://www.engadget.com/2015-12-21-peter-sunde-kopimashin.h... It's obviously absurd to enforce copyright as bytes are copied around instead of as it is used. Training an LLM is a different thing than re-hosting and giving away copies to other people. If you don't want people to transform your works - keep them private. You don't own ideas.
- golly_ned 2y agoAs the article says, Meta /was/ giving away copies to other people by seeding the libgen torrents. This isn't the usual case of "should companies be allowed to train on books".
- abigail95 2y agoThen it's a simple case of a rights holder taking them to court. What's the fuss about LLM training in this thread then?
- henriquemaia 2y agoThanks for the link. I wondered what that word meant. From the article: Kopimashin, as in Copy Machine.
- ocean_moist 2y agoAt least they seeded!
- stevage 2y agoWow, I'm actually a bit shocked that senior levels of management at Meta were fine with torrenting pirated books. WTaF. Meta does a lot of stuff I disagree with, but they're usually not just straight breaking the law.
- thunder-blue-3 2y agoYou know the wierd thing is - I've never used Meta AI. I've never thought of using it. The only product of FB i use is whatsapp, however I've not seen/heard any of my friends using Meta AI for FB,IG,Whatsapp. I really don't understand what their ROI here is...
- lewdev 2y agoIt's okay when large corporations download cars. But when you do it, you'll be in trouble.
- lakomen 2y ago[flagged]
- dang 2y agoYou can't post like this here, regardless of whom you want to kill. Between this and other abusive comments, we've banned the account. https://news.ycombinator.com/item?id=42946919 https://news.ycombinator.com/item?id=42946919 https://news.ycombinator.com/item?id=42690711 https://news.ycombinator.com/item?id=42690711 If you don't want to be banned, you're welcome to email hn@ycombinator.com and give us reason to believe that you'll follow the rules in the future. They're here: https://news.ycombinator.com/newsguidelines.html https://news.ycombinator.com/newsguidelines.html.
- elzbardico 2y agoNothing is gonna happen. Just a slap on the hand. And we all from the intelectual work class, writers, journalists, programmers will be proletarized by LLMs that have been: a) Financed via inflation/"cantillon effect" due to ZRP/Stimulus that absolutely flooded the market with funny money in the hand of the sharks. b) Trained upon copyrighted work without compensation. c) Trained upon open source without even asking politely for authorization. The Robber Barons from the last century can't even get close to our modern Feudal Tech Lords. Unless you're one of us that have amassed multi-generation wealth in a exit in the last 20 years, you're completely fucked.
- waltercool 2y agoBased. Free knowledge to the people
- buyucu 2y agoI love this. Large corpos should torrent more. Maybe we'll get better copyright law as a result.
- asjir 2y agoI thought about it for a full day, and I have one idea for how to handle copyrighted data training. It would need to be open / regulated and training till double descent would need to be disallowed, to make sure that the model is not memorizing the data.
- cratermoon 2y agoWe're starting to find out that Meta ruined LibGen for the rest of use who used it like a library. Just like how Google screwed over libraries by sending interns to the Stanford library to checkout books they scanned into Google Books. Not to increase shared knowledge or preserve human artificats, but to put them all in a museum and, to paraphrase Joni Mitchell, charge the people a dollar and a half just to see 'em.
- pjfin123 2y agoCopyright law needs major reform. We need to figure out a way to let authors monetize their work while not making complying with the law so burdensome. We've created a system where people who (understandably) ignore the law benefit at the expense of people trying to do the right thing.
- bloopbloopscoop 2y agoDeath to intellectual property!
- MoneCollinsgdjd 2y ago[dead]
- pilimi_anna 2y agoWe're grateful to Meta for helping seed and backup our torrents. The more copies the better. Thank you Meta, for helping preserve humanity's legacy! :)
- esarbe 2y agoIt's okay - they are multi-billion company. Rules don't apply to them. Rules are just for us peasants.
- kelseyfrog 2y agoThe usual copyright cartel is up in arms, crying theft. But here’s the truth: intellectual property is a state-enforced monopoly, not real property. Property is based on scarcity - if you take my car, I no longer have a car. But if you copy my book, I still have my book. No loss, no theft, just an outdated legal fiction designed to stifle innovation and enrich rent-seeking middlemen. An no, loss of potential sales doesn't count - it's like being able to claim a lottery ticket has real value. Copyright was never about protecting creators—it’s about locking down ideas, preventing competition, and extracting endless fees. Shakespeare borrowed, tech companies iterate, and science thrives on free exchange. The idea that knowledge should be locked away indefinitely is absurd. Meta’s mistake wasn’t using the data - it was pretending copyright still matters. AI is exposing the system for what it is: obsolete. The future belongs to those who create without asking permission.
- nullfield 2y agoI think everyone can see that whatever (imo not in accordance with the Constitution, after absurdities like deciding “limited time” the way mathematicians might define something of some order of infinity) the alleged social contract was is not functional the way it was intended, and we see who benefits and who loses. mass dynamic editing for vitriol and profanity occurred while writing this comment in order to remain within site rules
- flojo 2y agoDid they at least seed back?