5 ms·
AI training and the thing search engines do to make a search index are essentially the same thing. Hasn't the latter generally been regarded as fair use, or els
by zrm 11mo ago
AI training and the thing search engines do to make a search index are essentially the same thing. Hasn't the latter generally been regarded as fair use, or else how do search engines exist?
- Kye 11mo agoThere was a relatively tiny but otherwise identical uproar over Google even before they added infoboxes that reduced the number of people who clicked through.
- tpmoney 11mo agoThere was also the lawsuit against google for the Google Scholar project, which is not only very similar to how AI use ingest copyright material, but even more than AI actually reproduced word for word (intentionally so) snippets of those works. Google Scholar is also fair use.
- AnthonyMouse 11mo ago> There was a relatively tiny but otherwise identical uproar over Google even before they added infoboxes that reduced the number of people who clicked through. But is that because it isn't fair use or because of the virulent rabies epidemic among media company lawyers?
- Kye 11mo agoThis was normal people, as much as bloggers on the pre-social media early web could be considered normal.
- AnthonyMouse 11mo agoNormal people that aren't media companies were objecting to search engines indexing websites? That seems more likely to have been media companies using the fact that they're media companies to get people riled up over a thing the company is grumpy about.
- deleted 11mo ago[deleted]
- freejazz 11mo agoI don't think regular people pay attention to copyright decisions (they don't even pay attention to the cases to make it to the supreme court) but there are plenty of lawyers who don't work for media companies who disagree with the findings. I also think your characterization is ridiculous and pejorative.
- AnthonyMouse 11mo agoThey disagree with search engines being fair use? The general problem is that both the structure of copyright and the legacy media business model were predicated on copying being a capital-intensive process. If a printing press is expensive then reproduction is a good place to collect royalties, because you could go after that expensive piece of equipment if they don't pay. And if a printing press is expensive then a publisher who has one is offering a scarce service in a market with a high barrier to entry. The internet made copying free and that pretty well devastated the publishing industry, more as a result of the second one than the first. If your product isn't scarce -- if your news reporting is in competition with every blog and social media post -- you're not getting the same margins you used to. But there's no plausible way the incumbents are going to convince people that reporters with a website instead of a printing press need to be excluded from the market so they can have less competition, and that by itself and nothing more means their traditional business model is gone. They're competing for readers and advertisers against Substack and Reddit and the cat's not going back in the bag. Meanwhile copyright infringement got way easier and that's much more plausible to frame as a problem, so the companies want to sic their lawyers on it, except that the bag is here on the ground and the cat is still over there getting a million hits. There is no obviously good way to solve it (but plenty of bad ways to not solve it) and solving it still wouldn't put things back the way they were anyway. So their lawyers are constantly under pressure to do something but none of their options are good or effective which means they're constantly demanding things that are oppressive or asinine or, like the anti-circumvention clause in the DMCA, own-goals that tech megacorps use against content creators to monopolize distribution channels. Which is why it's an epidemic. If you can see the target the pressure is on to pull the trigger even when all you have is a footgun.
- justapassenger 11mo agoMost important part of fair use is does it harm the market for the original work. Search helps to brings more eyes to the original work, llms don't.
- tpmoney 11mo agoThe fair use test (in US copyright law) is a 4 part test under which impact on the market for the original work is one of 4 parts. Notably, just because a use has massively detrimental harms to a work's market does not in and of itself constitute a copyright violation. And it couldn't be any other way. Imagine if you could be sued for copyright infringement for using a work to criticize that work or the author of that work if the author could prove that your criticism hurt their sales. Imagine if you could be sued for copyright infringement because you wrote a better song or book on the same themes as a previous creator after seeing their work and deciding you could do it better. Perhaps famously, emulators very clearly and objectively impact the market for a game consoles and computers and yet they are also considered fair use under US copyright law. No one part of the 4 part test is more important than the others. And so far in the US, training and using an LLM has been ruled by the courts to be fair use so long as the materials used in the training were obtained legally.
- willis936 11mo ago1. Character of the use. Commercial. Unfavorable. 2. Nature of the work. Imaginative or creative. Unfavorable. 3. Quantity of use. All of it. Unfavorable. 4. Impact on original market. Direct competition. Royalty avoidance. Unfavorable. Just because the courts have not done their job properly does not mean something illegal is not happening.
- tpmoney 11mo agoAll of these apply to emulators. * The use is commercial (a number of emulators are paid access, and the emulator case that carved out the biggest fair use space for them was Connectix Virtual Game Station a very explicitly commercial product) * The nature of the work is imaginative and creative. No one can argue games and game consoles aren't imaginative and creative works. * Quantity of use. A perfect emulator must replicate 100% of the functionality of the system being emulated, often times including bios functionality. * Impact on market. Emulators are very clearly in direct competition with the products they emulate. This was one of Sony's big arguments against VGS. But also just look around at the officially licensed mini-retro consoles like the ones put out by Nintento, Sony and Atari. Those retro consoles are very clearly competing with emulators in the retro space and their sales were unquestionably affected by the existence of those emulators. Royalty avoidance is also in play here since no emulator that I know of pays licensing fees to Nintendo or Sony. So are emulators a violation of copyright? If not, what is the substantial difference here? An emulator can duplicate a copyrighted work exactly, and in fact is explicitly intended to do so (yes, you can claim its about the homebrew scene, and you can look at any tutorial on setting up these systems on youtube to see that's clearly not what people want to do with them). Most of the AI systems are specifically programmed to not output copyrighted works exactly. Imagine a world where emulators had hash codes for all the known retail roms and refused to play them. That's what AI systems try to do. Just because you have enumerated the 4 points and given 1 word pithy arguments for something illegal happening does not mean that it is. Judge Alsup laid out a pretty clear line of reasoning for why he reached the decision he did, with a number of supporting examples [1]. It's only 32 pages, and a relatively easy read. He's also the same judge that presided over the Oracle v. Google cases that found Google's use of the java APIs to be fair use despite that also meeting all 4 of your descriptions. Given that, you'll forgive me if I find his reasoning a bit more persuasive than your 52 word assertion that something illegal is happening. [1]: https://fingfx.thomsonreuters.com/gfx/legaldocs/jnvwbgqlzpw/ANTHROPIC%20fair%20use.pdf https://fingfx.thomsonreuters.com/gfx/legaldocs/jnvwbgqlzpw/...
- asdefghyk 11mo agore ".....AI training and the thing search engines do to make a search index are essentially the same thing. ...." Well, AI training has annoyed LOTS people. Overloaded websites.. Done things just because they can . ie Facebook sucking up content of lots pirate books Since this AI race started our small website is constantly over run by bots and it is not usable by humans because of the load.. NEWER HAD this problem before AI , when just access by search engine indexing .....
- AnthonyMouse 11mo agoThis is largely because search engines are a concentrated market and AI training is getting done by everybody with a GPU. If Google, Bing, Baidu and Yandex each come by and index your website, they each want to visit every page, but there aren't that many such companies. Also, they've been running their indexes for years so most of the pages are already in them and then a refresh is usually 304 Not Modified instead of them downloading the content again. But now there are suddenly a thousand AI companies and every one of them wants a full copy of your site going back to the beginning of time while starting off with zero of them already cached. Ironically copyright is actually making this worse, because otherwise someone could put "index of the whole web as of some date in 2023" out there as a torrent and then publish diffs against it each month and they could all go download it from each other instead of each trying to get it directly from you. Which would also make it easier to start a new search engine.
- soco 11mo agoGoogle doesn't offer for own gains copies of existing websites (except they do that lately as well)
- freejazz 11mo agoWeird, AI companies insist that AI models are not just indexes but instead something the model has "learned". So, again, to answer my question, it's certainly not a settled matter of law that AI models and/or their "training" is actually akin to a search engine such that it amounts to a fair use. So how is it that the EFF is reporting it like a fact?