4 ms·
I wonder what legal gymnastics are needed for "I can scrape anything off the web ignoring copyright and build a product from this, but you can’t even display wh
by rich_sasha 18d ago
I wonder what legal gymnastics are needed for "I can scrape anything off the web ignoring copyright and build a product from this, but you can’t even display what’s on my webpage elsewhere”.
Perhaps that is it in fact. The act of protecting it from scraping means you object. 99% of the blogged contents etc. Big AI helped themselves to was just… there. Public. Not free from copyright but still not paywalled.
- embedding-shape 18d agoTransformation. Taking something someone else made and showing it as-is, bypassing their own restrictions: No no. Taking something someone else made, modify it or use parts of it in some bigger thing or completely change it: Fine, if you have money and/or run a company
- bluefirebrand 18d agoSo in theory if you took twitter content and then transformed it so it "summarizes" all tweets with an AI rather than posting the exact text, would that be allowed? Because that's stupid. These laws are stupid.
- brainwad 18d agoThat's exactly the way UK courts are heading, see Getty vs Stability AI. The court ruled that there's no infringment because the model doesn't store exact copies, just derived weights, and therefore when it generates new images those aren't copies of protected works.
- bluefirebrand 18d agoThat's stupid, these courts are stupid It should have nothing to do with storing copies it should have to do with what the models can produce. And it's clear they can produce copyrighted works, they've just been tuned so they don't. That shouldn't satisfy anyone.
- deleted 18d ago[deleted]
- immibis2 17d ago> they can produce copyrighted works, they've just been tuned so they don't. In other words... they can't. A different one can but this one can't. The court is not stupid, and will consider this fact.
- butlike 17d agoThey can't unless the tuner produces a copyrighted work for whatever their purpose would be, because they can step in and "detune" the thing at 3am for a competitive advantage.
- embedding-shape 17d agoThis assumes the "tuning" (I assume you mean fine-tuning?) is lossless, which it isn't.
- scottyah 17d agoYou can produce copyrighted works and have just been tuned not to lol. I fail to see how limiting an ability to comply with the law is any different from just complying with the law?
- bluefirebrand 17d agoProprietary Computer systems shouldn't get the same leeway that humans do. Simple as that.
- dismalaf 18d agoThe point is that you can't steal someone else's content 1:1. But you can use it for a different use (say, display the tweet in an article, then comment on it).
- 8note 17d agotwitter didnt make it though. they have a license to it
- brainwad 18d agoPrecedent is pretty clear: competitive uses bad, transformative uses good. Xcancel scrapes and then competes directly with X, whereas LLM labs scrape the internet to make an agentic intelligent bot, a transformative use of the scraped content.
- bonsai_spool 18d ago> Precedent is pretty clear: What cases are you citing when you say this?
- immibis2 18d agoPerhaps https://en.wikipedia.org/wiki/Warner_Bros._Entertainment_Inc._v._WTV_Systems,_Inc https://en.wikipedia.org/wiki/Warner_Bros._Entertainment_Inc....
- brainwad 18d agoBartz v Anthropic. Though the plaintiffs did get something, it was because of the piracy to the original works (competing against the legal market for the books), not the use of them to train the LLM.
- deleted 18d ago[deleted]
- lesuorac 18d agoBartz is an author though. Is X claiming ownership of the posts people make because pretty much every single social media site doesn't so they have section 230 protection.
- brainwad 18d agoThey can just round up some friendly users and sue under their names. Starting with their own corporate accounts?
- htrp 18d agoexcept a bunch of paywalled stuff did end up in training corpora