6 ms·
Thanks Elon, now we know that scraping is illegal! Very good to clarify that for future proceedings against the AI thieves.
by zqwt3k 20d ago
Thanks Elon, now we know that scraping is illegal! Very good to clarify that for future proceedings against the AI thieves.
- wahnfrieden 20d agoIt’s the rehosting of the content not the scraping Edit: not a moral stance
- zqwt3k 20d agoThat is what all LLMs could do in 2023, verbatim, before they trained it out of them in order to keep up the pretense that there is no plagiarism. Now they all obfuscate the original or refuse to cite.
- echelon 20d agoGrok can do that if you give it an HN thread or other websites.
- jm4 20d agoI love how the content belongs to them when someone else reposts it but it belongs to the user if the content is illegal. Such a double standard with these social media and AI companies. Why do we put up with it?
- embedding-shape 20d ago> Why do we put up with it? You let it happen. Once people stop letting it happen, it'll stop. But social media is apparently the new "opium of the masses" so here we are and no one wants to do anything.
- flaburgan 20d agoI actually do want to do something. I started to scrape Twitter to make it freely available. The web should stay open.
- immibis2 20d agoI am also doing something, by hosting one of the public instances of Nitter. Nitter is useful for sporadic random access to tweets, but for public feeds like municipal authorities etc. it would be useful if someone scraped the feed and re-hosted the feed from their own server, without being hobbled by rate limits. Is that what you're doing?
- simianparrot 20d agoDid the users of X consent to XCancel copying their posts to their servers?
- conception 20d agoEveryone who goes to twitter copies the posts to their computers.
- 27183 20d agoSpot on. This is where a lot of these "terms and conditions" break down logically. Viewing some content on the internet is literally copying it. So is the distinction that xcancel served the content? But when I run mtr xcancel.com I see a bunch of hops between me and them. Every one of those hops is literally copying and retransmitting all the content. Are they not also serving it?
- immibis2 20d agoNo, this is where programmers rules-lawyer in ways that actual lawyers don't and then get law stuff hilariously wrong. No judge thinks that viewing an HTML page is downloading it, because downloading means saving a copy to your computer, not just looking at it. Even having an internet cache folder doesn't count as downloading. Even copying the file from the internet cache folder to somewhere might not count as downloading, although it'd still be a copy. Same as when LG said their TVs don't record you and then Hacker News said "how can they detect voice commands if they don't record your voice"... facepalm.
- 27183 20d agoI don't pretend to understand law, mostly it just doesn't make sense at all.
- rimunroe 20d agoCould you elaborate in what way you find the law mostly doesn't make sense? It has to be flexible in order to work with actual humans. Why should visiting a page on your computer count as copying? Usually when we talk about copying it's someone making a duplicate so it can be accessed later. Only a very technical user is going to be diving into their cache to view that content after the fact. The vast majority of people don't understand that the browser is storing anything on their computer, much less how to access it before it's purged.
- Levitz 20d agoWe put up with it mostly because it enables a lot of communication to happen at all. You want to hold the corporations legally liable for the content their host, you can say goodbye to basically reddit as a whole, any twitter clone, youtube comment sections and a whole more stuff.
- HumblyTossed 20d agoThere's a button on X that allows me to repost someone else's content.
- bel8 20d agothere's a setting to disable that
- wahnfrieden 20d agoThat’s not rehosting Edit: iPhone autocorrected my OP which meant to say rehosting not reposting
- blitzar 20d agoThats "Billionaire Use" -- its like a "Fair Use" exemption but for billionaires.
- HumblyTossed 20d agoAaron Swartz died so that "AI" billionaires can live.
- rich_sasha 20d agoI wonder what legal gymnastics are needed for "I can scrape anything off the web ignoring copyright and build a product from this, but you can’t even display what’s on my webpage elsewhere”. Perhaps that is it in fact. The act of protecting it from scraping means you object. 99% of the blogged contents etc. Big AI helped themselves to was just… there. Public. Not free from copyright but still not paywalled.
- embedding-shape 20d agoTransformation. Taking something someone else made and showing it as-is, bypassing their own restrictions: No no. Taking something someone else made, modify it or use parts of it in some bigger thing or completely change it: Fine, if you have money and/or run a company
- bluefirebrand 20d agoSo in theory if you took twitter content and then transformed it so it "summarizes" all tweets with an AI rather than posting the exact text, would that be allowed? Because that's stupid. These laws are stupid.
- brainwad 20d agoThat's exactly the way UK courts are heading, see Getty vs Stability AI. The court ruled that there's no infringment because the model doesn't store exact copies, just derived weights, and therefore when it generates new images those aren't copies of protected works.
- bluefirebrand 20d agoThat's stupid, these courts are stupid It should have nothing to do with storing copies it should have to do with what the models can produce. And it's clear they can produce copyrighted works, they've just been tuned so they don't. That shouldn't satisfy anyone.
- immibis2 20d agoEverything is both legal and illegal until a lawsuit happens. Then it collapses depending a little bit on the facts and mostly on who has the better lawyers. I suspect Nitter's first round with lawyers pointed out that scraping is legal, but now they have been threatened with something else than scraping - Elon claims something else the way Nitter runs is illegal, such as the use of fake accounts to circumvent an access control device (DMCA 1201).
- dannyw 20d agoAt the end of the day, as an individual or a team or a company, regardless of the statue and case law, you have to perform the calculus on your monetary and legal resources versus your counterparty. Obviously ungrounded and frivolous cases tend to be easier to defend asymmetrically, but if I was X's legal team, there's no shortage of semi-plauisble claims I could throw at the wall and see what sticks. As an example of this imbalance in action, BrightData is a 'gray area' company that basically does this exact kind of scraping. They have somehow won against Meta Platforms suing them, and even got X's lawsuit against them for scraping -- identical (?) activity to XCancel -- dismissed. According to Wikipedia: > In May 2024, a federal judge dismissed the suit, ruling that Bright Data did not violate X's terms of service or copyright by scraping publicly accessible data.[24] The judge emphasized that such scraping practices are generally legal and that restricting them could lead to information monopolies But does XCancel have the resources of a company like Bright Data, that's funded and used by companies like Deloitte and Moodys?
- immibis2 20d agoNotice it says "terms of service or copyright". If X's lawyers have any intelligence, they'll have a reason why XCancel is not identical to Bright Data. Perhaps this time, instead of claiming it's a copyright violation, they'll claim it's wire fraud because multiple accounts are used.
- tehwebguy 20d agoIf the name Bright Data is ringing a bell to anyone, it’s probably because they are a (the?) primary offender running the LG TV “residential proxy” (botnet)
- znpy 20d agoscraping content is mostly legal, redistributing content is not. if you started doing the same to, say, instagram content both meta and individual creators would sue you as well. sites like archive.ph are in a similar bucket btw, and yet nobody's complaining (except websites seeing people evading their paywall). but at the end of the day it's not really fair to apply laws differentially on the basis of whose political ideas we like more.
- immibis2 20d agoMeta has no exclusive rights to the content on Instagram, and X has no exclusive rights to the content on X. They have a non-exclusive license to republish it, etc.
- alterom 20d ago>scraping content is mostly legal, redistributing content is not. So, I take it some Twitter users took an issue with their tweets being redistributed by another platform and sued XCancel? Because surely you're not claiming that Twitter has any ownership of what gets posted on that platform, are you? >if you started doing the same to, say, instagram content both meta and individual creators would sue you as well. By that logic, Meta could sue me for posting my photos on other platforms. Meta doesn't have a case here, and neither does Twitter, but something tells me Twitter is involved in this nevertheless.
- Nemo_bis 19d agoThe copyright implications of displaying Instagram content elsewhere are still unclear. There were many lawsuits over Instagram embeds. https://copyrightlately.com/legal-embed-instagram-photos-website/ https://copyrightlately.com/legal-embed-instagram-photos-web...
- dismalaf 20d agoPretty sure it's legal when you do it for your own use (same as browsing a website) but it's illegal to redistribute web scraped results.
- burnte 20d agoIt's not that cut and dry or else search engines wouldn't be legal. It depends on how much is used, for what context, etc. This very well may wind up being fair use.
- dismalaf 20d agoSearch engines "modify" it ie. show snippets + direct to the actual site. In general fair use pretty much always requires it to be transformative and/or point to the source. Simply scraping it to prevent people from going to X isn't free use in any definition I've heard.
- gosub100 20d agoRemember ~10 years ago when Google had "cached" versions of the websites? I believe they removed that feature for this reason.
- gosub100 20d agoNow I want to know, what happens if you redistribute an "AI summary" of the copyright material?
- Quitschquat 20d agoThin skins over at there at X