3 ms·
I disagree, consent is important here, and the "move fast and break things" mentality is what's causing problems. #1 It's not all that certain AI training is f
by ADeerAppeared 2y ago
I disagree, consent is important here, and the "move fast and break things" mentality is what's causing problems.
#1 It's not all that certain AI training is fair use. There could be significant damages if it is found to not be fair use.
#2 "We're going to steal your shit even if you won't give us consent to do that" is a bad idea. We're already seeing the practical results: People stop/reduce posting their works to the clearnet. What's your scraper bot going to do? Create an account on every 'private' forum and immediately get sued? Pay a subscription fee to every single Patreon page?
And the big one, #3: It destroys public support for your technology, which is essential if you want to survive the oncoming government regulation.
Look at the public response to various tech companies stopping their AI rollout in europe over EU regulations. Varying from "Yes, this is exactly what we asked for" to "Good fucking riddance".
And it will get only worse yet. Europe's privacy authorities are already issuing quiet statements that the scraping of social media posts, such as is done for AI data collection, is not legal under the GDPR. (The law's pretty clear on this, it doesn't matter that it was "posted publicly", you're not allowed to use personal data like that) There's already whispers of going after the LLMs themselves, as they contain and continue to process personal data as well.
AI needs public support, and the lack of consent is slowly bleeding it dry.
- CharlieDigital 2y ago> What's your scraper bot going to do? Create an account on every 'private' forum and immediately get sued? More likely, that the owners of any repository of note will simply broker that data directly (a la Reddit).
- ADeerAppeared 2y agoYes, though this scenario isn't all that interesting to discuss: * In the case of big IP holders (e.g. media companies, news organizations) this is just obtaining consent. The only fun quirck is that OpenAI's purchasing of this drastically lowers the "AI training is fair use" claim by proving the existance of a market. * In the case of platforms like Reddit, it just kicks the problem one layer down. The platform does obtain "consent" through it's ToS. (Beware that this consent is legally weak, and won't protect your ass from anything outside copyright) But users will still see it as "stealing" and may flee the platform. There's still a notable shift away from the clearnet.
- godelski 2y ago> the "move fast and break things" mentality is what's causing problems. I want to add nuance. I don't think it is the "move fast and break things" mentality that creates all the shit, but that that there is no "time to clean up, everybody do your share" mentality to complement it. Doing things often creates a mess, and doing hard things often creates a bigger mess. You can't make a fancy meal without dirtying a bunch of dishes. But are we seriously not hiring "dishwashers"? Creating a mess is unavoidable, and certainly it shouldn't be too large of a mess, but the dishes are piling up and we can't hide it anymore. It isn't a linear problem because the mess compounds and the mess itself generates more mess. We won't refactor. We won't rewrite. So we just have patchwork on top of patchwork. That's enshitification. We also have a status quo that we sprint into a sprint and try to move as fast as we can but only measure how fast we're going by looking at how far we moved in a quarter. There's no long term measurement because "that's too hard." But this is like trying to circumnavigate the world and choosing to walk. You'll make progress every day and more importantly, __measurable__ (and easily measurable) progress. But you could spend 11 months building a fucking Cesna and still beat the person that could walk on water. You need to move fast (like the Cesna), but to move fast requires also slowing down. Who is willing to slow down? I think most people don't have a problem with people using public data to do research or similar activities. That people wouldn't be up in arms if OpenAI scraped all of YouTube, trained on it, proved internally that they could do cool stuff with it, AND THEN either started to generate their own data for training or started to purchase data. Even though this would still be costly to YouTube and be a weird ethical ground (like getting a "free trial" (or theft) of the data and pay only if it works). As someone with anxiety, I can assure you, it is not a good idea to constantly be rushing around chasing everything that needs to be solved. You just make more messes because you sloppily "fix" the issues, trying to move onto the next. The trick is to fight your own mind, slow down, triage, and solve anything that isn't a literal fucking fire with calm and care, no matter how much your own mind wants to convince you it is an emergency. But when everything is an emergency, nothing is. And that's the problem. We created an economy based on a business strategy that is functionally equivalent to an untreated and severe anxiety disorder.
- deleted 2y ago[deleted]
- JohnFen 2y ago> We're already seeing the practical results: People stop/reduce posting their works to the clearnet. It's really interesting. Ever since I started publicly stating that I've removed all my works from the public web because there is no realistic method of defending myself against AI scrapers, I've been encountering a surprising number of people who say they've done the same thing. I don't know how far this will go, but at least I know I'm not alone.