3 ms·
And from the following day, their paraphrasing of https://www.theverge.com/2023/11/4/23946353/generative-ai-copyright-training-data-openai-microsoft-google-meta
by ineptech 3y ago
And from the following day, their paraphrasing of https://www.theverge.com/2023/11/4/23946353/generative-ai-copyright-training-data-openai-microsoft-google-meta-stabilityai https://www.theverge.com/2023/11/4/23946353/generative-ai-co... is succinct:
> Meta: We stole so much stuff that, actually, we didn't steal any stuff
> Google: If things were different, things would be totally different
> Andreeson Horowitz: But we already spent so much money
> Microsoft: Think about how this would hurt the little guys, like us
> Anthropic: Fucking shut up about it
> Hugging Face: It's super legal but honestly it might not be idk
> StabilityAI: It's legal in places that aren't here
- ineptech 3y agoThis is a little oblique, but isn't there a parallel to the music piracy debates of the 90s here? The defense of those who steal training data, "It's not stealing because I'm only taking a little tiny piece from a whole bunch of people", seems like a neat inversion of the defense offered by people torrenting mp3s in their dorm rooms. It's worth noting that, back when it suddenly became technically feasible to copy all songs for free, the response from the rightsholders of music was, essentially, to demand that we all collectively pretend it isn't. It might make sense for the rightsholders of training data to try to do the same, if they could speak with one voice and if they had lobbyists, but I guess the fact that they can't and don't settles that.
- petercooper 3y agoI like a good metaphor, so I see your point, though it doesn't quite fit for me. To me, training a model is less like redistribution and more like reading something and having it influence your thoughts. You may be able to reproduce certain memorable sections of things you've read verbatim (common with poetry, for example) and you may well be able to reproduce a 'style', but you don't have the entirety of the source material memorized in a way that it could be redistributed in full (but if you did, that may be well an infringement at such a time).
- ineptech 3y agoI agree that copying music is not all that similar to an AI model using training data, but they're both examples of a superficially similar activity being transformed by automation and scale. The difference between manually copying a song on cassette for a friend and uploading that song to napster resembles the difference between a human learning to draw in an established artist's style to an AI model learning to draw in an established artist's style. In both cases, the offender's defense is a variation on, "It was okay when humans did this slowly..."