3 ms·
> The issue is a completely different matter if you made a copy on a harddrive. But there are no copies. For example, the LAION-2b training data is a total of
by usrbinbash 4y ago
> The issue is a completely different matter if you made a copy on a harddrive.
But there are no copies. For example, the LAION-2b training data is a total of 240 TB. The pruned SD model based on this dataset, is less than 5GB.
The data isn't copied into the models, it is used to teach the models, letting them learn patterns in the dataset.
- freejazz 4y agoYou are just anthropomorphizing the model by calling it teaching and then implicitly equating that it's the same thing happening in the human mind. That's your burden to establish when you say it's the same.
- unusualmonkey 4y agoActually, isn't the burden of the plaintiffs to prove their copyrights are being violated? Whether or not it's identical to human brains isn't the matter, they'd need to prove how a small 5GB model trained from a huge dataset infringes their rights specifically.
- freejazz 4y agoWell if it's their defense they need to substantiate it. Obviously the plaintiff has a theory, they filed the case.
- usrbinbash 4y agoI am not anthropomorphizing anything, because this is literally what happens. The model is taught, by having its predictions tested against examples, how images work.
- freejazz 4y agoYou are, that's not what teaching means and it's not what learning means.
- usrbinbash 4y agoOkay, then what do these 2 terms mean?
- crote 4y agoOkay, but can you prove that? I have a 1.9 GB mp4 file on my harddrive. It contains 2 hours and 15 minutes of 1080p video data at 24 fps. Assuming it was generated from 4096x2160 16-bit color depth source material, the "training data" was 10.32 TB. I bet I could even get a similar size reduction as LAION-2b if I recompressed it to 720p. Could I not also claim that I created an advanced AI model, which did not copy but learned patterns in the dataset? Modern video compression algorithms are getting quite complicated, after all. I think no reasonable person would agree with this, but can you prove that the AI model is doing something substantially different?
- usrbinbash 4y ago> Could I not also claim that I created an advanced AI model, which did not copy but learned patterns in the dataset? Such patterns would enable the video file to decode into a multitude of pictures not originally in the training data. Obviously, a video file cannot do that...it's just compressed data. Generative models however can generate things that are not in its training set. And of course, there is a fundamental difference in the source data between compressed video and a generative model: video codecs work with a sorted sequence of images, where most images are slight variations of the ones before them. The training for generative AI doesn't have these properties, the input is not an ordered sequence, and even similar pictures are not sequential variations of one another.
- anankaie 4y agoTo expand a bit on this: Relying solely "uncompressed" size does not a really good metric make (this is analogous to the raw input size of the LAION dataset): one could make a reasonable argument that there are not billions (1) of image-pairs that are effectively identical up to a minute shift. I would posit the correct basis would be the Shannon entropy of the "best fit" ordering (minimize inter-frame diff), versus the lossy-compressed video, and a similar "best fit" ordering for the LAION dataset vs. the model. My suspicion is that one will find that the relative number of "smooth transition" pairs in LAION viz the whole will be very different from the video. ------- (1) - Napkin math: There are about 194400 frames, so ~37 billion (37791165600) frame-pairs. Assuming you have runs of about 1 second between hard cuts throughout, so an incidence rate of 1/24 for non-smooth transitions, gives us about ~36 billion "smooth transition" frame-pairs. I think it is safe to assume "on the order of" 1 billion, then. This ignores long "action" scenes with significant variance in images throughout, but also ignores longer-than-1-second slower scenes, hence the order-of-magnitude shrink in the assumption as buffer.