3 ms·
> Could I not also claim that I created an advanced AI model, which did not copy but learned patterns in the dataset? Such patterns would enable the video file
by usrbinbash 4y ago
> Could I not also claim that I created an advanced AI model, which did not copy but learned patterns in the dataset?
Such patterns would enable the video file to decode into a multitude of pictures not originally in the training data. Obviously, a video file cannot do that...it's just compressed data.
Generative models however can generate things that are not in its training set.
And of course, there is a fundamental difference in the source data between compressed video and a generative model: video codecs work with a sorted sequence of images, where most images are slight variations of the ones before them. The training for generative AI doesn't have these properties, the input is not an ordered sequence, and even similar pictures are not sequential variations of one another.
- anankaie 4y agoTo expand a bit on this: Relying solely "uncompressed" size does not a really good metric make (this is analogous to the raw input size of the LAION dataset): one could make a reasonable argument that there are not billions (1) of image-pairs that are effectively identical up to a minute shift. I would posit the correct basis would be the Shannon entropy of the "best fit" ordering (minimize inter-frame diff), versus the lossy-compressed video, and a similar "best fit" ordering for the LAION dataset vs. the model. My suspicion is that one will find that the relative number of "smooth transition" pairs in LAION viz the whole will be very different from the video. ------- (1) - Napkin math: There are about 194400 frames, so ~37 billion (37791165600) frame-pairs. Assuming you have runs of about 1 second between hard cuts throughout, so an incidence rate of 1/24 for non-smooth transitions, gives us about ~36 billion "smooth transition" frame-pairs. I think it is safe to assume "on the order of" 1 billion, then. This ignores long "action" scenes with significant variance in images throughout, but also ignores longer-than-1-second slower scenes, hence the order-of-magnitude shrink in the assumption as buffer.