4 ms·
Video will suffer even more than still imagery or text from the inherent lack of continuity/self-consistent memory that these autocomplete/prediction algorithms
by throwaway920102 4y ago
Video will suffer even more than still imagery or text from the inherent lack of continuity/self-consistent memory that these autocomplete/prediction algorithms have.
So cool for abstract art but not for storytelling or following a script. Unless you are OK with the content being visually inconsistent like an acid trip.
- FairlyInvolved 4y agoI think that's just a scaling issue, fundamentally there's no reason why a model trained on video couldn't come to create coherent motion in the same way that image models can now product coherent lighting/themes. Smaller image models had the same problems with logical inconsistency just because they didn't have sufficient general understanding of how visual concepts. The same is almost certainly true of video - early smaller models will likely create janky movements/motion, however once they've seen enough video to understand how a person walks, how a scene is framed etc.. there's no reason we couldn't get to the same level of maturity as today's image models. I think the real issue will come from labelling - most video is only going to be labelled simply with basic info/captions without detailed descriptions of the camera pan, movement of subjects. The amount of text required to accurately describe a scene is much larger than a still image and I'm not sure how once would go about collecting this.
- babyshake 4y agoIt does seem to be the case that this type of generative image AI makes somewhat surreal imagery unless given very specific literal directions (aka, Tom Cruise hugging Ben Affleck). If using this AI right now it is probably best to work within these constraints.