3 ms·
The "initial frame -> video" problem seems way harder than generating video from a text input. Once a good dataset for that problem is assembled, it seems like
by pc2g4d 4y ago
The "initial frame -> video" problem seems way harder than generating video from a text input. Once a good dataset for that problem is assembled, it seems like Stable Diffusion would naturally accept an additional "time" dimension, and generate cohesive output, with corresponding massively increased hardware requirements.
Though I'd venture that the first "novel to feature film" or "novel to TV series" algorithm won't just be an upscaling of this tech....