3 ms·
The input poses appear to be generated from OpenPose [0], which uses regular images as input. With the creation of stable diffusion video, you could theoretical
by netruk44 3y ago
The input poses appear to be generated from OpenPose [0], which uses regular images as input. With the creation of stable diffusion video, you could theoretically prompt it to make a video of what you wanted, then run it through OpenPose.
But I think the more realistic approach is to take a video of yourself doing the motions you want the AI to generate, and run that through OpenPose instead.
Using just words leaves a lot to the model’s interpretation. I feel like you might wind up spending a lot of time manually fixing little things, similar to how you might infill the “wrong” parts of an AI generated image. It might be easier to just take a 15 second video to get the exact skeleton animation you want.
[0] https://github.com/CMU-Perceptual-Computing-Lab/openpose https://github.com/CMU-Perceptual-Computing-Lab/openpose