2 ms·
Cool to see the Coqui team dive into this too. I made a video using Stable Diffusion, Coqui TTS with a VITS model, Frame Interpolation for Large Motion, and Wa
by nanonomad 4y ago
Cool to see the Coqui team dive into this too. I made a video using Stable Diffusion, Coqui TTS with a VITS model, Frame Interpolation for Large Motion, and Wav2Lip a couple weeks ago:
https://youtu.be/BnrnJ_6wiYs https://youtu.be/BnrnJ_6wiYs
Theres links to the walkthroughs if anyone wants to set up the open source components and try it out.
My voice model wasnt great for this, so it shouldnt be considered ideal performance by any means.
Heres Bill Gates rapping: https://youtu.be/1ztm7aBssgA https://youtu.be/1ztm7aBssgA