3 ms·Zero-shot lip-to-speech synthesis with face image based voice control1 points by programd 3y ago