3 ms·
Are these text-to-speech models manually controllable in terms of prosody, etc., or is it all transformer-based text-to-audio? I've followed some of the resear
by ipsin 4y ago
Are these text-to-speech models manually controllable in terms of prosody, etc., or is it all transformer-based text-to-audio?
I've followed some of the research on prosody transfer, etc., but it still seems bad in the TTS systems I've heard.