35 ms·
This paper focuses on statistical parametric speech synthesis (SPSS). SPSS is only 1/2 of the text to speech problem. SPSS is the problem of going from linguis
by nicklo 11y ago
This paper focuses on statistical parametric speech synthesis (SPSS). SPSS is only 1/2 of the text to speech problem.
SPSS is the problem of going from linguistic features, phenomes, etc, to speech audio. These features are more or less golden, either derived from the audio itself or hand-labeled. So things like tonality, cadence, emphasis on words is already encoded as features which is why these samples sound so good.
Deriving these features from pure text is very hard, and this failing is the main reason most text to speech systems sound so dull and tone-dead.
That being said, these results are seriously impressive, sounding very natural. Would love to see someone try and train an end-to-end system from pure text to speech. I think we'd see some big improvements like what Baidu has done for end-to-end speech to text.