4 ms·
The problem is mostly about model training and architecture. I was doing TTS like 2.5/3 years ago and most models were train on fixed (+/- 5-10s) clips with lik
by machinekob 4y ago
The problem is mostly about model training and architecture. I was doing TTS like 2.5/3 years ago and most models were train on fixed (+/- 5-10s) clips with like avg of 80 words or so, there were few attempts for fixing that and if I remember correctly few RNN-based models were good at ignoring input length and generate "good" audio but new flow-based and diffusion based models are out of my domain as I'm in CV for past few years and only read some new cool paper once in a while :)
You can also search for postags (and token ids for them) that are especially placed for "pause" audio as they often fix problem with weird transition when you split the sentences.
This repo -> https://github.com/TensorSpeech/TensorFlowTTS https://github.com/TensorSpeech/TensorFlowTTS was very good few years back.