4 ms·
Loosely related, but what's the state of the art in natural text to speech ML/AI models?
by rixrax 4y ago
Loosely related, but what's the state of the art in natural text to speech ML/AI models?
- forgingahead 4y agoTortoise-TTS (stylised as "TorToiSe") is pretty amazing: https://github.com/neonbjb/tortoise-tts https://github.com/neonbjb/tortoise-tts
- snakers41 4y agoSilero TTS works fast even on one CPU thread, this is the point
- fxtentacle 4y agotext to WAV: You predict mel spectrums with a transformer architecture (so word embedding + attention decoder) and then convert them into audio signals with Parallel WaveGAN or Hifi-GAN. FastSpeech2 (Microsoft) is extremely good. WAV to text: You detect wave shapes with convolutions to generate an embedding, then attention layers to turn it into an encoding, then convert that to logits. Logits go into language model and that predicts the final sentence with a beam search decoder. wav2vec 2.0 (Facebook) is amazing.