4 ms·
Text to speech.... several years ago I was playing with a pet project to make procedurally generated number station broadcasts. For that I needed text 2 speech
by audiometry 6y ago
Text to speech.... several years ago I was playing with a pet project to make procedurally generated number station broadcasts. For that I needed text 2 speech with ?euphony? The available tools were horrid. There was some library MaryTTS (?) that had a poor python wrapper to talk to it. All the web based systems were super stingy in their demo resources and their use cases seemed geared more for TTS phrases that you’d have on a phone menu or something, not a paragraph of prose. I used Watson’s tool for a while before they became very annoying and I abandoned it.
Is there any better TTS Tool people could recommend?
- deleted 6y ago[deleted]
- shakna 6y agoAmazon Polly [0] and Google's Neural Voices [1] make Watson's voice seem pretty awful by comparison. Unfortunately, they can be costly, depending on what you're doing, and you have to hope neither of those companies cut off your access for arbitrary reasons. If/when that happens, and you go to look at the rest of the tools... They're pretty crap. Still. [0] https://aws.amazon.com/polly/ https://aws.amazon.com/polly/ [1] https://cloud.google.com/text-to-speech/ https://cloud.google.com/text-to-speech/
- nmfisher 6y agoI use Azure neural TTS for (Chinese) voice synthesis in my app, and honestly it's amazing. I did experiment with Google's API at the beginning, the audio would contain recurring artefacts that made it painfully obvious it wasn't real. If I had to describe it, I'd say "50ms of the sound your computer makes when it hangs while playing an audio file". No such problems with Azure. Zero artefacts and nice, crisp, natural enunciation, though unfortunately it is still quite unreliable when synthesizing single words (as compared with full sentences). You can actually test it out by using the "Read Aloud" feature inside Edge. I don't know how well this translates to English, though. I definitely enjoy the benefit of working with a gender/language combination (female/Chinese) that's received the largest total investment of effort and resources. Many well-resourced Chinese companies (banks, tech services, gaming, etc) all use similar synthesized female voices in their phone/online/B&M services.
- criddell 6y agoIf your software works on a Mac, maybe the built-in TTS functionality would work.
- audiometry 6y agoThe term I meant to use above was 'prosody' not euphony.
- kalico 6y agoI've been working on a TTS API connector and SSML editor. Voxabot.com Take note that unfortunately the Neutral TTS for Polly isn't fully exposed. That’s arguably the best sounding TTS.
- tpoacher 6y agoI can't access the demo. It says "Sign in with Google".
- billfruit 6y agoWhat does emacspeak use for example?
- nshm 6y agoFor open source offline TTS with more or less recent algorithms you can check https://github.com/TensorSpeech/TensorFlowTTS https://github.com/TensorSpeech/TensorFlowTTS