3 ms·
I thought I'd just add another fun fact/data point here. This is obviously my personal opinion. I have to use TTS to use my computer with a screen reader, and f
by ClawsOnPaws 5y ago
I thought I'd just add another fun fact/data point here. This is obviously my personal opinion. I have to use TTS to use my computer with a screen reader, and for that, I mostly prefer more synthetic speech. When I read long form text like books, articles, etc. I do prefer more natural voices, but for doing actual work like reading code or simply using user interfaces, I like the predictability of more synthetic/algorithmic speech. Apple added the neural Siri voices to the new VoiceOver. They sound incredible but the quality of the voice also brings latency with it. Something like ESpeak is much, much more performant and predictable, and it speeds up much better. I use my TTS at a very fast rate and I find that the more natural a voice, the harder it is to understand at that speech rate. Neural voices speak the same phrase of text differently every time it's uttered. Slightly different intonation, slightly different speech rhythm. This makes it hard to listen out for patterns. So for me there's definitely still a place for synthetic speech.
- mwcampbell 5y agoIn fact (as I'm sure you know), one of the most beloved speech synthesizers among English-speaking blind users is a closed-source product called ETI-Eloquence that has been basically dead for nearly 20 years. (It was ported to Android several years ago, but that port was discontinued because they couldn't update it for 64-bit.) No recent speech synthesizer has quite matched its consistent intelligibility, particularly at high speeds. espeak-ng comes close, but it has a bad reputation (mostly, I think, leftover from earlier versions of espeak that really weren't very good). Edit: Sample of ETI-Eloquence at my preferred speed: https://mwcampbell.us/audio/eloquence-sample-2021-09-25.mp3 https://mwcampbell.us/audio/eloquence-sample-2021-09-25.mp3 (yes, it mispronounces "espeak") Edit 2: To elaborate on what I mean by "mostly dead": In 2009 I was tasked with adding support for ETI-Eloquence to a Windows screen reader I developed. At that time, Nuance was still selling Eloquence to companies like the one I worked for back then. When I got the SDK, the timestamps on the files, particularly the main DLLs, were from 2002. As far as I know, an updated SDK for Windows was never released. I'm thankful for Windows's legendary emphasis on backward compatibility, particularly compared to Apple platforms and even Android. Finally, a sample of espeak-ng (in the NVDA screen reader) at my preferred speed: https://mwcampbell.us/audio/espeak-ng-sample-2021-09-25.mp3 https://mwcampbell.us/audio/espeak-ng-sample-2021-09-25.mp3 I use the default British pronunciation even though I'm American, because the American pronunciation is noticeably off.
- ClawsOnPaws 5y ago> In fact (as I'm sure you know), one of the most beloved speech synthesizers among English-speaking blind users is a closed-source product called ETI-Eloquence that has been basically dead for nearly 20 years. This is exactly the speech synthesizer I use daily. I've gotten so used to it over the years that switching away from it is painful. On Apple platforms, though, using it is not an option. So I use Karen. Used to use Alex, but Karen appears to be slightly more responsive and tries to do less human stuff when reading. Responsiveness is a very important factor, actually. Probably more so than people might realize. Eloquence and ESpeak react pretty much instantly whereas other voices might take 100 MS or so. This is a very big deal for me. Just like how one would like instant visual feedback on their screen, it's the same for me with speech. The less latency, the better. My problem with ESpeak is that it sounds very rough and metallic whereas Eloquence has a much warmer sound to it. I pitch mine down slightly to get an even warmer sound. Being pleasant on the ears is super important if you listen to the thing many, many hours a day.
- mwcampbell 5y agoI agree with you that Eloquence sounds warmer than eSpeak. I wish there was an open-source speech synthesizer comparable to Eloquence or even DECtalk. That approach to speech synthesis is old enough now that I'm sure there are published algorithms whose patents have expired. The problem, of course, would be funding the work on a good open-source implementation.
- app4soft 5y agoWhat about RHVoice?[0,1] [0] https://github.com/RHVoice/RHVoice https://github.com/RHVoice/RHVoice [1] https://rhvoice.org/en-voices/ https://rhvoice.org/en-voices/ [2] https://f-droid.org/en/packages/com.github.olga_yakovleva.rhvoice.android/ https://f-droid.org/en/packages/com.github.olga_yakovleva.rh...
- machawinka 5y agoThat is a bit too fast for me. I am surprised your brain can process at such high speed. I guess it is a matter of practice.
- chrismorgan 5y agoI’ve heard exactly this from a couple of blind people I’ve interacted with too.
- machawinka 5y agoGreat insight, it is monotonous but it really helps to be in the flow for a few hours of productive work.