5 ms·
As I've learned over time (and other people in these comments have clarified) it turns out that evaluating "quality" of Text To Speech is somewhat dependent on
by follower 2y ago
As I've learned over time (and other people in these comments have clarified) it turns out that evaluating "quality" of Text To Speech is somewhat dependent on the domain in which the audio output is being used (obviously with overlaps), broadly:
* accessibility
* non-accessibility (e.g. voice interfaces; narration; voice over)
The qualities of the generated speech which are favoured may differ significantly between the two domains, e.g. AIUI non-accessibility focused TTS often prioritises "realism" & "naturalness" while more accessibility focussed TTS often prioritizes clarity at high words-per-minute speech rates (which often sounds distinctly non-"realistic").
And, AIUI espeak-ng has historically been more focused on the accessibility domain.
- vlovich123 2y agoI don't have any disabilities so I don't know if espeak-ng is better on the pure accessibility axis. But given that MacOS tends to be received quite well by the accessibility crowd & it's definitely a focus from what I observed internally, given that MacOS has much higher realism & naturalness out of the box, I'm going to posit that it's not the linear tradeoff argument you've made & that espeak-ng defaults aren't tuned well out of the box.