4 ms·
https://speech.botiumbox.com https://speech.botiumbox.com just a small server, hope that it wont crash when posting the link here
by ftreml 7y ago
https://speech.botiumbox.com https://speech.botiumbox.com
just a small server, hope that it wont crash when posting the link here
- bArray 7y ago> just a small server, hope that it wont crash when posting > the link here In all honesty I was just looking for some pre-processed examples. The API itself seems intuitive enough, good job. What really stands out to me is that open-source text-to-speech is really awful compared to commercial solutions - which surprises me. Does anybody know why this is the case? (And not just "money", I'm talking technology)
- z3t4 7y agoI tested the text-to-speech and it was acceptable, on pair with the robotic voice many screen readers use. For me it's not that important that it sounds like a real human, it's more important that it's accurate, and that you can hear what it say. Also pre-processed examples is not really any useful, as the author will likely pick examples that turned out good. This API testing page was very useful though! As it was very easy for me to test myself.
- ftreml 7y agowhen using right ssml formatting, output quality is not so bad. but it cannot compete commercial engines, thats right. i guess thats for the same reason that google is dominating the speech recogniction world: they have tons of training data available. not smarter algorithms, just more data.
- z3t4 7y agoCool! Thanks! I've tried both wav2letter and DeepSpeech, and now this, but I get very poor results even with short sentences (compared to Google's proprietary services). Would it be possible to also make an API for passing in training data and automatically update the model? I'm thinking that the results might get better if they are trained with the specific audio/hardware/settings and dialect of the end user.
- magicalhippo 7y agoI just played with DeepSpeech (v0.6.1) and I found significant improvements by using a custom language model. The language model is built from sentences and is rather quick to build (~seconds), at least for the small number of sentences I used. This can then be combined with the pre-trained neural net. Though I hear DeepSpeech is currently fairly US-influenced when it comes to recognizing accents. So if you're not a native speaker, consider contributing to the open-source dataset over at https://voice.mozilla.org/ https://voice.mozilla.org/
- ftreml 7y agoits always an issue with domain specific utterances. for the freely available data they are oov and have to be handled somehow (fe generate pronounciation with seqitur)
- ftreml 7y agoit requires hundreds of hours of speech data to make any difference, at least for a context-free asr. and gpu-powered high-end hardware, and several days for training. training an asr model is totally different requirement than using it as a client