7 ms·
Show HN: Text-to-speech and speech-to-text open-source software stack
- ftreml 7y agoThis project is the result of a one year long learning process in speech recognition and speech synthesis. The original task was to automate the testing of a voice-enabled IVR system. While we started with real audio recordings, very soon it was clear that this approach is not feasible for a non-trivial app and it will be impossible to reach a satisfying test coverage. On the other hand, we had to find a way to transcribe the voice app response to text for doing our automated assertions. As cloud-based solutions where not an option (company policy), we very quickly got frustrated as there was no "get shit done" Open Source stack available for doing medium-quality text-to-speech and speech-to-text conversions. We learned how to train and use Kaldi, which is according to some benchmarks the best available system out there, but mainly targeting academic users and research. We made heavy-weight MaryTTS work to synthesize speech in reasonable quality. And finally, we packaged all of this in a DevOps-friendly HTTP/JSON API with a Swagger definition. As always, feedback and contributions are welcome!
- dillonmckay 7y agoWhat was the IVR software you were using, was that bespoke?
- ftreml 7y agono, it was a special call center software called vacapo.
- yorwba 7y agoYou should indicate somewhere which languages are supported. You mention Deepspeech German in the credits, but did you use it to implement support for German or just to inform your architectural choices?
- ftreml 7y agogood point, will add this information. in short, german and english, because those are the languages i am comfortable with. i hope to find native speakers contributing more languages. deepspeech was evaluated, but right now kaldi provides better performance, thats why we stick to kaldi by default.
- magicalhippo 7y ago> deepspeech was evaluated, but right now kaldi provides better performance Was this before the streaming API was added to DeepSpeech? I recently did some testing with it, and it provides text within ~100ms of last audio block on my PC. edit: that is, the most significant latency I had was from having to wait a bit to detect end of speech.
- ftreml 7y agowith "performance" i meant the error rate, not the speed. speed was not a criteria for us, so we didnt evaluate it. i read that deepspeech is way quicker in training phase as it is smarter in using gpu computing power.
- magicalhippo 7y agoAhh gotcha. Yes I got rather poor "hit rate" until I made my own language model. Fortunately for me I only needed it for command recognition, so the process was quite quick, and results were very good. No need to retrain the net.
- ftreml 7y agowe used a context-free data set, augmented with additional domain-specific sound samples and it worked out fine, although the additiinal samples made nearly no difference
- StudentStuff 7y agoMozilla DeepSpeech has seen significant improvements in the quality of results it provides (and RAM/CPU usage) since early December, the accuracy of transcriptions has substantially improved in our testing.
- ayumu722 7y agoI've built a speech synthesis system with Marytts before, it works, but unfortunately, the quality is not very good, HMM and unit selection speech synthesis are now very old approaches, they are far from the current state of the art, you should try open-source implementations of Tacotron2 or wavnet, you will surely achieve better quality.
- deleted 7y ago[deleted]
- monkpit 7y agoWhat exactly does “low-key” mean in this context?
- ftreml 7y agoit means: easy to install, easy to use, medium performance. no further know-how needed. compared to the total effort for selecting, training, deploying speech recogniction and speech synthesis it provides an extremly quick boilerplate to add voice to your pipeline. i wish there was something like that when we started the project.
- pouta 7y agoI built something quite similar on my own product. Is there any interest on adding more STT/TTS backends to the software? Think services like Lyrebird or Trint. I could contribute towards it since I have done it before. Thank you for building this!
- ftreml 7y agoyes absolutly! it should be a good mix of freely available packages, with meaningful default configuration.
- grizzles 7y agoFacebook has released wav2letter++. I'd wager that will outperform kaldi by a wide margin.
- ftreml 7y agoin our tests the performance was comparable, so we had no reason to switch from kaldi to something else (german data only). after all, what matters the most is the available training data. are there ready trained models available for wav2letter ?
- grizzles 7y agoOops, there I go being Anglocentric again. The datasets are out there; but idk if there are German trained models. Fortunately wav2letter++ looks like it's been built to be very friendly to train a more extensive one if there isn't. My understanding is they are still releasing parts of it.
- dsteinman 7y agoAny idea how it compares to Mozilla's DeepSpeech?
- ftreml 7y agofrom my experience, when trained with the same data, kaldi is slightly better and with custom recipes adaptable to changing conditions. deepspeech has way better documentation and is more developer friendly. wav2letter seems to be the quickest. i guess there is no real winner here ... depends what criteria are applied. just an example, kaldi is a weird mixture of c++, python2, python3, shell scripts, java, perl. hard to oversee. deepspeech is python. wav2letter is an exe file.
- magicalhippo 7y ago> deepspeech is python I was under the impression DeepSpeech was native (C++), with bindings for Python and others. Personally I've used it with Node.js so far, and I couldn't see any dependencies on Python. edit: I was talking client, you're talking training I guess.
- z3t4 7y agoWould be cool with a web demo.
- ftreml 7y agohttps://speech.botiumbox.com https://speech.botiumbox.com just a small server, hope that it wont crash when posting the link here
- bArray 7y ago> just a small server, hope that it wont crash when posting > the link here In all honesty I was just looking for some pre-processed examples. The API itself seems intuitive enough, good job. What really stands out to me is that open-source text-to-speech is really awful compared to commercial solutions - which surprises me. Does anybody know why this is the case? (And not just "money", I'm talking technology)
- z3t4 7y agoI tested the text-to-speech and it was acceptable, on pair with the robotic voice many screen readers use. For me it's not that important that it sounds like a real human, it's more important that it's accurate, and that you can hear what it say. Also pre-processed examples is not really any useful, as the author will likely pick examples that turned out good. This API testing page was very useful though! As it was very easy for me to test myself.
- ftreml 7y agowhen using right ssml formatting, output quality is not so bad. but it cannot compete commercial engines, thats right. i guess thats for the same reason that google is dominating the speech recogniction world: they have tons of training data available. not smarter algorithms, just more data.
- z3t4 7y agoCool! Thanks! I've tried both wav2letter and DeepSpeech, and now this, but I get very poor results even with short sentences (compared to Google's proprietary services). Would it be possible to also make an API for passing in training data and automatically update the model? I'm thinking that the results might get better if they are trained with the specific audio/hardware/settings and dialect of the end user.
- tianshuo 7y agoIs this using google's tacotron2 or wavenet anywhere? How does this compare to them?
- polishdude20 7y agoSo why is 40gigs of free space needed?
- ftreml 7y agospeech recogniztion model filed are quite big
- polishdude20 7y agoBut 40gb? I feel like that includes training data or something. A model can't just be 40 GB or else all of the audio would have to be passed through all 40gb of the model during inference. That seems huge.
- ftreml 7y ago40GB is maybe too much, but when building the docker images there is some space wasted. The image size after building is around 20GB (12GB marytts, 3GB kaldi de, 6GB kaldi en)
- hajimemash 7y agoHere's a sample wav output from using their swagger endpoint: https://drive.google.com/file/d/15y83NSXOCrEW9v9eQVCy6oHcWJ8DXGE0/view?usp=sharing https://drive.google.com/file/d/15y83NSXOCrEW9v9eQVCy6oHcWJ8... Why does the voice/pronunciation have such drastic volume spikes and dips?
- ftreml 7y agomarytts supports a high number of voices in several languages. you could try to use another voice.
- DonHopkins 7y ago"Now let's have a little taste of that old computer generated swagger." -Max Headroom https://www.youtube.com/watch?v=WTN1WsUCyQc&t=3m26s https://www.youtube.com/watch?v=WTN1WsUCyQc&t=3m26s
- tomcam 7y agoThat's fantastic work and the demo is very well done. Thanks for sharing it. You obviously put a lot of hard work into it. Feels super polished.
- hardwaresofton 7y agoCan anyone in the space expand on why it's increasingly rare to see people using/building on Sphinx[0]? Do people avoid it simply because of an impression that it won't be good enough compared to deep learning driven approaches? [0]: https://cmusphinx.github.io/ https://cmusphinx.github.io/
- ftreml 7y agofor me it was exactly that, yes. we had powerful hardware and budget available so there was no reason to stick to statistical models. the available benchmarks showed that kaldi easily outperforms cmusphinx - when starting from scratch you typically select the leader i would say
- shakna 7y agoI've avoided Sphinx after trying to use it, because: 1. Compiling it is hit and miss. Sometimes it works, sometimes it doesn't. There is no official package in any Linux distribution, so packaging anything with it is incredibly painful. There's no easy way to cross-compile your project, so you'll end up working around the build process. 2. The documentation is woeful. > Recent CMUSphinx code has noise cancellation featur. In sphinxbase/pocketsphinx/sphinxtrain it’s ‘remove_noise’ option. In sphinx4 it’s Denoise frontend component. So if you are using latest version you should be robust to noise in some degree already. Most modern models are trained with noise cancellation already, if you have your own model you need to retrain it. > The algorithm impelmented is spectral subtraction in on mel filterbank. There are more advanced algorithms for sure, if needed you can extend the current implementation. Inconsistent methods across the codebases prevents you knowing where to look, and if it is documented, it may involve spelling errors which you have to guess around (like above). Also plenty of vague references to other documents that may or may not even exist.
- sgt101 7y agoI think that sphinx is about as good as it can get. Without really significant technical change there is no prospect of the performance that would allow successful application to the kind of applications that people want to try.
- sandreas 7y agoCould you explain, what's the difference to - https://github.com/gooofy/zamia-speech#asr-models https://github.com/gooofy/zamia-speech#asr-models - https://github.com/mpuels/docker-py-kaldi-asr-and-model https://github.com/mpuels/docker-py-kaldi-asr-and-model in regards of speech recognition except the fact that its easier to use?
- ftreml 7y agozamia-speech: asr training scripts for research purposes, several ready trained asr models for download, based on voxforge data. zamia-speech is the (very hard in terms of know-how, hardware and software requirements) training part to be done where projects like botium speech processing are be based upon. the other one is an example for packaging kaldi in a docker container.
- sandreas 7y agoThank you :-) I tried to get in touch with https://www.vorleser.net/ https://www.vorleser.net/ in the past to provide Speech training data, but they were not really interested. I thought it would be great to improve STT having a real good and HUGE set of german audiobooks based on Text, that is publicly available... unfortunately i had no success trying to script something for this purpose (mainly lack of time).
- ftreml 7y agoI point you to this article: https://medium.com/@klintcho/creating-an-open-speech-recognition-dataset-for-almost-any-language-c532fb2bc0cf https://medium.com/@klintcho/creating-an-open-speech-recogni... It basically describes the thing you mentioned - matching freely available audio books with the source text and using some tools to preprocess the data suitable for ASR training (alignment, splitting).
- ajaviaad 7y agoWhich languages are supported?
- ftreml 7y agocurrently included german and english. contributions for other languages welcome, native speakers will have better insights into quality of speech output and recognition model
- briga 7y agoIs MaryTTS still as good as it gets for free TTS? I've been researching this topic and it seems like there are some open-source implementations of Tacotron, but the quality isn't necessarily great.
- ftreml 7y agowith ssml formatting it is good enough for a customer facing ivr, though clearly recognizable as robotic voice - if this should be an issue i would not recommend it
- lunixbochs 7y agoThe nvidia tacotron implementation was much better out of the box than the wav in the neighbor thread.
- CommanderData 7y agoAny recommendations for a real time solution? I maintain a platform which features live video events we'd like to add captioning and so far can only see IBM Watson providing a websockets interface for near real time stt.
- ftreml 7y agothe project includes a websocket endpoint for realtime decoding. will add it to the docs. we are already using it for a callcenter with around 50 parallel audio streams.
- ftreml 7y agoforgot to mention: when doing realtime parallel processing, the default configuration of this project is not a feasible setup. you have to run way more decoder workers, maybe distributed on various machines.
- bobmaxup 7y agoIf marytts is so good, why are we in many linux distros still using https://en.wikipedia.org/wiki/Festival_Speech_Synthesis_System https://en.wikipedia.org/wiki/Festival_Speech_Synthesis_Syst... as our default tts system?
- ftreml 7y agojust a guess: marytts is rather heavy weight
- monkeydust 7y agoAre there any performance metrics of this versus other offline and cloud based services?
- ftreml 7y agoyes there are plenty of them. just google for something like "kaldi vs google". in short: not surprising the blockbuster cloud services provide better results as they have way more training data. tradeoff between price, privacy, quality.
- mariushn 7y agoWould love to see a live demo. MaryTTS demo link is broken.
- ftreml 7y agosee here: https://speech.botiumbox.com https://speech.botiumbox.com
- mariushn 7y agoThanks. Hard to use and poor result :( Good example: https://cloud.google.com/text-to-speech/ https://cloud.google.com/text-to-speech/
- ftreml 7y agoIt's an API, of course it is hard to use without any real user interface ... but as an STT/TTS API, it won't get more easy than that ... Of course it is not a competitor to Google in any sense.
- klft 7y agoThanks for providing the test service. It's really easy to use if one is familiar with swagger UIs. First tests show that the results are good.
- dmos62 7y agoI'd like to have my laptop read out epubs or articles. Recommendations for speech synthesis (TTS) on the command line?
- ftreml 7y agopicotts is a command line tool. marytts is a client/server tool. (both included in botium speech processing and callable with curl). high quality with google cloud speech and amazon polly.
- dmos62 7y agopicotts looks cool, but it's not maintained (maybe it's perfect already?). nanotts is a fork of picotts with improvements to its cli interface, but it's not very maintained either (I was not able to compile it, due to it expecting Alsa; somehow -noalsa switch didn't help). I also discovered gtts (and the simpler google_speech), both available through pip. They're interfaces to google's apis, and have routines for handling large texts properly.