11 ms·
Ask HN: Are there any good open source text-to-speech tools?
There are paid services that offer this (e.g. resemble.ai), and a few colab notebooks that I haven't found very helpful, but I wanted to know whether anyone here has had any luck with free text to speech (tts, t2s) tools. Thank you!
- sampo 4y agoMimic3 from the Mycroft project https://github.com/MycroftAI/mimic3 https://github.com/MycroftAI/mimic3
- theCrowing 4y agoThe best is probably tortoise but you have to run it yourself https://github.com/neonbjb/tortoise-tts https://github.com/neonbjb/tortoise-tts here are some demos https://nonint.com/static/tortoise_v2_examples.html https://nonint.com/static/tortoise_v2_examples.html
- smcameron 4y agoOoh, that is a nice one!
- loa_in_ 4y agoEnglish only from a cursory glance
- ipnon 4y agoNLP is going to have this problem for a long time. Obviously most original research is done by Americans in English. There are really only valid training sets for languages that NLP researchers or engineers speak.
- throwaway290 4y agoChatGPT works in Russian for example, don't know about other languages
- jack_pp 4y agoI suspect it might be translated
- throwaway290 4y agoUnless it was trolling I saw evidence it was trained on Russian texts, how else it could do convincing style transfer from Russian poets for example. But as always only successful prompts are shared so I don't know how hit or miss it is
- Cu3PO42 4y agoIt also works in German and I'm relatively certain it's not translated outside the model itself. I've asked it to generate puns incorporating certain words and while the English results were subjectively somewhat better, the German ones were still "fine" and definitely wouldn't work in English.
- underlines 4y agoChatGPT is so crazy it even works in fluent Thai. That's better than any machine translation I've ever tried so far. It even takes cultural differences into account. For example when you ask it to translate "I love you" into Thai, it mentions, that normally you would not say this in the same circumstances as you would say it to your lover in the West, correctly explaining in what circumstances people would really use it, and what to use instead. That's revolutionary for minority languages without a lot of learning material available online. Also I am a native Swiss German speaker. For those who don't know: Swiss German is a dialect continuum, very very different from standard German to an extend, that most untrained German speakers don't understand us. There is no orthography (writing rules), no grammar rules etc. It's a mostly undocumented/unofficial writing system. Only spoken, and the varieties are vast. And guess what, I can write in completely random, informal Swiss German dialect and ChatGPT understands everything, but answers in standard German.
- zone411 4y agoChinese is well-represented among ML researchers.
- ls15 4y agoBecause 14% of the world's population speaks Mandarin Chinese. But what about Yoruba, Burmese or even Hakka Chinese?
- jack_pp 4y agoSpeech will have this problem but text based NLP can be translated and we have pretty good translators
- yewenjie 4y agoDoes anybody run Tortoise on cloud serverless GPUs? If yes, can you please recommend a setup?
- nine_k 4y agoFrom thee link: > Tortoise is a bit tongue in cheek: this model is insanely slow. It leverages both an autoregressive decoder and a diffusion decoder; both known for their low sampling rates. On a NVidia Tesla K80, expect to generate a medium sized sentence every 2 minutes. I suspect that for a real(-ish) time TTS system, something else is needed. OTOH if you want to record some voice acting for a game or other multimedia product, it still may be more cost-effective than recording a bunch of live humans. (K8 = NVidia Tesla K80, GPU, $800-900 for a 24GB version right now.)
- shaklee3 4y agoa k80 is extremely old by now, so I'd expect this to be maybe an order of magnitude faster.
- user3939382 4y agoExamples 4 and 5 sound like George Clooney for some reason.
- godelski 4y agoAs someone that works in generative modeling (vision) there's something that sparks suspicion. They note that these are hand picked results. Has anyone used this and can report the actual quality of results? I bring this up because anyone that has used Stable Diffusion or DallE will know why. Hand picked results are good, but median results matter a lot too.
- echelon 4y agoI'm the author of FakeYou.com and can speak to Tortoise and the TTS field. Tortoise produces quality results with limited training data, but is an extremely slow model that is not suitable for real time use cases. You can't build an app with it. It's good for creatives making one-off deepfake YouTube videos, and that's about it. You're looking for Tacotron 2 or one of its offshoots that add multi-speaker, TorchMoji, etc. You'll want to pair it with the Hifi-Gan vocoder to get end-to-end text to speech. (Avoid Griffin-Lim and WaveGlow.) Your pipeline looks like this at a high level: Input text => Text pre-processing => Synthesizer => Vocoder => [ Optional transcoding ] => Output audio TalkNet is also popular when a secondary reference pitch signal is supplied. You can mimic singing and emotion pretty easily. These three models are faster than real time, and there's a lot of information available and a big community built up around them. FakeYou's Discord has a bunch of people that can show you how to train these models, and there are other Discord communities that offer the same assistance. If you want to train your own voice using your own collected sample data, you can experiment with it on Google Colab and on FakeYou, then reuse the same model file by hosting it in a cloud GPU instance. We can also do the hosting for you if that's not your desire or forte. In any case, these models are solid choices for building consumer apps. As long as you have a GPU, you're good to go. If you're not interested in building or maintaining your own, you can use our API! I'd be happy to help.
- godelski 4y agoThanks for this, I actually appreciate the honesty. It is always difficult for me to parse the actual quality of things I don't have intimate experience with. Can I ask another question? If I wanted to hack around with STT and TTS (inference only) on a pi (4B+) is there anything that is approximately appropriate and can be done on device? (I could process on my main machine but I'd love to do it on the pi even with a decent delay)
- jimmySixDOF 4y ago@yacineMTB (Twitter) used Tortoise to diy his own podcast replicating Joe Rogan (by ChatGPT) & the results are amazing worth a quick listen to get the gist [1] I wrote a script that - pulled @_akhaliq's last 7 days of tweets - fished out the arxiv links - downloaded raw paper .tex - parsed out intros & conclusions - automated a podcast dialogue about the papers w/ web automation & GPT - generated a podcast [1] https://scribepod.substack.com/p/scribepod-1#details https://scribepod.substack.com/p/scribepod-1#details
- 082349872349872 4y agoI've had good luck with https://github.com/espeak-ng/espeak-ng https://github.com/espeak-ng/espeak-ng (for very specific non-english purposes, and I was willing to wrangle IPA)
- smcameron 4y agopico2wave with the -l=en-GB flag to get the British lady voice is not too bad for offline free TTS. You can hear it in this video: https://www.youtube.com/watch?v=tfcme7maygw&t=45s https://www.youtube.com/watch?v=tfcme7maygw&t=45s
- cloverlake 4y agoI've had good results with larynx: https://github.com/rhasspy/larynx https://github.com/rhasspy/larynx
- khobragade 4y agoYeah me too! Sadly unmet dependencies in Debian Sid; it doesn't work anymore :/
- woodlander87 4y agoI've been playing with mimic 3 from Mycroft lately. It's pretty usable out of the box and is self hostable. https://mycroft.ai/mimic-3/ https://mycroft.ai/mimic-3/
- zamalek 4y agoThe bonus with mimic is that it runs reasonably well on constrained hardware (such as an RPi).
- popey 4y agoHey! That ships with my voice! Well, a synthetic version of it anyway. It used to be horribly robotic, but they've improved it quite a bit for mimic-3. Worth a look. I wrote a blog post recently which talks a little about the origin of my voice in Mycroft, in case anyone is interested. https://popey.com/blog/2022/10/blog-to-speech-in-my-voice/ https://popey.com/blog/2022/10/blog-to-speech-in-my-voice/
- thom 4y agoBonus points for models that work well offline on mobile devices.
- jslakro 4y agoA handy way with python https://pyttsx3.readthedocs.io/en/latest/engine.html https://pyttsx3.readthedocs.io/en/latest/engine.html
- vram22 4y agohttps://news.ycombinator.com/item?id=34215278 https://news.ycombinator.com/item?id=34215278
- mindcrime 4y agoDepends on how you define "good". Espeak-ng, for example, works just fine as such. But the quality of the freely available voices is nothing close to the Siri / Google Assistant / Alexa / whatever standard. Understandable? Yes. Usable? Yes. But "good"? Mmmmm... YMMV.
- friend_and_foe 4y agoYeah, I have used espeak, flite tts, RH voice, and a couple of others and they work very well.
- nmfisher 4y agoNot sure if you’re looking to train your own model or just run inference on pretrained models, but if it’s the former, you can find espnet, TensorflowTTS and coqui on GitHub.
- LarryMullins 4y agoI'm not sure about the licensing of all the models/etc, but Coqui AI's 'TTS' python package is fairly good. https://github.com/coqui-ai/TTS https://github.com/coqui-ai/TTS
- IceHegel 4y agoWhen I last compared (about a year ago) Google was the best of the commercial solutions. Is that still the case?
- glandium 4y agoThe TTS in Google Translate is not exactly great, do they have something better? Other than that, I've never used the product, but I've seen Youtube ads for Speechelo, which seems to be quite decent (and a bunch of Youtube ads for other things that quite obviously were using Speechelo (same voice))
- glyphy177 4y ago[dead]
- mijoharas 4y agoI actually looked into this a month or so ago, what I was looking for was just reasonable sounding simple tts driven by the cli (so heavyweight things were out, as were most things with a local server, though I looked at some I think). I ended up going with pico-tts[0]. I remember looking at a few other things and left myself the following comment: # checked out mimic as well. Didn't seem great, espeak is like nails on a chalkboard # haven't checked out marytts or larynx or anything, but this is good enough™ [0] https://github.com/Iiridayn/pico-tts https://github.com/Iiridayn/pico-tts
- richardfeynman 4y agoThis doesn't answer the question, but I thought it might be relevant to mention here that I've been using chatGPT + resemble.ai to create what I believe is the first kid's stories podcast created entirely by AI. Here's how it works: - Kid requests a story about a, b, c on www.makedupstories.com - chatGPT generates the text for a story, summary, and title - we send this to resemble.ai (sounds like Tortoise TTS would work just as well), which has a clone of my voice - the audio file then gets sent to anchor.fm you can listen to example episodes here on Spotify: https://open.spotify.com/show/6liL4T3kJf1scHq134s0mJ https://open.spotify.com/show/6liL4T3kJf1scHq134s0mJ And here on Apple podcasts: https://podcasts.apple.com/us/podcast/kidscast-kids-stories-by-maked-up-stories/id1659969123 https://podcasts.apple.com/us/podcast/kidscast-kids-stories-...
- ancientworldnow 4y agoThere is no joy in this process.
- richardfeynman 4y agoAre you kidding? there's tons of joy. I have a backlog of hundreds of kids story requests and now instead of being able to satisfy one kid per day I can satisfy as many as I want. moreover, kids can't tell the difference and love the stories.
- generalizations 4y agoThey're too young to articulate the difference caused by a lack of emotion in the storytelling; but as kids, they are still early enough in their developmental process that I imagine hearing massive amounts of spoken audio, which lacks emotional depth, will harm them. I'd be cautious.
- richardfeynman 4y agoOn the one hand, there's clear empirical evidence from parents that this helps improve kids' imagination and storytelling ability, and on the other hand we have your pure conjecture that "spoken audio which lacks emotional depth" can harm kids. I'm not buying it, but even if it were true it's pretty clear by the rate of improvement in voice cloning that soon there will be more emotional depth in this form of audio.
- TylerLives 4y agoI have no recommendations, but I'm curious if someone has tried to train a TTS on the data made by one of the commercial services. Generating data would be very cheap, labels perfect, and there would be less noise than in the human datasets.
- charcircuit 4y agoWhy do that instead of using an existing dataset like Common Voice?
- TylerLives 4y agoSince you would be learning from another AI and not humans, there would be much less variation in the way words are pronounced and you would have a lot more data.
- mcint 4y agoWhat's the goal? What would be the benefit of training on a single TTS speaker
- TylerLives 4y agoTo have your own AWS TTS, which you can run offline and for free.
- inoffensivename 4y agodoes festival count as good these days? https://www.cstr.ed.ac.uk/projects/festival/ https://www.cstr.ed.ac.uk/projects/festival/ it's the only one I have any experience with
- windthrown 4y agoI have heard good things about Mozilla's TTS: https://github.com/mozilla/TTS https://github.com/mozilla/TTS
- speedgoose 4y agoIt’s dead unfortunately.
- hifikuno 4y agoFairly certain the team working on it spun off and made Coqui TTS.
- amelius 4y agoPapers-with-code would be the first place to look: https://paperswithcode.com/task/text-to-speech-synthesis https://paperswithcode.com/task/text-to-speech-synthesis
- metiscus 4y agoCoquiAi seems very good from the work I've done with it.
- brianshaler 4y agoI've have pretty good luck with flowtron after watching an nvidia screencast on it. CPU only inference performance isn't great though.
- culi 4y agoThere's... the web platform. No really, there's a SpeechSynthesis API: https://developer.mozilla.org/en-US/docs/Web/API/SpeechSynthesis https://developer.mozilla.org/en-US/docs/Web/API/SpeechSynth...
- jillesvangurp 4y agoIt works, and sounds absolutely terrible on Firefox.
- deleted 4y ago[deleted]
- cahoot_bird 4y agoIt's such an obvious answer perhaps is why nobody has commented it. But depending on the use, you might try web speech API synthesis. For example a Windows user might see a Cortana option whereas a Mac user might see Siri. Demo Here: https://mdn.github.io/dom-examples/web-speech-api/speak-easy-synthesis/ https://mdn.github.io/dom-examples/web-speech-api/speak-easy... Read more here https://github.com/mdn/dom-examples/tree/main/web-speech-api https://github.com/mdn/dom-examples/tree/main/web-speech-api
- n8henrie 4y agoI have a ton of fun using the "say" program on MacOS to write toy programs with my kids, have often wanted a version that could run on my eldest's Manjaro laptop. Are any of the above analogously simple to use?
- dannymi 4y agoespeak (festival)
- geenat 4y agohttps://github.com/gnat/text-to-speech-ubuntu https://github.com/gnat/text-to-speech-ubuntu
- infinite8s 4y agoAre there any that use TensorflowLite?
- deleted 4y ago[deleted]
- dopidopHN 4y agoCoqui is a open source text to speech solution. I haven’t used it in a while but I seen a lot a of new feature listed over the last year or so. Give it a try https://github.com/coqui-ai/TTS https://github.com/coqui-ai/TTS
- mozman 4y agoI’m interested in the opposite: I want to transcribe meetings at work because my memory and note taking are inadequate. I’m familiar with things like otter.ai but I am not risking my job by sharing data with something I don’t control.
- davidamaro 4y agohttps://github.com/chidiwilliams/buzz https://github.com/chidiwilliams/buzz
- sunbum 4y agoOpenAI’s whisper[1] should do the job for you. [1] - https://github.com/openai/whisper https://github.com/openai/whisper
- vram22 4y agoOld, but may be of interest: Speech synthesis in Python with pyttsx https://jugad2.blogspot.com/2014/03/speech-synthesis-in-python-with-pyttsx.html https://jugad2.blogspot.com/2014/03/speech-synthesis-in-pyth...
- tslmy 4y agoGiven a URL, this service return an audio file / stream (in WAV format) that reads out the main content of the webpage. https://github.com/tslmy/tts https://github.com/tslmy/tts
- blacklight 4y agoI've been quite happy with Mimic3 lately (https://github.com/MycroftAI/mimic3 https://github.com/MycroftAI/mimic3), the engine that powers Mycroft. It also comes with an easy-to-install Docker image.
- mmcwilliams 4y agoIf your use case allows for a web API, I've had good experience running OpenTTS[0]. It packages several models, including Coqui AI's TTS which I tend to use the most. There's a handy Docker image, too. [0] https://github.com/synesthesiam/opentts https://github.com/synesthesiam/opentts
- okokwhatever 4y agoFunny how "everybody" is working into the same ChatGPT projects right now (Speech to Text, API integration, TTS...) Somehow it's a nice to observe this trend to start working in other areas.
- var_cw 4y agoFounder of dubverse.ai here. As someone who has done production level TTS(deployed on India's largest news network) I can say there is alot of room for improvement in terms of intelligibility. Most of these open source toolkits/models offer only a certain quality of TTS which is IMO good to play around with but damn too tough to make it sound studio-quality
- o_____________o 4y agoSo what alternatives are there?
- tetmin 4y agoWe just released & open sourced this as a UI & API: https://tts.themetavoice.xyz/ https://tts.themetavoice.xyz/ It's free up to $30 & then cost price after that. It's exceptionally realistic, but can take a bit of time to synthesise as a result.