4 ms·
Was somewhat annoying to get everything to work as the documentation is a bit spotty, but after ~20 minutes it's all working well for me on WSL Ubuntu 22.04. So
by eigenvalue 3y ago
Was somewhat annoying to get everything to work as the documentation is a bit spotty, but after ~20 minutes it's all working well for me on WSL Ubuntu 22.04. Sound quality is very good, much better than other open source TTS projects I've seen. It's also SUPER fast (at least using a 4090 GPU).
Not sure it's quite up to Eleven Labs quality. But to me, what makes Eleven so cool is that they have a large library of high quality voices that are easy to choose from. I don't yet see any way with this library to get a different voice from the default female voice.
Also, the real special sauce for Eleven is the near instant voice cloning with just a single 5 minute sample, which works shockingly (even spookily) well. Can't wait to have that all available in a fully open source project! The services that provide this as an API are just too expensive for many use cases. Even the OpenAI one which is on the cheaper side costs ~10 cents for a couple thousand word generation.
- eigenvalue 3y agoTo save people some time, this is tested on Ubuntu 22.04 (google is being annoying about the download link, saying too many people have downloaded it in the past 24 hours, but if you wait a bit it should work again): git clone https://github.com/yl4579/StyleTTS2.git cd StyleTTS2 python3 -m venv venv source venv/bin/activate python3 -m pip install --upgrade pip python3 -m pip install wheel pip install -r requirements.txt pip install phonemizer sudo apt-get install -y espeak-ng pip install gdown gdown https://drive.google.com/uc?id=1K3jt1JEbtohBLUA0X75KLw36TW7U1yxq 7z x Models.zip rm Models.zip gdown https://drive.google.com/uc?id=1jK_VV3TnGM9dkrIMsdQ_upov8FrIymr7 7z x Models.zip rm Models.zip pip install ipykernel pickleshare nltk SoundFile python -c "import nltk; nltk.download('punkt')" pip install --upgrade jupyter ipywidgets librosa python -m ipykernel install --user --name=venv --display-name="Python (venv)" jupyter notebook Then navigate to /Demo and open either `Inference_LJSpeech.ipynb` or `Inference_LibriTTS.ipynb` and they should work.
- deleted 3y ago[deleted]
- degobah 3y agoVery helpful, thanks!
- wczekalski 3y agohave you tested longer utterances with both ElevenLabs and with StyleTTS? Short audio synthesis is a ~solved problem in the TTS world but things start falling apart once you want to do something like create an audiobook with text to speech.
- wingworks 3y agoI can say that the paid service from ElevenLabs can do long form TTS very well. I used it for a while to convert long articles to voice to listen to later instead of reading. It works very well. I only stopped because it gets a little pricey.
- stavros 3y agoThe OpenAI API is ten times cheaper and a fair bit faster. Also, ElevenLabs keeps diverging for me, and starts mispronouncing words after two or three sentences.
- wczekalski 3y agoOne thing I've seen done for style cloning is a high quality fine tuned TTS -> RVC pipeline to "enhance" the output. TTS for intonation + pronunciation, RVC for voice texture. With StyleTTS and this pipeline you should get close to ElevenLabs.
- KolmogorovComp 3y agoRVC? R… Voice Model?
- eigenvalue 3y agoI suspect they are doing many more things to make it sounds better. I certainly hope open source solutions can approach that level of quality, but so far I've been very disappointed.
- sandslides 3y agoThe LibriTTS demo clones unseen speakers from a five second or so clip
- eigenvalue 3y agoAh ok, thanks. I tried the other demo.
- eigenvalue 3y agoI tried it. Sounds absolutely nothing like my voice or my wife's voice. I used the same sample files as I used 2 days ago on the Eleven Labs website, and they worked flawlessly there. So this is very, very far from being close to "Eleven Labs quality" when it comes to voice cloning.
- thot_experiment 3y agoAh that's disappointing, have you tried https://git.ecker.tech/mrq/ai-voice-cloning https://git.ecker.tech/mrq/ai-voice-cloning ? I've had decent results with that, but inference is quite slow.
- jsjmch 3y agoElevenLabs are based on Tortoise-TTS which was already pre-trained on millions of hours of data, but this one was only trained on LibriTTS which was 500 hours at best. If you have seen millions of voices, there are definitely gonna be some of them that sound like you. It is just a matter of training data, but it is very difficult to have someone collect these large amounts of data and train on it.
- lewismenelaws 3y agoYep. Tried as well. Tried a little clip of Tony Sopranos and it came out as a british guy. xTTSv2 does it much better. But the quality on the trained voices are great though.
- eigenvalue 3y ago