16 ms·
StyleTTS2 – open-source Eleven-Labs-quality Text To Speech
- zsoltkacsandi 3y agoIs it possible to optimize somehow the model to run a Raspberry with 4 GB of RAM?
- deleted 3y ago[deleted]
- zsoltkacsandi 3y agoI was able to get it work with libjemalloc.
- GaggiX 3y agoHow fast is it on your raspberry?
- zsoltkacsandi 3y agoSuper slow. On my Mac Mini the inference was running in seconds, on Raspberry, minutes.
- synesthesiam 3y agoYou may want to try Piper for this case (RPi 4): https://github.com/rhasspy/piper https://github.com/rhasspy/piper
- modeless 3y agoI made a 100% local voice chatbot using StyleTTS2 and other open source pieces (Whisper and OpenHermes2-Mistral-7B). It responds so much faster than ChatGPT. You can have a real conversation with it instead of the stilted Siri-style interaction you have with other voice assistants. Fun to play with! Anyone who has a Windows gaming PC with a 12 GB Nvidia GPU (tested on 3060 12GB) can install and converse with StyleTTS2 with one click, no fiddling with Python or CUDA needed: https://apps.microsoft.com/detail/9NC624PBFGB7 https://apps.microsoft.com/detail/9NC624PBFGB7 The demo is janky in various ways (requires headphones, runs as a console app, etc), but it's a sneak peek at what will soon be possible to run on a normal gaming PC just by putting together open source pieces. The models are improving rapidly, there are already several improved models I haven't yet incorporated.
- lucubratory 3y agoHow hard on your end does the task of making the chatbot converse naturally look? Specifically I'm thinking about interruptions, if it's talking too long I would like to be able to start talking and interrupt it like in a normal conversation, or if I'm saying something it could quickly interject something. Once you've got the extremely high speed, theoretically faster than real time, you can start doing that stuff right? There is another thing remaining after that for fully natural conversation, which is making the AI context aware like a human would be. Basically giving it eyes so it can see your face and judge body language to know if it's talking too long and needs to be more brief, the same way a human talks.
- modeless 3y agoYes, I implemented the ability to interrupt the chatbot while it is talking. It wasn't too hard, although it does require you to wear headphones so the bot doesn't hear itself and get interrupted. The other way around (bot interrupting the user) is hard. Currently the bot starts processing a response after every word that the voice recognition outputs, to reduce latency. When new words come in before the response is ready it starts over. If it finishes its response before any more words arrive (~1 second usually) it starts speaking. This is not ideal because the user might not be done speaking, of course. If the user continues speaking the bot will stop and listen. But deciding when the user is done speaking, or if the bot should interrupt before the user is done, is a hard problem. It could possibly be done zero-shot using prompting of a LLM but you'd want a GPT-4 level LLM to do a good job and GPT-4 is too slow for instant response right now. A better idea would be to train a dedicated turn-taking model that directly predicts who should speak next in conversations. I haven't thought much about how to source a dataset and train a model for that yet. Ultimately the end state of this type of system is a complete end-to-end audio-to-audio language model. There should be only one model, it should take audio directly as input and produce audio directly as output. I believe that having TTS and voice recognition and language modeling all as separate systems will not get us to 100% natural human conversation. I think that such a system would be within reach of today's hardware too, all you need is the right training dataset/procedure and some architecture bits to make it efficient. As for giving the model eyes, actually there are already open source vision-language models that could be used for this today! I'd love to implement one in my chatbot. It probably wouldn't have social intelligence to read body language yet, but it could definitely answer questions about things you present to the webcam, read text, maybe even look at your computer screen and have conversations about what's on your screen. The latter could potentially be very useful, the endgame there is like GitHub Copilot for everything you do on your computer, not just typing code.
- causality0 3y agoWhat are the chances this gets packaged into something a little more streamlined to use? I have a lot of ebooks I'd love to generate audio versions of.
- carbocation 3y agoHaving now tried it (the linked repo links to pre-built colab notebooks): 1) It does a fantastic job of text-to-speech. 2) I have had no success in getting any meaningful zero-shot voice cloning working. It technically runs and produces a voice, but it sounds nothing like the target voice. (This includes trying their microphone-based self-voice-cloning option.) Presumably fine-tuning is needed - but I am curious if anyone had better luck with the zero-shot approach.
- visarga 3y agoYes, please integrate it with Mistral and Whisper. This has got to get into the LLM frontends.
- modeless 3y agoDone: https://apps.microsoft.com/detail/9NC624PBFGB7 https://apps.microsoft.com/detail/9NC624PBFGB7 It's mostly just a demo for now and a little bit janky but it's fun to chat with and you can see the promise for 100% local voice AI in the future.
- exizt88 3y agoThe weights aren’t MIT-licensed, so this is not usable in commercial applications, right?
- _lvbh 3y agoIt is usable in commercial applications given you disclose the use of AI. This applies only to the pre-trained models. You can train your own from scratch without these restrictions. You can fine tune it on your own voice and also not be required to disclose the use of AI.
- mazoza 3y agomeh this is not that good. Sounds quite boring.
- ChildOfChaos 3y agoAgreed, this isn't Eleven labs quality at all.
- Havoc 3y agoThose sound incredibly good. Though would def like to clone a pleasant voice on it before using. Those sound good but not my cup of tea
- GaggiX 3y agoThey really should have uploaded the models on Huggingface than Gdrive.
- jasonjmcghee 3y agoOut of curiosity - to folks that have had success with this... This voice cloning is... nothing like XTTSv2, let alone ElevenLabs. It doesn't seem to care about accents at all. It does pretty well with pitch and cadence, and that's about it. I've tried all kinds of different values for alpha, beta, embedding scale, diffusion steps. Anyone else have better luck? Sure it's fast and the sound quality is pretty good, but I can't get the voice cloning to work at all.
- carbocation 3y agoI had the same experience as what you described (with a lot of experimentation with alpha and beta, as well as uploading different audio clips).
- dsrtslnd23 3y agoSee the conclusion remarks in the paper - they acknowledge that voice cloning is not that good (yet).
- jsjmch 3y agoSee my previous comment about this point. ElevenLabs are based on Tortoise-TTS which was already pre-trained on millions of hours of data, but this one was only trained on LibriTTS which was 500 hours at best. XTTS was also trained with probably millions of speakers in more than 20 languages. If you have seen millions of voices, there are definitely gonna be some of them that sound like you. It is just a matter of training data, but it is very difficult to have someone collect these large amounts of data and train on it.
- lossolo 3y ago> It is just a matter of training data, but it is very difficult to have someone collect these large amounts of data and train on it. It's really not that difficult, they are trained mostly on audiobooks and high quality audio from yt videos. If we talk about EV model then we are talking about around 500k hours of audio, but Tortoise-TTS is only around 50k from what I remember.
- 3y ago
- wanderingmind 3y agoAs a tangent away from LLMs, is there an integration available to be used in Android as TTS Engine?. The TTS voice that I have now (RHVoice) for OSMAnd is really driving me crazy and almost makes me want to go back to Google Maps.
- lxe 3y agoWow this thing is wicked fast!
- lfmunoz4 3y agoBeen looking for a speech to text that can work in real time and run locally, anyone know which are the best options available?
- kats 3y agoThis is really harmful and unethical work. It will be used to hurt millions of elderly people with scams. That's the real application that will happen 100x more than anything else. It's unethical and harmful to release tools that will be overwhelmingly used to hurt elderly people. What they should do about it is: Stop releasing models. Only release a service so that scammers will not use it. Also, only released audio that is watermarked, so that apps can tell that a phone call might be a scam. When they share models with researchers, use previous best practices: post a Google Form to request access.
- flarg 3y agoMillions of elderly people are already getting scammed by overseas call centers so unless we do something more significant this tech will not make one iota of a difference.
- kats 3y agoThat's not really true, most scammers have a male voice with a heavy accent. When they have tools that easily disguise their voice, scammers can reach many more elderly people.
- slow_numbnut 3y agoThat might have been true about a year ago, but I've been getting calls from well-spoken native-level scammers for about two months now. They are so frequent that I can put them on speaker during family gatherings to raise awareness. Sample sizes of 1 are never representative but they definitely have full access to native speakers or tech that can generate very passable speech.
- maeil 3y agoIt seems quite possible that the change you've seen in these last two months is because some have started using these models. More likely than a sudden huge shift in either the country of origin or English skills of the scammers.
- deknos 3y agoIs this really opensource and/or free software? like code, data(set/s) and models? I am quite tired to see some "open-source" advertisement, where the half or more is not really free. general psa: please be honest in your announcements :|
- _lvbh 3y agoMIT licensed. Models, code, and everything is available right there when you click the link. Maybe actually check it out before complaining.
- mx20 3y agoBut you are wrong the trained models are separate on Google Drive and have following Text that seems to be an additional License Agreement that also includes using the software and any trained Modell. License Part 2 Text: "Before using these pre-trained models, you agree to inform the listeners that the speech samples are synthesized by the pre-trained models, unless you have the permission to use the voice you synthesize. That is, you agree to only use voices whose speakers grant the permission to have their voice cloned, either directly or by license before making synthesized voices pubilc, or you have to publicly announce that these voices are synthesized if you do not have the permission to use these voices."
- _lvbh 3y agoThat’s for the pre-trained models. Train one up yourself.
- Monicjames 3y agoSo, we've got this open-source TTS wizardry going on, which is kinda like if Siri had a caffeine overdose - faster, snappier, and way more fun at parties. This thing is running on gaming rigs with beefy GPUs, and it's apparently so user-friendly, even your grandma could set it up without accidentally summoning a digital demon. But here's the real kicker - it's got the manners of a Victorian gentleman. You can rudely interrupt it mid-sentence, and it'll just stop and listen. Politeness level 100. The reverse, though - getting Mr. Bot to interrupt you - is still in the 'that's too much brain for my silicon' phase. Like, how do you teach a bunch of 1s and 0s to know when you're just taking a dramatic pause or actually done with your TED talk? And get this - they're talking about making this bot read body language. Imagine your laptop judging you for your slouchy posture or that 'I haven't slept properly in days' look. Creepy? Maybe a bit. Cool? Absolutely. In conclusion, StyleTTS2 is shaping up to be the cool new kid on the block, but it's still learning the ropes of human conversation. It's like that super smart friend who knows everything about quantum physics but can't tell when you're sarcastically saying 'Yeah, sure, let's invade Mars tomorrow.
- _lvbh 3y agoI am an introvert: I rarely socialize, listen to podcasts at 2x speed, and mostly use subtitles rather than listening to audio for movies; therefore having a below average ability to differentiate humans/robots. I asked someone to play the recordings for me to differentiate. I could not tell which was human (only between StyleTTS2 and Ground truth. The others were obvious)
- ideasman42 3y agoWhen trying to input a larger amount of text I get the error: The expanded size of the tensor (4293) must match the existing size (512) Any way to fix this from the IPython notebook examples?
- ideasman42 3y agoOnce this is working, is there a simple way to switch voices with the default downloaded models? Or does this require downloading other models or generating them?
- rsbeare 3y agoThis is great! Nice work. I made my own whisper & auto-typer which types what you say (forked whisper-typer). I added OpenAI Q/A and RAG query feature so I could ask it questions (instead of auto keystroke typing) by voice command. For responses to questions, I used Eleven Labs - but even with latency optimized & streaming, it was slow, so disabled it. I just swapped from OpenAI to Mistral 7b for Q/A querying. Much more responsive. Stoked to explore StyleTTS2 now! Really glad that I came across your post. Thank you for sharing!
- sandslides 3y agoJust tried the collab notebooks. Seems to be very good quality. It also supports voice cloning.
- fullstackchris 3y agoGreat stuff, took a look through the README but... what are the minimum hardware requirements to run this? Is this gonna blow up my CPU / harddrive?
- sandslides 3y agoNot sure. The only inference demos are colab notebooks. The models are approx 700mb each so I imagine it will run on modest gpu
- bbbruno222 3y agoWould it run in a cheap non-GPU server?
- dmw_ng 3y agoSeems to run about "2x realtime" on 2015 4 core i7-6700HQ laptop, that is, 5 seconds to generate 10 seconds of output. Can imagine that being 4x or greater on a real machine
- deleted 3y ago[deleted]
- thot_experiment 3y agoI skimmed the github but didn't see any info on this, how long does it take to finetune to a particular voice?
- progbits 3y ago> MIT license > Before using these models, you agree to [...] No, this is not MIT. If you don't like MIT license then feel free to use something else, but you can't pretend this is open source and then attempt to slap on additional restrictions on how the code can be used.
- deleted 3y ago[deleted]
- sandslides 3y agoYes, I noticed that. Doesn't seem right does it
- weego 3y agoI think you mis-parsed the disclaimer. It's just warning people that cloned voices come with a different set of rights to the software (because the person the voice is a clone of has rights to their voice).
- chrismorgan 3y ago(Don’t let’s derail the conversation, please, but “disclaimer” is completely the wrong word here. This is a condition of use. A disclaimer is “this isn’t mine” or “I’m not responsible for this”. Disclaimers and disclosures are quite different things and commonly confused, but this isn’t even either of them.)
- gosub100 3y agoThis always annoys me when people put "disclaimers" on their posts. IANAL, so tired of hearing that one. It's pointless because even if you were a lawyer, you cannot meaningfully comment on a case without the details, jurisdiction, circumstance, etc. Next, it's meaningless because is anyone going to blindly bow down and obey if you state the opposite? "Yes, I AM a lawyer, you do not need to pay taxes, they are unconstitutional." Thirdly, when they "disclaimer" themselves as working at google, that's not a dis-claimer, thats a "claimer", asserting the affirmative. I know their companies require them to not speak for the company without permission, but I hardly ever hear that one, usually its just some useless self-disclosure that they might be biased because they work there. Ok, who isn't biased? What bugs me overall is that it's usually vapid mimicry of a phrase they don't even understand.
- mlsu 3y agoWe're now at "free, local, AI friend that you can have conversations with on consumer hardware" territory. - synthesize an avatar using stablediffusion - synthesize conversation with llama - synthesize the voice with this text thing soon - VR - Video wild times!
- imiric 3y agoI'm looking forward to this tech being used in video games, as well as generative models in general. Interacting with smart NPCs will make everyone's experience different. The avatars themselves could be dynamically generated, and entire environments for that matter. Truly game changing technology for interactive entertainment.
- jpeter 3y agoWhich consumer gpu runs llama 70B?
- sroussey 3y agoProsumer gear. MacBook Pro M3 Max.
- mlsu 3y agoA Mac with a lot of unified RAM can do it, or a dual 3090/4090 setup gets you 48gb of VRAM.
- jadbox 3y agoDoes this actually work? I had thought that you can't use SLI to increase your net memory for the modal?
- speedgoose 3y agoIt works. I use ollama these days, with litellm for the api compatibility, and it seems to use both 24GB GPUs on the server.
- 3y ago
- godelski 3y agoWhy name it Style<anything> if it isn't a StyleGAN? Looks like the first one wasn't either. Interesting to see moves away from flows, especially when none of the flows were modern. Also, is no one clicking on the audio links? There are some... questionable ones... and I'm pretty sure lots of mistakes.
- gwern 3y ago> Looks like the first one wasn't either. The first one says it uses AdaIN layers to help control style? https://arxiv.org/pdf/2205.15439.pdf#page=2 https://arxiv.org/pdf/2205.15439.pdf#page=2 Seems as justifiable as the original StyleGAN calling itself StyleX...
- godelski 3y agoSee my other comment. StyleGAN isn't about AdaIN. StyleGAN2 even modified it.
- lhl 3y agoIt's not called a GAN TTS right? StyleGAN is called what it is because of a "style-based" approach and StyleTTS/2 seems to be doing the same (applying style transfer) through different method (and disentangling style from the rest of the voice synthesis). (Actually, looked at the original StyleTTS paper and it actually even partially uses AdaIN in the decoder, which is the same way that StyleGAN injected style information? Still, I think is besides the point for the naming.)
- godelski 3y agoYeah no I get this but the naming convention has become so prolific that anyone working in generative space hears "Style<thing>" and you should think "GAN". (I work in generative vision btw) My point is not that it is technically right, it is that the name is strongly related with the concept now. Such that if you use a style based network and don't name it StyleX that it's odd and might look like you're trying to claim you've done more. Not that there aren't plenty of GANs that are using Karras's code and called something else. > AdaIN Yes, StyleGAN (version 1) uses AdaIN but StyleGAN2 (and beyond) doesn't. AdaIN stands for Adaptive Instance Normalization. While they use it in that network, to be clear, they did not invent AdaIN and the technique isn't explicit to style, it's a normalization technique. One that StyleGAN2 modifies because the standard one creates strong and localized spikes in the statistics which results in image artifacts.
- api 3y agoIt should be pretty easy to make training data for TTS. The Whisper STT models are open so just chop up a ton of audio and use Whisper to annotate it, then train the other direction to produce audio from text. So you’re basically inverting Whisper.
- nmfisher 3y agoI think you’re talking about just using Whisper to annotate audio for a TTS pipeline but someone from Collabora actually created a TTS model directly from Whisper embeddings https://github.com/collabora/WhisperSpeech https://github.com/collabora/WhisperSpeech
- eginhard 3y agoSTT training data includes all kinds of "noisy" speech so that the model learns to recognise speech in any conditions. TTS training data needs to be as clean as possible so that you don't introduce artefacts in the output and this high-quality data is much harder to get. A simple inversion is not really feasible or at least requires filtering out much of the data.
- satvikpendem 3y agoFunnily enough, the TTS2 examples sound better than the ground truth [0]. For example, the "Then leaving the corpse within the house [...]" example has the ground truth pronounce "house" weirdly, with some change in the tonality that sounds higher, but the TTS2 version sounds more natural. I'm excited to use this for all my ePub files, many of which don't have corresponding audiobooks, such as a lot of Japanese light novels. I am currently using Moon+ Reader on Android which has TTS but it is very robotic. [0] https://styletts2.github.io/ https://styletts2.github.io/
- qingcharles 3y agoFirst Wife is a professional voice-over actor. I saw someone left her a bad review saying "Clearly an AI." 2023. There is no way to win.
- coconut08 3y agohow are you planning on using this with epubs? i'm in a similar boat. would really like to leverage something like this for ebooks.
- satvikpendem 3y agoI wonder if you can add a TTS engine to Android as an app or plugin, then make Moon+ Reader or another reader to use that custom engine. That's probably how I'd do it for the easiest approach, but if that doesn't work, I might just have to make my own app.
- a_wild_dandan 3y agoI’m planning on making a self-host solution where you can upload files and the host sends back the audio to play, as a first pass on this tech. I’ll open source the repo after fiddling and prototyping. I’ve needed this kinda thing for a long time!
- coconut08 3y ago
- lhl 3y agoI tested StyleTTS2 last month, my step-by-step notes that might be useful for people doing local setup (not too hard): https://llm-tracker.info/books/howto-guides/page/styletts-2 https://llm-tracker.info/books/howto-guides/page/styletts-2 Also I did a little speed/quality shootoff with the LJSpeech model (vs VITS and XTTS). StyleTTS2 was pretty good and very fast: https://fediverse.randomfoo.net/notice/AaOgprU715gcT5GrZ2 https://fediverse.randomfoo.net/notice/AaOgprU715gcT5GrZ2
- rahimnathwani 3y agoThanks. Following the instructions now. BTW mamba is no longer recommended (for those like me who aren't already using it), and the #mambaforge anchor in the link didn't work.
- lhl 3y agoI switched from conda to mamba a while ago and never looked back (it's probably saved dozens of hours from waiting for conda's slow as molasses package resolution). I'm looking at the latest docs and it doesn't look like there's any deprecation messages or anything (it does warn against installing mamba inside of conda, but that's been the case for a long time): https://mamba.readthedocs.io/en/latest/installation/mamba-installation.html https://mamba.readthedocs.io/en/latest/installation/mamba-in... It looks like miniforge is still the recommended install method, but also the anchor has changed in the repo docs, which I've updated, thx. FWIW, I haven't run into any problems using mamba. While I'm not a power user, so there are edge cases I might have missed, but I have over 35 mamba envs on my dev machine atm, so it's definitely been doing the job for me and remains wicked fast (if not particularly disk efficient).
- rahimnathwani 3y agoI had somehow missed the introduction of mamba, and have been using the default conda solver (which I think is the 'classic' one). Apparently conda now supports using the mamba solver: https://www.anaconda.com/blog/a-faster-conda-for-a-growing-community https://www.anaconda.com/blog/a-faster-conda-for-a-growing-c... conda update -n base conda conda install -n base conda-libmamba-solver conda config --set solver libmamba
- jasonjmcghee 3y agoI've been playing with XTTSv2 and on my 3080ti, and it's sightly faster than the length of the final audio. It's also good quality, but these samples sound better. Excited to try it out!
- gjm11 3y agoHN title at present is "StyleTTS2 – open-source Eleven Labs quality Text To Speech". Actual title at the far end doesn't name any particular other product; arXiv paper linked from there doesn't mention Eleven Labs either. I thought this sort of editorializing was frowned on.
- modeless 3y agoIt is editorializing and it is an exaggeration. However I've been using StyleTTS2 myself and IMO it is the best open source TTS by far and definitely deserves a spot on the top of HN for a while.
- stevenhuang 3y agoEleven Labs is the gold standard for voice synthesis. There is nothing better out there. So it is extremely notable for an open source system to be able to approach this level of quality, which is why I'd imagine most would appreciate the comparison. I know it caught my attention.
- yreg 3y agoBut is this even approaching Eleven? Doesn't seem like it from the other comments here.
- lucubratory 3y agoOpenAI's TTS is better than Eleven Labs, but they don't let you train it to have a particular voice out of fear of the consequences.
- huac 3y agoI concur that, for the use cases that OpenAI's voices cover, it is significantly better than Eleven.
- GaggiX 3y agoYes, it's against the guidelines. In fact, when I read the title, I didn't think it was a new research paper but a random GitHub project.
- stevenhuang 3y agoI really want to try this but making the venv to install all the torch dependencies is starting to get old lol. How are other people dealing with this? Is there an easy way to get multiple venvs to share like a common torch venv? I can do this manually but I'm wondering if there's a tool out there that does this.
- stavros 3y agoI generally try to use Docker for this stuff, but yeah, it's the main reason why I pass on these, even though I've been looking for something like this. It's just too hard to figure out the dependencies.
- amelius 3y ago> is starting to get old lol. If it's starting to get old, then this means that an LLM like Copilot should be able to do it for you, no?
- stevenhuang 3y agoI mean that I already have like 10 different torch venvs for different projects all with various pinned versions and CUDA variants. Still worth the trade-off of not having to deal with dependency hell, but you start to wonder if there is a better way. All together this is many GBs of duplicated libs, wasted bandwidth and compute.
- wczekalski 3y agoI use nix to setup the python env (python version + poetry + sometimes python packages that are difficult to install with poetry) and use poetry for the rest. The workflow is: > nix flake init -t github:dialohq/flake-templates#python > nix develop -c $SHELL > # I'm in the shell with poetry env, I have a shell hook in the nix devenv that does poetry install and poetry activate.
- lukasga 3y agoCan relate to this problem a lot. I have considered starting using a Docker dev container and making a base image for shared dependencies which I then can customize in a dockerfile for each new project, not sure if there's a better alternative though.
- victorbjorklund 3y agoThis only works for English voices right?
- e12e 3y agoNo? From the readme: In Utils folder, there are three pre-trained models: ASR folder: It contains the pre-trained text aligner, which was pre-trained on English (LibriTTS), Japanese (JVS), and Chinese (AiShell) corpus. It works well for most other languages without fine-tuning, but you can always train your own text aligner with the code here: yl4579/AuxiliaryASR. JDC folder: It contains the pre-trained pitch extractor, which was pre-trained on English (LibriTTS) corpus only. However, it works well for other languages too because F0 is independent of language. If you want to train on singing corpus, it is recommended to train a new pitch extractor with the code here: yl4579/PitchExtractor. PLBERT folder: It contains the pre-trained PL-BERT model, which was pre-trained on English (Wikipedia) corpus only. It probably does not work very well on other languages, so you will need to train a different PL-BERT for different languages using the repo here: yl4579/PL-BERT. You can also replace this module with other phoneme BERT models like XPhoneBERT which is pre-trained on more than 100 languages.
- modeless 3y agoThose are just parts of the system and don't make a complete TTS. In theory you could train a complete StyleTTS2 for other languages but currently the pretrained models are English only.
- svapnil 3y agoHow fast is inference with this model? For reference, I'm using 11Labs to synthesize short messages - maybe a sentence or something, using voice cloning, and I'm getting it at around 400 - 500ms response times. Is there any OS solution that gets me to around the same inference time?
- wczekalski 3y agoIt depends on hardware but IIRC on V100s it took 0.01-0.03s for 1s of audio.
- eigenvalue 3y agoWas somewhat annoying to get everything to work as the documentation is a bit spotty, but after ~20 minutes it's all working well for me on WSL Ubuntu 22.04. Sound quality is very good, much better than other open source TTS projects I've seen. It's also SUPER fast (at least using a 4090 GPU). Not sure it's quite up to Eleven Labs quality. But to me, what makes Eleven so cool is that they have a large library of high quality voices that are easy to choose from. I don't yet see any way with this library to get a different voice from the default female voice. Also, the real special sauce for Eleven is the near instant voice cloning with just a single 5 minute sample, which works shockingly (even spookily) well. Can't wait to have that all available in a fully open source project! The services that provide this as an API are just too expensive for many use cases. Even the OpenAI one which is on the cheaper side costs ~10 cents for a couple thousand word generation.
- eigenvalue 3y agoTo save people some time, this is tested on Ubuntu 22.04 (google is being annoying about the download link, saying too many people have downloaded it in the past 24 hours, but if you wait a bit it should work again): git clone https://github.com/yl4579/StyleTTS2.git cd StyleTTS2 python3 -m venv venv source venv/bin/activate python3 -m pip install --upgrade pip python3 -m pip install wheel pip install -r requirements.txt pip install phonemizer sudo apt-get install -y espeak-ng pip install gdown gdown https://drive.google.com/uc?id=1K3jt1JEbtohBLUA0X75KLw36TW7U1yxq 7z x Models.zip rm Models.zip gdown https://drive.google.com/uc?id=1jK_VV3TnGM9dkrIMsdQ_upov8FrIymr7 7z x Models.zip rm Models.zip pip install ipykernel pickleshare nltk SoundFile python -c "import nltk; nltk.download('punkt')" pip install --upgrade jupyter ipywidgets librosa python -m ipykernel install --user --name=venv --display-name="Python (venv)" jupyter notebook Then navigate to /Demo and open either `Inference_LJSpeech.ipynb` or `Inference_LibriTTS.ipynb` and they should work.
- deleted 3y ago[deleted]
- degobah 3y agoVery helpful, thanks!
- Evidlo 3y agoWhat's a ballpark estimate for inference time on a modern CPU?
- beltsazar 3y agoIf AI will render some jobs obsolete, I suppose the first one will be audio book narrators and voice actors.
- riquito 3y agoI can see a future where the label "100% narrated by a human" (and similar in other industries) will be a thing
- washadjeffmad 3y agoHardly. Imagine licensing your voice to Amazon so that any customer could stream any book narrated in your likeness without you having to commit the time to record. You could still work as a custom voice artist, all with a "no clone" clause if you chose. You could profit from your performance and craft in a fraction of the time, focusing as your own agent on the management of your assets. Or, you could just keep and commit to your day job. Just imagine hearing the final novel of ASoIaF narrated by Roy Dotrice and knowing that a royalty went to his family and estate, or if David Attenborough willed the digital likeness of his voice and its performance to the BBC for use in nature documentaries after his death. The advent of recorded audio didn't put artists out of business, it expanded the industries that relied on them by allowing more of them to work. Film and tape didn't put artists out of business, it expanded the industries that relied on them by allowing more of them to work. Audio digitization and the internet didn't put artists out of business; it expanded the industries that relied on them by allowing more of them to work. And TTS won't put artists out of business, but it will create yet another new market with another niche that people will have to figure out how to monetize, even though 98% of the revenues will still somehow end up with the distributors.
- tomcam 3y agoVery impressive. It would take me a long time to even guess that some of these are text to speech.
- carbocation 3y agoCurious if we'll see a Civitai-style LoRA[1] marketplace for text-to-speech models. 1 = https://github.com/microsoft/LoRA https://github.com/microsoft/LoRA
- swyx 3y agosilicon valley is very leaky, eleven labs is widely rumored to have raised a huge round recently. great timing because with OpenAI's TTS and now this thing the options in the market have just expanded greatly.
- readyplayernull 3y agoSomeone please create a TTS with marked-down emotions/intonations.
- wg0 3y agoThe quality is really really INSANE and pretty much unimaginable in early 2000s. Could have interesting prospects for games where you have LLM assuming a character and such TTS giving those NPCs voice.
- beachy 3y agoThis is a big thing for one area I'm interested in - golf simulation. Currently playing in a golf simulator has a bit of a post-apocalyptian vibe. The birds are cheeping, the grass is rustling, the game play is realistic, but there's not a human to be seen. Just so different from the smacktalking of a real round, or the crowd noise at a big game. It's begging for some LLM-fuelled banter to be added.
- wahnfrieden 3y agoIs there a way to port this to iOS? Apple doesn't provide an API for their version of this.
- ddmma 3y agoWell done, been waiting for a moment like this. Will give it a try!
- motivence7856 3y ago[flagged]