6 ms·
This is HN, so I'm surprised that no one in the comments section has run this locally. :) Following the instructions in their repo (and moving the checkpoints/
by randkyp 3y ago
This is HN, so I'm surprised that no one in the comments section has run this locally. :)
Following the instructions in their repo (and moving the checkpoints/ and resources/ folder into the "nested" openvoice subfolder), I managed to get the Gradio demo running. Simple enough.
It appears to be quicker than XTTS2 on my machine (RTX 3090), and utilizes approximately 1.5GB of VRAM. The Gradio demo is limited to 200 characters, perhaps for resource usage concerns, but it seems to run at around 8x realtime (8 seconds of speech for about 1 second of processing time.)
EDIT: patched the Gradio demo for longer text; it's way faster than that. One minute of speech only took ~4 seconds to render. Default voice sample, reading this very comment: https://voca.ro/18JIHDs4vI1v https://voca.ro/18JIHDs4vI1v
I had to write out acronyms -- XTTS2 to "ex tee tee ess two", for example.
The voice clarity is better than XTTS2, too, but the speech can sound a bit stilted and, well, robotic/TTS-esque compared to it. The cloning consistency is definitely a step above XTTS2 in my experience -- XTTS2 would sometimes have random pitch shifts or plosives/babble in the middle of speech.
- bambax 3y agoI am trying to run it locally but it doesn't quite work for me. I was able to run the demos allright, but when trying to use another reference speaker (in demo_part1), the result doesn't sound at all like the source (it's just a random male voice). I'm also trying to produce French output, using a reference audio file in French for the base speaker, and a text in French. This triggers an error in api.py line 75 that the source language is not accepted. Indeed, in api.py line 45 the only two source languages allowed are English and Chineese; simply adding French to language_marks in api.py line 43 avoids errors but produces a weird/unintelligible result with a super heavy English accent and pronunciation. I guess one would need to generate source_se again, and probably mess with config.json and checkpoint.pth as well, but I could not find instructions on how to do this...? Edit -- tried again on https://app.myshell.ai/ https://app.myshell.ai/ The result sounds French alright, but still nothing like the original reference. It would be absolutely impossible to confuse one with the other, even for someone who didn't know the person very well.
- randkyp 3y agoI played with it some more and I have to agree. For actual voice _cloning_, XTTS2 sounds much, much closer to the original speaker. But the resulting output is also much more unpredictable and sometimes downright glitchy compared to OpenVoice. XTTS2 also tries to "act out" the implied emotion/tone/pitch/cadence in the input text, for better or worse. But my use case is just to have a nice-sounding local TTS engine, and current text-to-phoneme conversion quirks aside, OpenVoice seems promising. It's fast, too.
- echelon 3y agoAnd StyleTTS2 generalizes out of domain even better than that.
- dragonwriter 3y ago> but when trying to use another reference speaker (in demo_part1), the result doesn’t sound at all like the source I’ve noticed the same thing and I wonder if there is maybe some undocumented information about what makes a good voice sample for cloning, perhaps in terms of what you might call “phonemic inventory”. The reference sample seems really dense. > Indeed, in api.py line 45 the only two source languages allowed are English and Chinese If you look at the code, outside of what the model does it relies on the surrounding infrastructure converting the input text to the international phonetic alphabet (IPA) as part of the process, and only has that implemented for English and Mandarin (though cleaners.py has broken references to routines for Japanese and Korean.
- epiccoleman 3y agoI have got to build or buy a new computer capable of playing with all this cool shit. I built my last "gaming" PC in 2016, so its hardware isn't really ideal for AI shenanigans, and my Macbook for work is an increasingly crusty 2019 model, so that's out too. Yeah, I could rent time on a server, but that's not as cool as just having a box in my house that I could use to play with local models. Feels like I'm missing a wave of fun stuff to experiment with, but hardware is expensive!
- beardedwizard 3y agoI would love a recommendation for an off the shelf "gpu server" good for most of this that I can run at home.
- lakomen 3y agoI'm clueless about AI, but here's a benchmark list https://www.videocardbenchmark.net/high_end_gpus.html https://www.videocardbenchmark.net/high_end_gpus.html Imo the 4070 super is the best value and consumes the least amount of Watts, 220 in all the top 10. So anything with one and some ECC RAM aka AMD should be fine. Intel non-xeons need the expensive w680 boards and very specific RAM per board. ECC because you wrote server. We're professionals here after all, right?
- zoklet-enjoyer 3y agoI forgot all about Vocaroo!
- causi 3y agoWe're so close to me being able to open a program, feed in an epub, and get a near-human level audiobook out of it. I'm so excited.
- aedocw 3y agoGive https://github.com/aedocw/epub2tts https://github.com/aedocw/epub2tts a look, the latest update enables use of MS Edge cloud-based TTS so you don't need a local GPU and the quality is excellent.
- causi 3y agoInteresting. Seems like a pain to get running but I'll give it a shot. Thanks.
- jurimasa 3y agoI think this is creepy and dangerous as fuck. Not worth the trouble it will be.
- CamperBob2 3y agoOther sites beckon.
- _zoltan_ 3y agoyou're gonna be REALLY surprised out there in the real world.
- lessolives 3y ago[dead]
- aftbit 3y agoI want to try chaining XTTS2 with something like RVCProject. The idea is to generate the speech in one step, then clone a voice in the audio domain in a second step.
- fellowniusmonk 3y agoI'm running it locally on my M1. The reference voices sound great, trying to clone my own voice it doesn't sound remotely like me.