6 ms·
Sopro TTS: A 169M model with zero-shot voice cloning that runs on the CPU
- brikym 9mo agoA scammers dream.
- jacquesm 9mo agoThat's exactly how I see it.
- soulofmischief 9mo agoUnfortunately, we have to prepare for a future where this kind of stuff is everywhere. We will have to rethink how trust is modeled online and offline.
- gosub100 9mo agounfortunately I think you're right, the cons massively outweigh the pros. One constructive use would be making on-demand audiobooks.
- CoastalCoder 9mo agoI agree. I'd be curious to hear why its advocates believe that this is a net win for society.
- Alex2037 9mo agoit doesn't need to be. are video games a net win for society? is porn?
- lukebechtel 9mo agoVery cool. I'd love a slightly larger version with hopefully improved voice quality. Nice work!
- sammyyyyyyy 9mo agoThanks! Yeah I kinda postponed publishing it until it was a bit better, but as a perfectionist, it would have never been published
- lukebechtel 9mo agounderstood! Glad you shipped.
- convivialdingo 9mo agoImpressive! The cloning and voice affect is great. Has a slight warble in the voice on long vowels, but not a huge issue. I'll definitely check it out - we could use voice generation for alerting on one of our projects (no GPUs on hardware).
- sammyyyyyyy 9mo agoCool! Yeah the voice quality really depends on the reference audio. Also mess with the parameters. All the feedback is welcome
- realityfactchex 9mo agoThat's cool and useful. IMO, the best alternative is Chatterbox-TTS-Server [0] (slower, but quite high quality). [0] https://github.com/devnen/Chatterbox-TTS-Server https://github.com/devnen/Chatterbox-TTS-Server
- iLoveOncall 9mo agoChatterbox-TTS has a MUCH MUCH better output quality though, the quality of the output from Sopro TTS (based on the video embedded on GitHub) is absolutely terrible and completely unusable for any serious application, while Chatterbox has incredible outputs. I have an RTX5090, so not exactly what most consumers will have but still accessible, and it's also very fast, around 2 seconds of audio per 1 second of generation. Here's an example I just generated (first try, 22 seconds runtime, 14 seconds of generation): https://jumpshare.com/s/Vl92l7Rm0IhiIk0jGors https://jumpshare.com/s/Vl92l7Rm0IhiIk0jGors Here's another one, 20 seconds of generation, 30 seconds of runtime, which clones a voice from a Youtuber (I don't use it for nefarious reasons, it's just for the demo): https://jumpshare.com/s/Y61duHpqvkmNfKr4hGFs https://jumpshare.com/s/Y61duHpqvkmNfKr4hGFs with the original source for the voice: https://www.youtube.com/@ArbitorIan https://www.youtube.com/@ArbitorIan
- sammyyyyyyy 9mo agoYou should try it! I wouldn’t say it’s the best, far from that. But also wouldn’t say it’s terrible. If you have a 5090, then yes, you can run much more powerful models in real time. Chatterbox is a great model though
- iLoveOncall 9mo ago> But also wouldn’t say it’s terrible. But you included 3 samples on your GitHub video and they all sound extremely robotic and have very bad artifacts?
- samuel-vitorino 9mo ago[dead]
- blitzar 9mo agoMission impossible cloning skills without the long compile time. "The pleasure of Buzby's company is what I most enjoy. He put a tack on Miss Yancy's chair ..." https://www.youtube.com/watch?v=H2kIN9PgvNo https://www.youtube.com/watch?v=H2kIN9PgvNo https://literalminded.wordpress.com/2006/05/05/a-panphonic-poem-for-mission-impossible-3/ https://literalminded.wordpress.com/2006/05/05/a-panphonic-p...
- btbuildem 9mo agoIt's impressive given the constraints! Would you consider releasing a more capable version that renders with fewer artifacts (and maybe requires a bit more processing power)? Chatterbox is my go-to, this could be a nice alternative were it capable of high-fidelity results!
- sammyyyyyyy 9mo agoThis is my side “hobby”. And compute is quite expensive. But if the community’s responsive is good, I will definitely think about it! Btw, chatterbox is a great model and inspiration
- bicepjai 9mo agoThanks can you share details about compute economics you dealt with ?
- sammyyyyyyy 9mo agoYeah sure. The training was about ~250 dollars, which is quite low by today’s standards. And I spent a bit more on ablations and research
- bicepjai 9mo agoI was on similar path and saw my bills going over 1000 dollars as interests to do research and ablations grew. Then I decided to get one Blackwell Pro 6000 and trying things with that :) If you have suggestions on how to manage metrics let us know. Currenty trying langfuse since its one click install on coolify
- btbuildem 9mo agoIs that something that could be done on a local setup? Eg, 2x RTX3090?
- littlestymaar 9mo agoVery cool work, especially for a hobby project. Do you have any plans to publish a blog post on how you did that? ?What training data and how much? Your training and ablations methodology, etc.
- elaus 9mo agoVery nice to have done this by yourself, locally. I wish there was an open/local tts model with voice cloning as good as 11l (for non-english languages even)
- sammyyyyyyy 9mo agoYeah, we are not quite there, but I’m sure we are not far either
- SoftTalker 9mo agoWhat does "zero-shot" mean in this context?
- nateb2022 9mo ago> Zero-shot learning (ZSL) is a problem setup in deep learning where, at test time, a learner observes samples from classes which were not observed during training, and needs to predict the class that they belong to. The name is a play on words based on the earlier concept of one-shot learning, in which classification can be learned from only one, or a few, examples. https://en.wikipedia.org/wiki/Zero-shot_learning https://en.wikipedia.org/wiki/Zero-shot_learning edit: since there seems to be some degree of confusion regarding this definition, I'll break it down more simply: We are modeling the conditional probability P(Audio|Voice). If the model samples from this distribution for a Voice class not observed during training, it is by definition zero-shot. "Prediction" here is not a simple classification, but the estimation of this conditional probability distribution for a Voice class not observed during training. Providing reference audio to a model at inference-time is no different than including an AGENTS.md when interacting with an LLM. You're providing context, not updating the model weights.
- woodson 9mo agoThis generic answer from Wikipedia is not very helpful in this context. Zero-shot voice cloning in TTS usually means that data of the target speaker you want the generated speech to sound like does not need to be included in the training data used to train the TTS models. In other words, you can provide an audio sample of the target speaker together with the text to be spoken to generate the audio that sounds like it was spoken by that speaker.
- coder543 9mo agoWhy wouldn’t that be one-shot voice cloning? The concept of calling it zero shot doesn’t really make sense to me.
- 9mo ago
- derefr 9mo agoIs there yet any model like this, but which works as a "speech plus speech to speech" voice modulator — i.e. taking a fixed audio sample (the prompt), plus a continuous audio stream (the input), and transforming any speech component of the input to have the tone and timbre of the voice in the prompt, resulting in a continuous audio output stream? (Ideally, while passing through non-speech parts of the input audio stream; but those could also be handled other ways, with traditional source separation techniques, microphone arrays, etc.) Though I suppose, for the use-case I'm thinking of (v-tubers), you don't really need the ability to dynamically change the prompt; so you could also simplify this to a continuous single-stream "speech to speech" model, which gets its target vocal timbre burned into it during an expensive (but one-time) fine-tuning step.
- vunderba 9mo agoI don’t know about open models, but ElevenLabs has had this idea of mapping intonation/emotion/inflections onto a designated TTS voice for a while. https://elevenlabs.io/blog/speech-to-speech https://elevenlabs.io/blog/speech-to-speech
- gcr 9mo agoChatterbox TTS does this in “voice cloning” mode but you have to implement the streaming part yourself. There are two inputs: audio A (“style”) and B (“content”). The timbre is taken from A, and the content, pronunciation, prosody, accent, etc is taken from B. Strictly soeaking, voice cloning models like this and chatterbox are not “TTS” - they’re better thought of as “S+STS”, that is, speech+style to speech
- qingcharles 9mo agoThere must be something out there that does this reliably as I often see/hear v-tubers doing it.
- lumerios 9mo agoyes, check out RVC (retrieval voice conversation) which I believe is the only good open source voice changer. Currently there's a bit of a conflict between the original creator and current developers. So don't use the main fork. I think you'll be able to find a more up-to-date fork that's in english.
- nunobrito 9mo agoMuito fixe. Now the next challenge (for me) is how to convert this to DART and run on Android. :-)
- sammyyyyyyy 9mo agoObrigado! Quando (e se fizeres isso) manda pm!
- woodson 9mo agoDoes the 169M include the ~90M params for the Mimi codec? Interesting approach using FiLM for speaker conditioning.
- sammyyyyyyy 9mo agoNo, it doesn’t.
- jacquesm 9mo agoWhat could possibly go wrong... Don't you ever think about what the balance of good and bad is when you make something like this? What's the upside? What's the downside? In this particular case I can only see downsides, if there are upsides I'd love to hear about them. All I see is my elderly family members getting 'me' on their phones asking for help, and falling for it. I've gotten into the habit of waiting for the other person to speak first when I answer the phone now and the number is unknown to me.
- sammyyyyyyy 9mo agoYes, you are right. However, there are many upsides to this kind of technology. For example, it can restore the voices of people who were affected by numerous diseases
- jacquesm 9mo agoOk, that's an interesting angle, I had not thought of that, but of course you'd still need a good sample of them from before that happened. Thank you for the explanation.
- Alex2037 9mo agoare you under the impression that this is the first such tool? it's not. it's not even the hundredth. this Pandora's box has been opened a long time ago.
- idiotsecant 9mo agoThere is no such thing as bad technology.
- jacquesm 9mo agoThat is simply not true. There is lots of bad technology.
- idiotsecant 9mo ago
- sergiotapia 9mo agoIt sounds a lot like RFK Jr! Does anyone have any more casual examples?
- guerrilla 9mo agoI don't understand the comments here at all. I played the audio and it sounds absolutely horrible, far worse than computer voices sounded fifteen years ago. Not even the most feeble minded person would mistake that as a human. Am I not hearing the same thing everyone else is hearing? It sounds straight up corrupted to me. Tested in different browsers, no difference.
- foolserrandboy 9mo agoI thought it was RFK
- serf 9mo agospasmodic dysphonia as a service.
- sammyyyyyyy 9mo agoAs I said, some reference voices can lead to bad voice quality. But if it sounds that bad, it’s probably not it. Would love to dig into it if you want
- guerrilla 9mo agoI mean I'm talking about the mp4. How could people possibly be worried about scammers after listening to that?
- sammyyyyyyy 9mo agoI didn’t specially cherry pick those examples. You can try it anyway for yourself. But thanks for the feedback anyway
- guerrilla 9mo agoNo shade on you. It's definitely impressive. I just didn't understand people's reactions.
- Gathering6678 9mo agoEmm...I played the sample audio and it was...horrible? How is it voice cloning if even the sample doesn't sound like any human being...
- sammyyyyyyy 9mo agoI should have posted the reference audio used with the examples. Honestly it doesn’t sound so different from them. Voice cloning can be from a cartoon too, doesn’t have to be from a human being
- nemomarx 9mo agoA before / after with the reference and output seems useful to me, and maybe a range from more generic to more recognizable / celebrity voice samples so people can kinda see how it tackles different ones? (Prominent politician or actor or somebody with a distinct speaking tone?)
- Gathering6678 9mo agoThat is probably a good idea. I was so confused listening to the example.
- sammyyyyyyy 9mo agoAlso, I didn’t want to use known voices as the example, so I ended up using generic ones from the datasets
- krunck 9mo agoI just had some amusing results using text with lots of exclamations and turning up the temperature. Good fun.
- yamal4321 9mo agoTried english. There are similarities. Really impressive for such budget Also increadibly easy to use, thanks for this
- xiconfjs 9mo agoBut its english-only - so what else could you have tried? Asking because I‘m interested in a german version :)
- VerifiedReports 9mo agoWhat is "zero-shot" supposed to mean?
- mikalauskas 9mo ago[dead]
- carteazy 9mo agoI believe in this case it means that you do not need to provide other voice samples to get a good clone.
- spwa4 9mo agoIt means there is zero training involved in getting from voice sample to voice duplicate. There used to be models that take a voice sample, run 5 or 10 training iterations (which of course takes 10 mins, or a few hours if you have hardware as shitty as mine), and only then duplicate the voice. This you give the voice sample as part of the input, and immediately it tries to duplicate the voice.
- x3haloed 9mo agoDoesn’t NeuTTS work the same way?
- onion2k 9mo agozero-shot is a single prompt (maybe with additional context in the form of files.) few-shot is providing a few examples to steer the LLM multi-shot is a longer cycle of prompts and refinement
- moffkalast 9mo agoif you had one-shot or one opportunity
- nake89 9mo ago
- LoveMortuus 9mo agoThis is very cool! And it'll only get better. I do wonder, if, at least as a patch-up job, they could do some light audio processing to remove the raspiness from the voices.
- armcat 9mo agoSuper nice! I've been using Kokoro locally, which is 82M parameters and runs (and sounds) amazing! https://huggingface.co/hexgrad/Kokoro-82M https://huggingface.co/hexgrad/Kokoro-82M
- machiaweliczny 9mo agoI tried Kokoro-JS that I think runs in browser and it was too way too slow with latency also not supporting language I wanted
- armcat 9mo agoI have a 5070 in my rig. What I'm running is Kokoro in a Python/FastAPI backend - I also use local quantized models (I swap between ministral-3 and Qwen3) as "the brains" (offload to GPT-5.2 inc. web search for "complex" tasks or those requiring the web). In the backend I use Kokoro and generate wav bytes that I send to the frontend. The frontend is just a simple HTML page with a textbox and a button, invoking a `fetch()`. I type, and it responds back in audio. The round-trip time is <1 second for me, unless it needs to call OpenAI API for "complex" tasks. I am yet to integrate STT as well and then the cycle is complete. That's the stack, and not slow at all, but it depends on your HW.
- machiaweliczny 9mo agoBTW does anyone know of good assistant voice stack that's Open Source? I used https://github.com/ricky0123/vad https://github.com/ricky0123/vad for voice activation -> works good, then just using Web Speech API as that's the fastest and then commercial TTS for speed as couldn't find good one.
- jokethrowaway 9mo agoSorry but the quality is too bad. I'm sure it has its uses, but for anything practical I think Vibe Voice is the only real OSS cloning option. F2/E5 are also very good but has plenty of bad runs, you need to keep re-rolling.
- jokethrowaway 9mo agoI'm sure it has its uses, but for anything with a higher requirement for quality, I think Vibe Voice is the only real OSS cloning option. F2/E5 are also very good but have plenty of bad runs, you need to keep re-rolling until you get good outputs.
- bcrl 9mo agoWhat measures are being taken to ensure that this model isn't used to lower the cost of fraudsters committing grandparent scams by mimicking the voices of grandchildren?
- burnt-resistor 9mo agoNone, obviously, and it's barking up the wrong tree. The genie is already out of the bottle as there are zillions of similar free services and software that do the same thing, and there's no quick-fix panacea technological solutions to social and legal problems. Legislation in every locality need to create extremely harsh penalties for impersonating other people, and elders need to be educated to ask questions of their family members that only the real people would know the answers to.
- bcrl 9mo agoAh yes, the "things are bad; we shouldn't try to fix them" argument. That isn't a philosophy which I subscribe to. People should very much consider the ethical implications of releasing software they created to the general public.