Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
sammyyyyyyy
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
1.
▲
by
sammyyyyyyy
8mo ago
Today, I'm releasing a new version of my side project: SoproTTS A 135M parameter TTS model trained for ~$100 on 1 GPU, running ~20× real-time on a base MacBook M3 CPU. v1.5 highlights (on CPU): • 250 ms TTFA streaming latency • 0.05 RT
2.
▲
by
sammyyyyyyy
8mo ago
Nice
3.
▲
I trained a 135M TTS model for ~$100, runs 20× real-time on CPU
(huggingface.co)
4 points
by
sammyyyyyyy
8mo ago
|
2 comments
4.
▲
by
sammyyyyyyy
8mo ago
Today, I'm releasing a new version of my side project: SoproTTS A 135M parameter TTS model trained for ~$100 on 1 GPU, running ~20× real-time on a base MacBook M3 CPU. v1.5 highlights (on CPU): • 250 ms TTFA streaming latency • 0.05 RT
5.
▲
Sopro v1.5: A 135M TTS model trained for ~$100, runs 20× real-time on CPU
(huggingface.co)
5 points
by
sammyyyyyyy
8mo ago
|
1 comments
6.
▲
by
sammyyyyyyy
8mo ago
Today, I'm releasing a new version of my side project: SoproTTS A 135M parameter TTS model trained for ~$100 on 1 GPU, running ~20× real-time on a base MacBook M3 CPU. v1.5 highlights (on CPU): • 250 ms TTFA streaming latency • 0.05 RT
7.
▲
by
sammyyyyyyy
9mo ago
You should try it! I wouldn’t say it’s the best, far from that. But also wouldn’t say it’s terrible. If you have a 5090, then yes, you can run much more powerful models in real time. Chatterbox is a great model though
8.
▲
by
sammyyyyyyy
9mo ago
Also, I didn’t want to use known voices as the example, so I ended up using generic ones from the datasets
9.
▲
by
sammyyyyyyy
9mo ago
I should have posted the reference audio used with the examples. Honestly it doesn’t sound so different from them. Voice cloning can be from a cartoon too, doesn’t have to be from a human being
10.
▲
by
sammyyyyyyy
9mo ago
I didn’t specially cherry pick those examples. You can try it anyway for yourself. But thanks for the feedback anyway
11.
▲
by
sammyyyyyyy
9mo ago
As I said, some reference voices can lead to bad voice quality. But if it sounds that bad, it’s probably not it. Would love to dig into it if you want
12.
▲
by
sammyyyyyyy
9mo ago
Yeah sure. The training was about ~250 dollars, which is quite low by today’s standards. And I spent a bit more on ablations and research
13.
▲
by
sammyyyyyyy
9mo ago
No, it doesn’t.
14.
▲
by
sammyyyyyyy
9mo ago
Yes, you are right. However, there are many upsides to this kind of technology. For example, it can restore the voices of people who were affected by numerous diseases
15.
▲
by
sammyyyyyyy
9mo ago
Obrigado! Quando (e se fizeres isso) manda pm!
16.
▲
by
sammyyyyyyy
9mo ago
Yeah, we are not quite there, but I’m sure we are not far either
17.
▲
by
sammyyyyyyy
9mo ago
This is my side “hobby”. And compute is quite expensive. But if the community’s responsive is good, I will definitely think about it! Btw, chatterbox is a great model and inspiration
18.
▲
by
sammyyyyyyy
9mo ago
Cool! Yeah the voice quality really depends on the reference audio. Also mess with the parameters. All the feedback is welcome
19.
▲
by
sammyyyyyyy
9mo ago
Thanks! Yeah I kinda postponed publishing it until it was a bit better, but as a perfectionist, it would have never been published
20.
▲
Sopro TTS: A 169M model with zero-shot voice cloning that runs on the CPU
(github.com)
360 points
by
sammyyyyyyy
9mo ago
|
123 comments