4 ms·
Anyone can clone a voice with access to a couple of seconds of it.
by cal85 2y ago
Anyone can clone a voice with access to a couple of seconds of it.
- throw101010 2y agoA couple seconds? Which project does this?
- InvertedRhodium 2y agoI’ve had success with elevenlabs.
- tazu 2y agoThere are many models, XTTS [1] is a good one. [1]: https://coqui.ai/blog/tts/open_xtts https://coqui.ai/blog/tts/open_xtts
- defamation 2y agoevery single voice cloning project I'm sure you could have used google https://elevenlabs.io/app/voice-lab https://elevenlabs.io/app/voice-lab https://app.resemble.ai/users/sign_in https://app.resemble.ai/users/sign_in https://github.com/neonbjb/tortoise-tts https://github.com/neonbjb/tortoise-tts https://coqui.ai/blog/tts/open_xtts https://coqui.ai/blog/tts/open_xtts it was even possible 5 years ago https://github.com/CorentinJ/Real-Time-Voice-Cloning https://github.com/CorentinJ/Real-Time-Voice-Cloning
- xyst 2y agoMission Impossible 3 was only the proof of concept
- baobabKoodaa 2y agoNope. If you actually tried those, you would quickly find out they don't work. It's actually really hard to clone a voice from a few seconds sample.
- knowaveragejoe 2y agoA few seconds, yeah. I've seen fairly convincing reproductions from 30 seconds of reading text though.
- deleted 2y ago[deleted]
- jazzyjackson 2y agohow many do you think "a couple" is?
- grugagag 2y agoCloning voice signature or timbre may need a bit more for a good quality. Then there are idiosyncracies in one’s voice. In addition to that, there are tiny verbal tics, expressions, cadence, feel, and some more to be able to say you have properly cloned someone’s voice. The two second sample is like a shallow clone of sorts and is indeed vector space.
- BoredPositron 2y agoThe last sentence is hilarious. What do you think "properly" cloned voices are? Not every model is few shot and not every model relies on their training set for paralanguage anymore. Easiest way to try it out is properly the pro voice cloning from elevenlabs.
- Izkata 2y ago"Expressions" in particular is about choice of words. A few seconds definitely isn't enough to duplicate that.
- BoredPositron 2y agoThe choice of words is usually yours for a tts model.
- deleted 2y ago[deleted]
- mitthrowaway2 2y agoWe'll have to gradually get used to the notion that a person's voice, like so many other things we once thought of as intimately personal, is just a coordinate in a high-dimensional vector space.
- mjburgess 2y agoThose are not contradictions.
- bee_rider 2y agoBut it is unintuitive to those of us who are from the past.
- trescenzi 2y agoOne of the common bad takes is that since personal things can be distilled to “just math” they are no longer personal or valuable in the sort of way we value personal things. Increasing our understanding of how the world works shouldn’t devalue the world. To put it another way, people having souls is not the only reason to treat people like people. Or as Dr Seuss says: A person’s a person no matter how small.
- bee_rider 2y agoBut I do think there are some psychological hurdles we’ll have to overcome, as it becomes possible to mechanically copy individual personal aspects. A person’s a person no matter how small, but we aren’t all used to thinking of ourselves as very small.
- visarga 2y ago> they are no longer personal or valuable in the sort of way we value personal things There were people with your face or voice before, you just didn't care. AI continues a trend: the internet has always been post-scarcity, free copying and offering an amazing array of choices. We are acting as if it's a new thing, but it's been here for 25 years.
- minimaxir 2y agoFun fact: many of the TTS API providers (Google, Azure) now have voice cloning capability. There's a reason it has gone under the radar.