8 ms·
Show HN: Voice Clones for Creators
- _josh_meyer_ 4y agoVoice Clones for: -> Video Games https://coqui.ai/video-games https://coqui.ai/video-games -> Dubbing https://coqui.ai/dubbing https://coqui.ai/dubbing -> Post-production https://coqui.ai/post-production https://coqui.ai/post-production Free access
- pointlessone 4y agoThe video games one is funny. The hero text is "Emotion on Call" and the video title is "hearing is Believing". The video is narrated by the most passionless voice imaginable. It barely has question inflections.
- jerf 4y agoI've wondered if this tech could be useful for video games... but it's going to take more than just recording one "voice" sample and typing some text. At the very least you're going to need some formal notifications of tone, because there's a world of difference between "[sarcastic] Yeah, really." and "[emphatic] Yeah, really." and at times its a super-AI-complete problem to know exactly what the author intended without annotations. (That is, even humans would get it wrong sometimes.) Unfortunately for the company, that video is indeed a great demo of the fact you can't just shove their technology at some text and "solve" the problem. However, for the first time, I feel like a half-decent solution to this problem is on the horizon, rather than absurdly far away.
- hhw3h 4y agoUsing the Brave browser, the examples don't work on your home page.
- reubenmorais 4y agoThanks for reporting, will take a look.
- jnurmine 4y agoWell, I'd like the voices of David Attenborough and Christopher Plummer for audio books please...
- meremortals 4y agohttps://github.com/BenAAndrew/Voice-Cloning-App https://github.com/BenAAndrew/Voice-Cloning-App
- meremortals 4y agoAlso, this timely announcement that Spotify has acquired Sonantic: https://techcrunch.com/2022/06/13/spotify-is-acquiring-sonantic-the-ai-voice-platform-used-to-simulate-val-kilmers-voice-in-top-gun-maverick/ https://techcrunch.com/2022/06/13/spotify-is-acquiring-sonan...
- 2Gkashmiri 4y agoDon't forget Morgan freeman
- jnurmine 4y agoIndeed! I'd like to add Lance Henriksen, too.
- inetsee 4y agoI'd like to be able to generate a voice like HAL, from "2001, A Space Odyssey".
- _josh_meyer_ 4y agowhy don't you try :D ? Make a clone where you're talking like Hal... see how it does ;)
- robbomacrae 4y ago15.ai has it
- 4y ago
- solardev 4y agoThat's pretty cool! Easy and fast to use, with reasonably good speaker reproduction. Right now the free version only allows a sentence or two of output. Are you moving towards a paid plan with longer outputs? And the training right now is just a paragraph of text. Does it get better after additional training?
- _josh_meyer_ 4y agoFor this zero-shot approach, there's actually no "fine-tuning", and from what we can tell ~30 seconds is optimal (the text isn't actually even used... you can say anything!) For longer outputs and legit fine-tuning, that could become an offering :D
- windsignaling 4y agoIs it possible to do the same thing (zero-shot on an arbitrary input voice) using one of the models on your Github?
- geuis 4y agoOn mobile: The page is too wide. The demos don't play.
- _josh_meyer_ 4y agothanks for flagging! what OS are you on? iOS seems working OK atm
- deleted 4y ago[deleted]
- probably_wrong 4y agoUnder Firefox on Android the demos play, but I actually need to click left of the play button for them to start. Looks like the radius of the time slider is too wide and overlaps the play button area. Perhaps that's what the parent comment is seeing?
- _josh_meyer_ 4y agothanks for flagging!
- puranjay 4y agohad the same problem on Safari iOS. Worked only after I accepted the cookie consent form.
- _josh_meyer_ 4y agowalk-though video --> https://www.youtube.com/watch?v=ri1U-bJ-6vc&t=4s https://www.youtube.com/watch?v=ri1U-bJ-6vc&t=4s
- leetrout 4y agoYou are really missing an opportunity by not including the actual audio recorded and the sample played back.
- _josh_meyer_ 4y agothe voices are demoed for [videogames](https://www.youtube.com/watch?v=x8tEdwll_CY&t=4s https://www.youtube.com/watch?v=x8tEdwll_CY&t=4s), [dubbing](https://www.youtube.com/watch?v=TBjdUY3_ccQ&t=1s https://www.youtube.com/watch?v=TBjdUY3_ccQ&t=1s), and [post-production](https://www.youtube.com/watch?v=ykoBFrb9itY https://www.youtube.com/watch?v=ykoBFrb9itY)... but I see your point - thanks!
- _josh_meyer_ 4y agovideo game demo --> https://www.youtube.com/watch?v=x8tEdwll_CY https://www.youtube.com/watch?v=x8tEdwll_CY
- ChicagoBoy11 4y agoI've tried a few of these over the past few years. This is by far the most impressed I've been by the generated result. I'm actually an educator and have used these services in the past with students as an entry point into a discussion about technology, AI, and media ethics. Is there anyone I can reach out to at Coqui for a possible collaboration?
- _josh_meyer_ 4y agoof course! we've got an open ethics discussion going here: https://github.com/coqui-ai/TTS/discussions/1036 https://github.com/coqui-ai/TTS/discussions/1036 and our chatrooms are probably the best/fastest way to get ahold of us: gitter.im/coqui-ai/TTS
- ratww 4y agoSame here. Not only the result is the best so far and instantly useable, the website is great to use and the fastest. The voices do lose a bit of character and accent, though. On the other hand that makes my own voice much more useable. That's probably for the better, and good for video narration, audiobooks and marketing. But I'm curious about the Video Game offering [1], where it says "You control the enunciation and emotion; the pitch and prosody; rate, duration, and contour". Is that planned for the future? [1] https://coqui.ai/video-games https://coqui.ai/video-games
- _josh_meyer_ 4y agomore fine-tuned control over enunciation/emotion and the like is very much in the works... stay tuned :)
- fxtentacle 4y agoGreat to see other Germans working on ASR and TTS :) I'm looking forward to having an offline and private voice assistant one day... In case you're not a member yet, take a look at LEAM.AI for potential partners and/or government funding.
- OtmaneBenazzou 4y ago"JSON.parse: unexpected character at line 2 column 1 of the JSON data" on firefox(bad request to graphql)
- leetrout 4y agoYea, 400 error on their graphql endpoint
- RileyJames 4y agoRemove step one from the video. Sign up is assumed, and it’s boring, and I dont need to be told how to put my email in the box that says email. Or organisation name in the… The point of the video is to show me the value. Show me what I can create. Within the first 10 seconds I should say “wow, I want this”. Start with a result, and then go through the key! Portions of the process, if you feel that’s the value. There’s no cloned voice in the whole video…
- mod 4y agoThere's three voice demos on the front page.
- TheTaytay 4y agoYes, but I agree wholeheartedly with the recommendation that you are responding to. I clicked on the first big “play” button I saw, thinking it would be a quick demo. I then started skipping through the video to find the meat. In particular, I expected to hear someone’s cloned voice and say, “wow”. When I realized it was primarily some stock footage and some silent recording of the UI, I moved on, and THEN realized that there were demo voices.
- _josh_meyer_ 4y agoyours is a solid point -- would be nice to have that "wow factor" right off-the-bat
- WhitneyLand 4y agoThey are not even visible on the first visible page on a mobile device.
- WhitneyLand 4y agoLearn from this….every second you have someone’s attention is precious.
- yellowapple 4y agoI would just scrap the video entirely.
- kevmo314 4y agoIs there an API for this? I actually have a programmatic use case I'd love to look into using this for :)
- _josh_meyer_ 4y agocurrently it's only user-facing, but still early days:)
- mattlondon 4y agoAre there are citations/verification for the "used by" logos? Basically they have every FAANG there, they're claiming that these companies use this, yet certainly at least for Google they have their own very active research and products going on in this area (DeepMind, wavenet etc) so it feels fairly implausible to me that they'd not be using their own tech that they are actively researching themselves, and instead use this? Feels a bit fraudulent to splash these names and logos around without any proof?
- 0des 4y agoI think sometimes folks put these on there for things like, an issue posted on github by a user whos email has a certain domain, or other very flimsy acknowledgements.
- wlesieutre 4y agoMaybe overly cynical, but I usually assume it's "someone with an email address from that company signed up for a free account because they wanted to spend 30 seconds trying out the service"
- 0des 4y agoI dont want us to be right, but damn. It could be what is happening in a lot of cases.
- patmcc 4y agoLooks like they've got some decently popular open source repos - https://github.com/coqui-ai/TTS https://github.com/coqui-ai/TTS for example - so I wouldn't be surprised at all if Google and others used that, even if only for their own research and comparison. That's my charitable interpretation anyway.
- stavros 4y agoThis is great! I see you've open-sourced some stuff, is there something I can use to train and generate these voices on my computer? I'd love a short tutorial somewhere of "you have an MP3 of your voice, here's how to generate more".
- _josh_meyer_ 4y agodelicious documentation for your reading pleasure:) https://tts.readthedocs.io https://tts.readthedocs.io
- stavros 4y agoExcellent, thank you!
- iamflimflam1 4y agoSeems to have turned me into an American... I would definitely be willing to use a service like this. I produce a lot of YouTube content and recording and editing good quality audio takes a lot of effort.
- nmstoker 4y agoGreat to see how things have progressed Josh. My sample was good - not likely to make anyone think it was me (it gave me a distinctly American twist), but it's an impressive zero shot result. Btw the website looks slick too - works well for me on Android 12 / Chrome
- _josh_meyer_ 4y agoThanks, Neil :D The current model definitely does an "implicit" accent conversion to American English, but in the near future we'll have something that keeps accents on separate sides of the pond, so to speak :)
- mynegation 4y agoI get an error at /voice/create: Unexpected token < in JSON at position 1 Google Chrome 102.0.5005.61 on MacOS Big Sur
- _josh_meyer_ 4y agostill getting the error? I can't replicate on MacOS Monterey
- andybak 4y agome too on an iPad using Firefox
- bobkazamakis 4y agoHasn't been linked yet so: https://github.com/coqui-ai/TTS https://github.com/coqui-ai/TTS Derived from Mozilla's TTS project. Big fan.
- _josh_meyer_ 4y agothanks for linking -- the core project is indeed open source! the founding team all worked at Mozilla on TTS and DeepSpeech, then we spun out to create Coqui <3
- _josh_meyer_ 4y agoProduct Hunt --> https://www.producthunt.com/posts/coqui https://www.producthunt.com/posts/coqui
- BasilPH 4y agoWhen Descript[^0] came out with automatic overdubbing, it was absolute magic and it stayed magic for a long time. Happy to see that there is now an open-source alternative, although I'm not sure how easy their models are to use. Does anybody know how they prevent you from creating a voice that isn't yours? I expect the text they make you read to be random, but if you have a large enough corpus of the person you want to imitate you could just piece it together. [^0]: https://www.descript.com https://www.descript.com
- dang 4y agoThis looks like a repeat from 6 months ago: Show HN: Clone your voice and speak a foreign language - https://news.ycombinator.com/item?id=29786132 https://news.ycombinator.com/item?id=29786132 - Jan 2022 (112 comments) Usually the cutoff for reposts is a year (see https://news.ycombinator.com/newsfaq.html https://news.ycombinator.com/newsfaq.html), unless the project has had a major update since then (see https://news.ycombinator.com/showhn.html https://news.ycombinator.com/showhn.html). In the latter case, the post should come with an explanation of what's new since the previous thread.
- _josh_meyer_ 4y agosame company, yes, but major updates since our demo: 1) Quality: the underlying model is much, much better, and so is the quality of the Voice Clone. 2) Productization: previously we just had a stand-alone demo, now we've launched the product (user accounts, multiple voices, etc.)
- dang 4y agoOk, that's probably fine.