12 ms·
Show HN: Clone your voice and speak a foreign language
- carbonx 5y agoI just tried it and it sounds nothing at all like me. shrug
- kubb 5y agoSame here. It even mispronounced the basic french words, and inserted some background music similar to what you can hear on the CDs with exercises that come with those "foreign language for beginners" textbooks.
- alonmln 5y agoCool, it's impressive how much can it do with a short sample, although this seems like an easy way for end users to deep fake their friends / enemies saying something.
- kdavis 5y agoCurrently we’re looking at possible solutions, see for example here[1]. If you have suggestions, feel free to chime in! In the demo we specifically disallowed bulk uploads to hinder such abuses. [1] https://github.com/coqui-ai/TTS/discussions/1036 https://github.com/coqui-ai/TTS/discussions/1036
- Philip-J-Fry 5y agoMaybe the solution is to have a randomly generated paragraph of text to read which expires in short amount of time. So you can't predict it and you don't have enough time to splice together a fake reading from something else.
- Gigachad 5y agoThe problem with any anti abuse measure is someone can create another project which does not have any of this. There are a handful of projects which can do pretty good voice synthesis right now. It would be about as easy as getting a consensus for all photo editing tools to place a watermark on the image to prevent abuse.
- tiborsaas 5y agoI tested it with your comment: https://sndup.net/mghy/ https://sndup.net/mghy/ :) It's also a new possibility to somewhat personalize the text to speech engines. The above example is not really close to my voice.
- sxv 5y agoMy 26 second training input perhaps wasn't enough. The result sounded like someone else. Is the result some kind of merger of my voice and a native speaker's?
- reubenmorais 5y agoSimilarity depends on many factors: recording quality, which language you're synthesizing in (models trained on more speakers do better), and diversity of prosody in your recording. Try recording for a bit longer and "acting out" a bit in your tone, that tends to give me interesting results :)
- jeroenhd 5y agoInteresting. I like the addition of music to make sure it's not just a raw voice sample. The output I get seems to be a mix of a native speaker and my voice, because my (thick) accent is being filtered out. I suppose that if I ever take proper English pronunciation classes, I now know what to strive for.
- SwiftyBug 5y agoI speak Brazilian Portuguese natively. I chose to record my voice saying a specific sentence and to "translate" it to Brazilian Portuguese using the exact same sentence. I was very pleased to find out that I became a Mineiro from the countryside, one of the coolest accents in Brazil!
- reubenmorais 5y agoThe Brazilian Portuguese model is a bit of an extreme showcase (and thus really cool!), as it was trained on a single speaker (entirely recorded by the main author of the paper, Edresson Casanova, who's Brazilian). The fact that it can do multi-lingual voice cloning at all in that case is already surprising. You can find more details in the project page [0] and paper [1]. And here's the corpus. [2] [0] https://edresson.github.io/YourTTS/ https://edresson.github.io/YourTTS/ [1] https://arxiv.org/abs/2112.02418 https://arxiv.org/abs/2112.02418 [2] https://edresson.github.io/TTS-Portuguese-Corpus/ https://edresson.github.io/TTS-Portuguese-Corpus/
- actually_a_dog 5y agoYou spoke Portuguese into it and it just changed your accent? That's kinda cool.
- themodelplumber 5y agoI can speak just enough to know how I sound, and was surprised to hear that accent too :D
- bernardom 5y agoThis is very cool. I recorded myself in Portuguese -> Portuguese and got the same result. I also did Pt -> En and sounded like... me speaking English, though with some artifacts. VERY cool.
- garfieldnate 5y agoThere was a thread a while back about the need for "accent correction," meaning that native speakers with one accent could more easily consume content in the same accent. It looks like the technology exists now! This is worth money. If you find an accent that many people really dislike, the odds are that's it's also very difficult for people to understand that accent (until they are accustomed to it).
- acqbu 5y agoGold!
- pcarolan 5y agoThis is incredibly impressive and does a great job of capturing my voice. Well done!
- akeck 5y agoIs it supposed to translate or just read with the target accent? For me, it's only reading the English input text with the target accent.
- reubenmorais 5y agoIt doesn't translate the text, you have to put in text in the target language. But you can record audio speaking in any language you want.
- momolo 5y agois the model available?
- _josh_meyer_ 5y agoDemo: https://coqui.ai https://coqui.ai Code: https://github.com/coqui-ai/tts https://github.com/coqui-ai/tts Blogpost: https://coqui.ai/blog/tts/yourtts-zero-shot-text-synthesis-low-resource-languages https://coqui.ai/blog/tts/yourtts-zero-shot-text-synthesis-l... Paper: https://arxiv.org/abs/2112.02418 https://arxiv.org/abs/2112.02418
- echelon 5y agoThis is so cool! Thank you! How do y'all intend to profit (succeed as a startup) if you're releasing so much publicly? I'd love to see you guys succeed. Really great to see where some of the Mozilla TTS folks wound up, too.
- bagels 5y agoIs there a static demo that I don't have to provide my own voice for?
- deleted 5y ago[deleted]
- kdavis 5y agoWe did not provide such a demo in part to hinder nefarious uses of the technology.
- crumpled 5y agoHonestly, how much of a hinderance is that? A person could just supply a recording of another person, couldn't they?
- trump_tts 5y agoYes, it's pretty easy to play back a video as the input text and then generate a reasonable fascimile. Here's Trump: https://sndup.net/zkbg/ https://sndup.net/zkbg/
- gruez 5y agobut it looks like you provide the source code here? https://github.com/coqui-ai/tts https://github.com/coqui-ai/tts. How much of a hindrance are you hoping to add?
- BoorishBears 5y agoDiscourages low-hanging, hit-and-run usage that's likely to get their site shut down. If someone wants to fake a statement there are already 100 ways to do it. Not making their servers the ones doing the deed puts a meaningful barrier in place for more casual misuse. And for serious cases like impersonation on a large scale, the resources are there to likely do better than this instant feedback model can.
- 5y ago
- deleted 5y ago[deleted]
- ceva 5y agoit says enter your text here ..
- kdavis 5y agoYou're free to enter any input sentence you want in the text box. The input sentence generally should be in the language you selected from the dropdown. For example, if the dropdown has "French" selected you could enter the text "Allons enfants de la Patrie, Le jour de gloire est arrivé!" Clicking "Submit" then generates a TTS reading of the sentence you input in the language selected from the dropdown. For fun you can mix and match. In other words, select a language from the drop down and enter text in the text box not in the language selected from the dropdown. (For example, the dropdown could have "French" selected and the sentence could be "O say can you see, by the dawn's early light". This gives interesting results, it sounds as if a native French speaker is speaking English.)
- wombatmobile 5y agoAwesome! How do I embed this?
- IanCal 5y agoVery interesting! Is the music an intentional blended track or an artifact of generation?
- _josh_meyer_ 5y agovery much intentional. Background music makes misuse/abuse less likely (both intentional and unintentional) Read more here about in our open discussion: https://github.com/coqui-ai/TTS/discussions/1036 https://github.com/coqui-ai/TTS/discussions/1036
- Gigachad 5y agoI appreciate the effort here, but it almost feels like this is hopeless as it seems so many groups are able to build voice synthesis right now that the tech has fallen in to the common persons hand and some of them won't make any effort to stop abuse. Maybe if we can get watermarked stuff out first and the average person gets up to speed with what tech can do, we can all adjust our expectations before the real wave of abuse hits.
- marcan_42 5y agoYou can probably run the output through Spleeter[1] and get rid of the background music very easily. Just throw more AI at the problem... It's very hard to curb intentional misuse. [1] https://github.com/deezer/spleeter https://github.com/deezer/spleeter
- winter_squirrel 5y ago
- bengalister 5y agoI am French and I did try it, recording my voice in English (I have a thick French accent to English speaking ears, ok for French ones). And the result back in French was kind of good even it did sound almost like me with a slight American English accent.
- tambourine_man 5y agoI'm curious but a bit afraid to test it out. The idea of having a model of my voice out there that can say whatever is written in a text box is scary.
- colecut 5y agoShould you never speak to be sure you aren't recorded?
- imapeopleperson 5y agoThis is a perfect example of when the law shouldn’t be so far behind the tech.
- creato 5y agoExactly which part of this do you think should be illegal?
- ehnto 5y agoIt doesn't have to be illegal but I think some defensive regulation here is smart. Things people are concerned about may already be illegal. Imitation, identity theft, slander and so on. Think about the new layer it adds to domestic disputes and criminal investigations. Perhaps a solution is a sound fingerprint requirement for voice imitation software so that it's easily identifiable in court if it's an imitation voice. It's somewhat of a new frontier, imagine during a divorce proceeding your ex-partner fabricates voice recordings of you threatening the kids so you don't get custody, how do you protect yourself against that, how to you prove that's what happened? Soon enough it'll just be an app on their phone that they use to record your voice during a discussion, then later spits out a sound file of you saying whatever they want you to say. That's clearly a socially dangerous tool.
- shanlalit 5y agoTomorrow a paid tool or a costly hidden company will allow anyone to get statement in your voice (based on sample). How you are going to proof, that it is not you? Fake calls to your relatives in your voice or even fake video with your face and voice asking for money! or illegal activities. Few years later a company will come and say we can detect if it's fake or not pay $10,000 for solution, or get ready to be in prison. Oops! legal system doesn't accept this as a proof, now what? Welcome to the prison. Both companies are making money, and you are paying by money and your life.
- grogenaut 5y agoThis is one of those ideas that seems obvious when you hear it and also I'm pissed I didn't think about it. It also seems like a key component to a universal translator. This + VTT + a phone sounds like it'd put UN translators out of business (:) yeah I know, nuance probbably matters there).
- forgotmyoldacc 5y agoThis is called end-to-end speech translation, and has been around since 2017. Here's an article from 2019: https://www.technologyreview.com/2019/05/20/103054/google-ai-language-translation/ https://www.technologyreview.com/2019/05/20/103054/google-ai...
- Ice_cream_suit 5y agoGreat opportunity for criminals and state actors to take identity theft to the next level.
- graderjs 5y agoScoff...As if they didn't have this for 10 years already.
- Ice_cream_suit 5y agoDid they have your voice model so easily available, hosted on a poorly secured servers , until you decided to try out this new free toy ?
- martopix 5y agoThere is a video of me talking for an hour straight on youtube, for example.
- trompetenaccoun 5y agoProfessional criminals surely do have something like it already: https://www.forbes.com/sites/thomasbrewster/2021/10/14/huge-bank-fraud-uses-deep-fake-voice-tech-to-steal-millions/?sh=86665f675591 https://www.forbes.com/sites/thomasbrewster/2021/10/14/huge-... There is no reason to blame the creators, this is going to go mainstream one way or another.
- raphman 5y agoNice! The first few seconds sound a lot like me. Afterwards, not so much.
- fnord77 5y ago"At Schwab, my voice is my password" [1] [1] https://www.schwab.com/voice-id https://www.schwab.com/voice-id
- echelon 5y agoThat's got to be among the worst ideas I've ever seen. https://fakeyou.com/tts/result/TR:eyfam30e255zxy69vn6a7z7yn9vqb https://fakeyou.com/tts/result/TR:eyfam30e255zxy69vn6a7z7yn9...
- zuhayeer 5y agoPretty neat! Was poking around on the site, and under the hood the interface to upload and render the audio is powered by Gradio: https://gradio.app/ https://gradio.app/
- gambiting 5y agoAs someone who actually speaks two languages - gave it a voice sample in Polish, then used it to synthesize the voice in English - sounds absolutely nothing like me. Meh.
- Bichote 5y agoVery cool project but just a little nitpick; maybe use a picture of a real Coqui frog frog on your site? https://en.m.wikipedia.org/wiki/Coqu%C3%AD https://en.m.wikipedia.org/wiki/Coqu%C3%AD
- patrec 5y agoHave you investigated whether this is useful for language learning? Presumably it ought to be easier to try to emulate (and compare and contrast) speech in "your own voice" (with a native accent) than someone else's. Another useful feature to this end might be to emulate how your voice sounds to you (rather than other people); not sure how difficult that is.
- jstsch 5y agoIndeed! I tried it with some French and was impressed. After recording in English and synthesizing a short sentence I tried to record and speak using the same intonation/speed as the generated French audio. It matches almost perfectly. Except of course for the bg music I don’t think anyone could discern which one was real and which one was fake. It didn’t work for all sentences, and there were some obvious glitches, but for the pieces where it did it was quite freaky. Also, hearing the French sentence in my own voice made it quite easy to pronounce it correctly. When I try this using for example the Google Translate TTS it’s much much harder.
- reaperducer 5y agoOff topic, but this reminded me: What ever happened to that thing that Google demoed where its robots would call restaurants and make reservations for you? Did that ever find its way into Android, or another product?
- lern_too_spel 5y agoI don't know who makes reservations these days, but it's available almost everywhere in the US now from Google Assistant. https://support.google.com/business/answer/7690269#zippy=%2Cdevices-supported%2Cavailability-in-the-united-states https://support.google.com/business/answer/7690269#zippy=%2C...
- netman21 5y agoTried it. Just a voice to text of French guy talking. Definitely not my voice.
- yosito 5y agoVery cool! If I were looking for a side project, I'd extend this, add a DeepL integration for automating translations, add some voice models for other languages/people and wrap it as a mobile app where people could pay to unlock the voice models.
- xiii1408 5y agoPretty interesting! I tried this both English -> French and French -> English. English -> French seemed to work best, with the AI output have a very similar timbre to my real voice. Not hyperrealistic for me, but decent enough given I gave it a ~20s sample. French -> English was less good in terms of the timbre and pitch of the voice---way higher than my real voice. It did have a bit of a Canadian accent, though, which is funny because I speak French with a Quebec accent. Maybe that's what I would sound like if I had a Canadian accent in English?
- mod50ack 5y agoFunnily, I (native American English speaker who learned French in QC, and whose accent in French indicates this) tried it both ways. I think the accent is basically built in both ways, which makes sense, although it would be more interesting if it based your accent in the output off the phonology in the input.
- thejosh 5y agoThis is great, my wife actually thought it was me speaking for a moment when she first came in!
- throw453221 5y agoThis will be great for foreign movies. While I still prefer subtitles, for those who watch with dubs, it’ll be amazing to hear the actor’s “real” voice.
- Rhinobird 5y agoMy question is, how long until we have automatically dubbed anime?
- var_cw 5y agoPretty soon. But won't you prefer subs over dubs for anime?
- matheist 5y agoTo help prevent malicious use, consider presenting the user with specific (randomly-generated) text to read aloud, and check (with speech to text) that they actually read that, instead of allowing them to say whatever they want. That will help ensure that this is only being used by the person visiting the web page. (That will only help with the hosted version, of course, not if you make the model code/weights available. I didn't generate this idea myself but also can't remember where I saw it. I think it was from someone offering a similar service.)
- FatDrunknStupid 5y ago
- abel_ 5y agoAn interesting reflection is how quickly research around TTS/STT has progressed. I remember reading [0] thinking we were a long ways away. And things will get way better with multi-task learning and multi-modal learning in the coming years (or months really). In fact, just a year after this post was written, CoquiAI started their open source projects [1]. [0] https://news.ycombinator.com/item?id=22869365 https://news.ycombinator.com/item?id=22869365 (https://thegradient.pub/towards-an-imagenet-moment-for-speech-to-text/ https://thegradient.pub/towards-an-imagenet-moment-for-speec...) [1] https://star-history.com/#coqui-ai/TTS&coqui-ai/STT https://star-history.com/#coqui-ai/TTS&coqui-ai/STT
- jijji 5y agowhat about a real time white english to ebonics/jive translator, when are we going to see this...
- Ninjinka 5y agoThis is amazing. I can't wait until this is used to dub TV shows, so we get the original actors' voices, especially for shows like Squid Game that had such terrible dubs.
- var_cw 5y agoDon't you think thats more of a translation problem rather than how well it was spoken?
- lostcolony 5y agoIt also misses that vocal inflection and timing is part of what makes a solid dub. Even if the translation was amazing, with all the subtleties of the language conferred somehow, you still have to get that right for it to be convincing. Otherwise you could end up with solid dialog such as Pride and Prejudice, as delivered by Tommy Wiseau or Christopher Walken or something ("Those. Who do not. ComPLAIN. Are never pitied.")
- perryizgr8 5y agoI spoke for about a minute in English, having no idea what is the ideal length for it to properly figure out my voice. The result sounded like someone else completely. There was also some strange music in the background, which made me think that it was playing back a recording of a real person speaking! A real person who's not me.
- mulholio 5y agoGoing to test this out on Hinge voice prompts and see what happens
- dutchbrit 5y agoWeird, my voice turned American...
- milkers 5y agoWhat kind of wizardy is this!? Congrats Coqui team!
- rambojazz 5y agoCan tools like this exist in offline mode only? Or do they require some high-power computers for the model?
- text2db 5y agoI recorded my voice in English and it converted into French, but I heard after converting it to french voice, a music was heard in the sound after my voice was played , what was it. Btw idea is really cool, its like how will you speak in same tone in other languages.
- garfieldnate 5y agoI put my voice in in English and asked for English back. The output mostly sounded the same, but the interesting thing was that it had some music playing in the background, like the ambient kind you might hear in a YT video while a narrator talks.