16 ms·
Lyrebird – An API to copy the voice of anyone
- celticninja 9y agoIs this enough to beat voice recognition software? If you thought fake news was bad before wait until these 'secret' recordings start getting released and reported on.
- _mhr_ 9y agoImagine when entire videos can be synthesized.
- halflings 9y agoAlready done: https://www.youtube.com/watch?v=ohmajJTcpNk https://www.youtube.com/watch?v=ohmajJTcpNk
- patrics123 9y agoI think its good if a tech like this is publicly available. It will be used by comedians and satire outlets and over time raise awareness about possible fakes - pics/video or it did not happen... Well, no problem anymore ;-)
- deleted 9y ago[deleted]
- brango 9y agoWow that's amazing. Combine the two and soon Hollywood stars will be redundant. Faceless session actors could just manipulate models of real people who've signed release forms with the vocal performed by similarly faceless vocal artists, or maybe even AI generated voices. Actors would lose their uniqueness and so end up being paid a pittance instead of being able to command the vast sums they can today. Another set of jobs soon to be made redundant by the rise of technology.
- throwaway29292 9y agoI wouldn't be comfortable watching a movie scene if I knew I was looking at computer-generated faces and voices.
- techdragon 9y agoAre you comfortable with Auto-Tune in music, not the t-pain / etc exaggerated style... the nearly universal application of Auto-Tune to recording and live performance to ensure a "consistent product", and "save on expensive studio time"? Because the market appears to have spoken on that one and it said "meh, I don't care" with an solid shrug of indifference. By the same logic, one can see artificially produced vocal performance combined with artificial overlaying of photorealistic 3d reproduction as a way to cost effectively maximise the performer and crew expenses, and ensure the consistency of a performance. The results may even be better than what they could have done with the real performer in the case of some attractive actors who are not very good at the acting part of being an actor/modern celebrity. That said, I'll definitely miss the days people had to actually be able to act, but then again I also miss the days people used to actually be able to play an instrument well and or sing well if they wanted to be a famous musician.
- kastnerkyle 9y agoJapan, as usual, is ahead of the game here! [0] [0] https://www.youtube.com/watch?v=pEaBqiLeCu0 https://www.youtube.com/watch?v=pEaBqiLeCu0
- TuringNYC 9y agoDid you feel that way when seeing Grand Moff Tarkin on "Rogue One"? https://www.wired.com/video/2017/02/how-rogue-one-recreated-grand-moff-tarkin/ https://www.wired.com/video/2017/02/how-rogue-one-recreated-...
- ClassyJacket 9y agoYes? I missed literally all his dialogue because he was so poorly animated I couldn't take my mind off it. So out of place and jarring.
- froindt 9y agoCan you imagine this for the next generation of cyber bullying? It could get super messy in high schools. Alice broke up with Bob. Bob grabs the YouTube videos from Alice and makes video and voice profiles. Bob then posts a video of Alice saying how breaking up was her biggest mistake and how she misses <list of every sexual thing you can think of> because Bob does it all best. That could end badly really easily.
- simcop2387 9y agoThat was one of my first thoughts too [1]. I doubt it currently will do so as it does sound fairly robotic. It does sound closer however than other demos I've heard so i think we're probably on the way. [1] https://youtu.be/-zVgWpVXb64?t=41 https://youtu.be/-zVgWpVXb64?t=41
- TeMPOraL 9y agoThrow some hard compression on it to get additional distortion, and I feel you could get away with claiming it was a recording done by a low-power bug.
- timlyo 9y agoI'm interested to see if it can fool Android's trusted voice.
- eps 9y agoImpressive. But also enabling the next gen of "Mom, I'm in Mexican jail. Quickly wire me $2,000 so I can get out." scams.
- d--b 9y agoAnd Black mirror's "would you like to speak with your dead husband again?"
- Ygg2 9y agoForget dead husbands, with this tech, it will be hard to trust anything a politician said. Basically, once they master adding this to video, ANYTHING could be construed against anyone. Want a video of a politician saying "Hitler was right" to cheering masses? Want a video about a president saying it's time to start Nuclear War One? You can make that.
- efaref 9y agoI can't wait to hear the new cassetteboy songs.
- la_oveja 9y agoIn the future: PGP signed speeches.
- TheOtherHobbes 9y agoYou joke, but I strongly suspect Authenticated Personal Trust is absolutely going to have become a thing.
- tobr 9y agoAnd any genuine recording can be plausibly denied.
- d00b 9y agohttps://www.youtube.com/watch?v=ohmajJTcpNk https://www.youtube.com/watch?v=ohmajJTcpNk hf
- yladiz 9y agoThis is pretty cool (although, I have no idea what other technologies exist for this kind of thing), but it's definitely not convincing enough to a human listener. This sounds like it might be convincing enough for some programs like "Hey, Siri" but it's not gonna convince your mom. You can listen to the samples on the page linked here and you can immediately tell that Obama and Trump don't sound quite human.
- patrics123 9y agoYet.
- brian-armstrong 9y agoText to speech is still pretty distinguishable as not-human, and that seems like an easier problem (only has to work for one specific voice, not an arbitrary voice). So just on the basis alone I wonder if this isn't still a ways out
- a12jun 9y agoInteresting thought: is it easier or more difficult to make a synthetic voice undistinguishable from a human one, compared to producing speech copying a real voice? To me, at least, the voices I heard in Lyrebird's demo actually sounded more 'real' than Microsoft Sam for example.
- yladiz 9y agoOf course, the voices produced by Lyrebird sound more "real" than Microsoft Sam, since I would define realness as sounding humanlike. However, I would strongly prefer Microsoft Sam over something generated by this algorithm for general use because this algorithm produces voices that are still in an uncanny valley, because it is almost human but not human enough, whereas Microsoft Sam is obviously not human.
- fudged71 9y agoI would argue that it depends on the sample rate. Over the telephone there are several TTS voices that are very convincing because the audio quality is lower.
- felipemesquita 9y agoThis site has a "demo" section featuring only Soundcloud clips. Uses to much the present tense "In a world first, Montreal-based startup Lyrebird today unveiled" and "Record 1 minute [...] and Lyrebird can [..]Use this key to generate anything" but has no actual product or beta version. Adobe had a much more impressive sneak peek of a similar product called VoCo: https://www.youtube.com/watch?v=I3l4XLZ59iw https://www.youtube.com/watch?v=I3l4XLZ59iw
- Gaelan 9y agoTo be fair, that demo could have been staged whereas we can be pretty darn sure those aren't Trump's actual words.
- felipemesquita 9y agoI agree. My main issue with them is that those clips could have been painstakingly produced using some very far from shipping software with lots of manual tinkering while the copy on the site mostly reads like there's a product out. The part about being far from shipping could also be true about adobe's software, but I think their presented result (assuming it is real) sounded better, and they were more honest about the stage the product is in.
- stevenh 9y agoLyrebird has real tech capable of doing what they claim, whereas Adobe's demo was completely fake. https://www.youtube.com/watch?v=I3l4XLZ59iw&t=2m34s https://www.youtube.com/watch?v=I3l4XLZ59iw&t=2m34s "Wife" sounds exactly the same in both places. All they did was copy the exact waveform from one point to another, like an automated cut and paste in Audacity. Nothing is being synthesized. https://www.youtube.com/watch?v=I3l4XLZ59iw&t=3m54s https://www.youtube.com/watch?v=I3l4XLZ59iw&t=3m54s The word "Jordan" is not being synthesized. The speaker was recorded saying "Jordan" beforehand for this insertion demo and they're trying to play it off as though it was synthesized on the fly. That's incredibly dishonest. This is a scripted performance and Jordan is feigning surprise. https://www.youtube.com/watch?v=I3l4XLZ59iw&t=4m40s https://www.youtube.com/watch?v=I3l4XLZ59iw&t=4m40s Again, the phrase "three times" here was prerecorded. This was a phony demonstration of a nonexistent product. Reporters parroted the claims and none questioned what they witnessed. Adobe falsely took credit and received endless free publicity for a breakthrough they had no hand in by staging this fake demo right on the heels of the genuine interest generated by Google WaveNet. They were hoping they'd have a real product ready before anyone else. If Adobe had a real product then they'd have proven it with a demo as alarming, undeniable, and straightforward as Lyrebird's. Instead they relied on aesthetics and flashy, polished, deceptive performance art with famous actors and de facto applause tracks to cover up the fact that they have nothing.
- cjlars 9y agoI was wondering when CG Sir David Attenborough would get here and start narrating my day to day.
- seanhandley 9y agoActually, when the sad day comes and he passes away I'd be very comforted to hear his voice on new nature documentaries. I just can't watch them unless he's narrating.
- mhandley 9y agoI agree. This leads to an interesting question: can the estate of a deceased person sell or license their voice rights for new future performances? I suspect the law has some catching up to do.
- anigbrowl 9y agoThis is not qualitatively different from existing situations, and will be only a minor legal wrinkle. Estates have been licensing the likeness of dead people for commercial purposes for a good while, now those likenesses are simply more sophisticated. I'll tell you what will get complicated, copyright holders complaining that their product was used as input for the training algorithm and demanding a slice of any profits because they made the famous individual more famous by casting them.
- barrystaes 9y agoI imagine Sir David Attenborough would basically just cite the intro of "The Gods Must Be Crazy" movie, casually explaining how humans technological progress is futile in regards of happiness. That text seems to holds up to present day.
- afinlayson 9y agoThis is how a lot of tech companies make proper text2speech, this was just done using the vast amount of audio that's out there for these people. Soon Trump will use this to state that things he's said are fake news. God help us all.
- Gaelan 9y agoThey claim they only needed one minute of sample audio. We really need to start requiring that all public announcements (news, press releases, etc) are digitally signed and put into the blockchain.
- backpropaganda 9y ago[deleted]
- IanCal 9y agoThey both sound right to me.
- amarant 9y agoIt's there any copyright protections for a person's voice? If not, David Attenborough and Morgan Freeman will be lead voice actors in my next game project
- Asooka 9y agoAFAIK not copyright per se in the traditional sense, but there is a "likeness right".
- pluma 9y agoThis. I would assume copying a voice pattern is treated no different from copying an appearance. If a 3D model is a very close approximation to a particular person's face and you use it in a game intentionally to benefit from that person's likeness, you can run into legal issues. And of course you really open yourself up for a lawsuit if you then attach the real person's name to it. EDIT: See the NPR Planet Money podcast linked elsewhere in the comments. Apparently there are fairly specific laws to protect someone's likeness -- including their voice -- thanks to the entertainment industry and Frank Sinatra.
- wiiittttt 9y agoLook at Crispin Glover's lawsuit for Back to the Future II as well.
- smnc 9y agoTom Waits sued for this reason multiple times [1]. I guess he had a strong case because in all three cases he was (or so he claimed) approached to do a commercial and the advertisers went with a Waits-like surrogate after he declined. https://en.wikipedia.org/wiki/Tom_Waits#Lawsuits https://en.wikipedia.org/wiki/Tom_Waits#Lawsuits
- praptak 9y agoIt's personality rights. You cannot copyright a naturally existing thing like a face or a voice.
- 9y ago
- scibolt 9y agoVoice Actors out of business! :D
- carlob 9y agoI wonder how dependent this is on language: can we make Trump speak Chinese using a one minute audio track of him speaking English?
- simcop2387 9y agoI'd imagine you might get close, but it'd probably work best if you can get the person to use all the phonemes that you want to reproduce. That said depending on how good it is, it might slur other phonemes together to approximate it, which would probably work to give it the accent that the speaker would likely have.
- Sunset 9y agoIt all sounded sort of slurry and muffled. Maybe if you're imitating a naturally slurred speaker it would be more effective.
- JustFinishedBSG 9y agoCooler : http://www.dtic.upf.edu/~mblaauw/IS2017_NPSS/ http://www.dtic.upf.edu/~mblaauw/IS2017_NPSS/ https://arxiv.org/abs/1704.03809 https://arxiv.org/abs/1704.03809
- backpropaganda 9y agoCanned voices? Robotic intonation? No thanks.
- JustFinishedBSG 9y agoDid you even listen to OP voices and compare them with for example the article Spanish generated voice?
- retox 9y agoVery impressive, but like someone else said there are definite robotic/synthetic moments. I wonder how easy it would be to combine all of the other projects linked in these comments, is there a common interface that could easily combine all of them? Probably not because of commercial concerns...
- kastnerkyle 9y agoThis model is quite cool, but also quite a bit different than what lyrebird.ai is doing. NPSS has a lot of extra information in the control inputs about pronunciation and timing (the part-of-phoneme timer feature) - this means that most of the "hard parts" (in my opinion) for naturalness are control inputs to NPSS/WaveNet style models, rather than variables the model must generate globally and consistently as in lyrebird. At generation time NPSS appears to generate each component autoregressively as well, but I am not clear on whether the demo samples do this or if they use "true" values for f0 at least - what forces the model to sing the exact same melody, if many melodies are possible given the underlying audio information? Also note that NPSS has some amount of post-processing, at least reverb and perhaps other common musical mixing - we don't really know how these samples are generated, and I have a hard time decyphering exactly what inputs are required, and what are generated from the paper alone. However, I really, really, really like NPSS - I just don't think the comparison you are making is valid here. These features (f0, duration, pronunciation) are some of the most difficult things to learn to model from datasets of speech and text directly, and I am not sure how they got the subset used (I think only f0 and pronunciation/phoneme) for this NPSS model. Giving creators fine-grained control of the performance (as in NPSS) is quite cool, and if these systems can get fast enough I think the possibilities are really exciting. The same things could likely be done with lyrebird as well - there is no real "tech reason" you couldn't add more conditional inputs, with finer grained information/control. The key part in my mind is deciding what amount of complexity to show to a user, and what amount to try and capture inside the model - some people may want to control (for example) duration and f0 directly for a performance, while others may want to just upload clips to an API and get reasonable results back, with less ability to control each sample (they can still curate themselves for the "best" samples). Lyrebird.ai is handling the latter case, while the former case would require quite a bit more intervention from the average user, almost becoming like an instrument ala the original voder [0]. However, you could potentially have both approaches as a kind of beginner/advanced mode, but advanced mode needs a user interface, and probably near-realtime feedback. I used to really strongly believe that the audio model was going to be the hard part of "neural" TTS (blame my background in DSP perhaps), but post-WaveNet the game has really changed a lot - conditional audio models are something we are starting to know how to do pretty well. The text pipeline of most TTS systems is still the craziest part in my mind, check out a "normal" feature extraction of 416 hand-specified features [1]! These extractions can be upwards of 1k features per timestep/frame, and generally require a lot of linguistic knowledge to specify for new languages. It seems (given Alex Graves' demo [2], char2wav [3], tacotron[4]) that we are making progress on learning this information directly from text, which in my mind is a key breakthrough for TTS in languages besides English, where lots of work on English pronunciation has been done already and is generally available. [0] https://www.youtube.com/watch?v=TsdOej_nC1M https://www.youtube.com/watch?v=TsdOej_nC1M [1] https://github.com/CSTR-Edinburgh/merlin/blob/master/misc/questions/questions-radio_dnn_416.hed https://github.com/CSTR-Edinburgh/merlin/blob/master/misc/qu... [2] https://www.youtube.com/watch?v=-yX1SYeDHbg&t=38m00s https://www.youtube.com/watch?v=-yX1SYeDHbg&t=38m00s [3] http://josesotelo.com/speechsynthesis/ http://josesotelo.com/speechsynthesis/ [4] https://google.github.io/tacotron/ https://google.github.io/tacotron/
- selbekk 9y agoScary.
- qeternity 9y agoWhile all of these vec2speech type models are impressive, I get the feeling that most of the comments didn't listen to any of the samples. It's still distinctly robotic sounding, probably has quite a bit of garbage output that needs to be filtered manually (as many of these nets often have) and is a far cry from fooling a human.
- cel1ne 9y agoIt doesn't sound too different from a voice coming over a walkie talkie or some kind of intercom. The problem might be that high frequencies, especially overtones, aren't properly constructed, but I'm certain that can be improved.
- StavrosK 9y agoThe main problem is that the algorithms don't yet know what to stress in a sentence. The problem is semantic, and not so much about the sound of the voice itself. You can synthesize someone's voice perfectly, but if it's stressing words incorrectly or not at all, it's not going to fool anyone. Then again, that's probably easier to work around by having humans annotate the sentences to be read.
- comex 9y ago> Then again, that's probably easier to work around by having humans annotate the sentences to be read. Or by starting with a recording of someone else reading the sentence. Then you get the research problem known as "voice conversion", which has been studied a fair amount, but mostly prior to the deep learning era - and mostly without the constraint of limited access to the target person's voice. (On the other hand, research often goes after 'hard' conversions like male-to-female, whereas if your goal is forgery, you can probably find someone with a similar voice to record the input.) Anyway, here's an interesting thing from 2016, a contest to produce the best voice conversion algorithm, with 17 entrants: http://vc-challenge.org/summary.html http://vc-challenge.org/summary.html
- truthexposer 9y ago
- backpropaganda 9y agoRelevant discussion from 17 hours ago: https://news.ycombinator.com/item?id=14177589 https://news.ycombinator.com/item?id=14177589
- hoodoof 9y agoIt feels like the future has arrived.
- ageofwant 9y agoOh yea. The Troll embedded deep in my soul giggles in glee. However, the day some shill tries to sell me travel insurance in departed nana's voice would be the day I start signing my voice convos' with a pgp key.
- joeblau 9y agoI wonder how accurately this would reproduce dead musicians voices. I've had this idea for about 8 years called the Notorious BIG project. I have about 20 acapellas that I was originally going to manually chop into a song. Neural Nets can pretty much solve this now.
- LegendaryPatMan 9y agoThis is pretty basic at the moment and it's terrifying. Yeah, it has an MS Sam feel to it, but as the tech improves and we know it will, you could use a service like this to put words in someone's mouth. Think about how you could trip up a CEO or a Politician by playing some random clip that they never said. When that gets into the Zeitgeist judgments will be made in the court of public opinion devoid of facts or real evidence. You could destroy democracy or people's lives with technology like this
- TheOtherHobbes 9y agoI don't think the tech will improve fast. I've been watching speech synthesis since the 80s, and progress hasn't accelerated over that time. Speech synthesis is one of those 90% problems - when you're 90% done, you find you only have 90% left to do. This level of synthesis is relatively easy. Getting to the 'Can reliably pass for the real thing" level is going to take a huge amount of extra work. It's not even about computational power - it's about the sophistication of the models, and their ability to parse words into phonemes correctly with some knowledge of social and linguistic context. "Good enough for some applications" - like phone switchboard systems - is a simpler problem. Virtual impersonation is very much harder.
- meesterdude 9y agoI think you over estimate the complexity and required work to get to virtual impersonation. This will be a problem sooner than you think.
- jedwards1211 9y agoI was pretty impressed by fake Obama's voice. Obviously it doesn't stand up to close scrutiny, but I think if I heard it playing in the background, I could be fooled. And the biggest giveaway was occasional weird intonation rather than the timbre of his voice. All they have to do is make it to where you say a sentence, and it matches your intonation with the other person's voice.
- return0 9y agoThere are human impersonators already. I suppose it's not that easy to fake a visible, high-ranking person for long.
- Markoff 9y agoit's interesting development but it sounds too robotic, there is zero intonation/punctuation, zero variantions in the voice depending on mood of speaker, etc., in the end extremely robotic and if someone really need to fake someone else voice convincingly it would be still easier to hire professional voice imitator
- abetusk 9y agoDoes anyone know of any free/open source alternatives to this? Is it too new to expect a FOSS library?
- bertlequant 9y agoLinda Tripp approves
- sctb 9y agoPlease post civilly and substantively on HN or not at all.
- dyu- 9y agoThis trump version [1] is quite believable. [1] https://soundcloud.com/user-535691776/trump-6 https://soundcloud.com/user-535691776/trump-6
- w8rbt 9y ago“Believe only half of what you see and nothing that you hear.” -- Edgar Allan Poe
- eadz 9y agoCombined with Face2Face[1] live video impersonation, it is truly time to be very careful verifying videos or even live streams. https://www.youtube.com/watch?v=ohmajJTcpNk https://www.youtube.com/watch?v=ohmajJTcpNk
- ericfrederich 9y agoAwesome... Not sure if the voice thing can be done in realtime yet, but you're right... the combination of these two would be awesome
- red023 9y agoHoly shit this is crazy!
- red023 9y agoYeah vote me down coward. Because I used a "evil" forbidden word. I even used in a positive context to show how amazed I am about this facial manipulation (much much more then about the voice thing) Flag me ban me I do not fucking care. I can make another account.
- grzm 9y agoYou were likely down-voted more because your first comment doesn't add anything substantive to the discussion rather than for the language you used. As the guidelines ask, please don't comment on being downvoted, as it makes for boring reading. And doing so in the manner you did is definitely uncalled for. https://news.ycombinator.com/newsguidelines.html https://news.ycombinator.com/newsguidelines.html
- Mz 9y agoIf you are referring to the word "shit," it is not forbidden here and is not likely the reason you were downvoted. I have a terrible potty mouth. I try to keep it PG-13ish online, but if I am tired or something, the way I actually talk tends to come out. My tendency to use the F word like other people use "very" does not appear to be in any way problematic per se. I suggest you rethink your assessment of what is happening here.
- kristaps 9y agoAs noted in other comments, all the samples still sound very robotic, so this is probably "just" a method to tune the parameters of an existing voice synthesizer to mimic a real persons voice as much as it allows.
- TillE 9y agoThat's exactly what it sounds like. The same old mediocre TTS with voices modified to mimic specific well-known voices. It's impressive for what it is, but a lot of people here seem way too excited. This isn't any kind of breakthrough, and only the shortest hand-picked snippet would fool anyone.
- ParadisoShlee 9y agoThe audio feed sounds like they're real and drunk.. so that's impressive
- vermontdevil 9y agoComing soon - fake videos of future political candidates saying outrageous things that will derail their campaigns. Maybe from now on - just learn ASL. Hard to fake a distinctive signing style.
- augustt 9y agoAny ideas on what the underlying technology looks like? Maybe some kind of GAN for audio...
- retox 9y agoThrough the tinny speaker of my mobile phone the Obama in the first sample is almost spot on. Some speed issues with Trump but really impressive.
- wirddin 9y agoIf they can pull this off with the API, this is worth millions of dollars on the table.
- olegkikin 9y agoUntil someone implements it in a few lines of Tensorflow and posts it on Github.
- redsummer 9y agoI can't wait for Richard Burton to read me the news.
- redsummer 9y agoI wonder if you could do this with singing? Feed it acappela Bowie, Sinatra, Elvis songs, then give it new text, and out comes a similar voice and melody.
- inetknght 9y agoSite doesn't load at all on my machine without some javascript from Cloudflare for Ajax. I guess this product isn't for me then.
- pbhjpbhj 9y agoLast week on BBC Radio 4 I heard of a woman who was losing her voice through disease (MND maybe?), a similar system was being anticipated and she was saving voice samples to seed it with. She had been a singer and strongly identified her self with her voice, she wanted to be able to use a speech synthesis system that had her own voice pattern. Apologies if this was already mentioned, but it seems to be a use others here hadn't considered.
- tommoor 9y agoActually fantastic to see there are legitimate uses for this tech beyond the obvious
- phaed 9y agoIf it seemed to you to be a use others hadn't considered, why would you apologize for mentioning it?
- pbhjpbhj 9y agoI hate it when threads are full of the same comment, I didn't diligently search to check it hadn't been mentioned already; ergo preemptive apology.
- dorian-graph 9y agoThis is actually quite inspiring!
- k2xl 9y agoI was just thinking that Stephen Hawking would perhaps be interested in using this to replace his current voice synthesizer (feeding in old interviews of him when he could talk). He has said that he has adopted the current voice since he has associated it with his own, but I wonder if he would prefer his old actual voice.
- Ensorceled 9y agoI think Hawking is now so firmly tied to that voice that he would probably never switch for public speaking engagements and the like. I could see him doing such a switch for personal interactions.
- Sunset 9y agoNow make it say the Navyseal copypasta with Trump's voice, but make him speak slowly and with emphasis.
- koolba 9y agoThe President Obama voice sounds decent. But the President Trump and Senator Clinton voices sound like robots. Reminds me of the crappy text to speech program that came with Windows.
- paraschopra 9y agoI appreciate the ethics link up there in the menu. Not sure if I noticed it on any other AI startup (or for that matter, any startup). Given how complex the world is becoming due to ever increasing co-dependence with tech, I can see how such pages could become as important as 'pricing' or 'sign up' pages. (The privacy issues with Unroll.me, Uber and a thousand other such services will only accelerate this trend). Good job, team Lyrebird. My feedback is that while the inclusion of ethics page is great, it could do with more content on your vision and what you will not let your tech be used for. I know others can develop similar tech, but it will be good to read about YOUR ethics. [Edited for clarity]
- feedjoelpie 9y ago> but it will be good to read about YOUR ethics Not just that, but ethical expectations on the users, backed up by legal policy, would seem important for this.
- tigerBL00D 9y agoI agree, it is reassuring to see that the team is thinking about ethical implications. Judging by the samples from the homepage there are audible artifacts in the recordings resulting from synthesis. I doubt these would pass scrutiny if presented as evidence in court. In some ways forging a voice is like forging a signature, truth can be exposed with enough effort.
- gwbas1c 9y agoNow we can't trust the news anymore. In a year or two we'll never know if recordings are real or not.
- got2surf 9y agoThis is exciting! If you look at historic speeches (ie from American Rhetoric http://www.americanrhetoric.com/top100speechesall.html http://www.americanrhetoric.com/top100speechesall.html), there are large variations in average characteristics between various styles/contexts (on average, pitch/volume/speed are different for inspirational vs somber speeches, for example). But there are also really large differences in the variation - an inspirational speech may be marked by large swings from quiet, reflective pieces to booming, rousing calls-to-action while a somber speech has fewer swings in delivery. For the examples given for various intonations from Obama/Trump, some intonations are much more natural than others. It would be interesting to decide how to parametrize a sentence for the intended intonation. (based on word2vec analysis of the words in the sentence, punctuation cues in the sentence, and perhaps a specified category of "emotional delivery"). It would be interesting at the sentence-level, but also at the macro speech-level to include the right "mix" of intonations for a specific context. On a related note, it would be interesting to study the patterns of intonations in successful vs unsuccessful outbound sales calls, for example, to learn how to best simulate a good human sales voice.
- jtbayly 9y agoCan we get these speeches in audio form now? https://medium.com/@samim/obama-rnn-machine-generated-political-speeches-c8abd18a2ea0 https://medium.com/@samim/obama-rnn-machine-generated-politi...
- leke 9y agoOMG I want to play with this so bad.
- deleted 9y ago[deleted]
- Ensorceled 9y agoThe samples all sound a little like Rich Little and Stephen Hawking's love child doing impressions: they won't fool very many people. But, you can certainly see where this is going and that's the worrisome part.
- dmix 9y agoI'm sure it will improve dramatically over the years. This seems to be a problem with all digital voice software, it's never entirely human sounding. Pretty good starting point though.
- kkotak 9y agoRIP Dan Castellaneta.
- rglover 9y agoThis is fucking terrifying.
- pinpeliponni 9y agoFunny thing is, this is approximately where CIA was with similar technology in closer to 2000. They did some demos for politicians about how they can given anyone's fake their messages. That stuff is golden for propaganda means, and for confusing stuff like military chains of command. Today the CIA probably has worked out all the robotic artifacts already, and their output is really indistinguishable.
- nojvek 9y agoNNets just recently got really good. You are correct though, politicians would love this. I believe technology would make a Judge's life really hard.
- nemo1618 9y agoClearly the solution is to employ a GAN, so that we simultaneously get artificial voices that are indistinguishable to the human ear, as well as judges that are able to reliably distinguish them.
- dmix 9y ago> Funny thing is, this is approximately where CIA was with similar technology in closer to 2000 Source?
- abdias 9y agoNot OP, but here is one related source: http://www.washingtonpost.com/wp-srv/national/dotmil/arkin020199.htm http://www.washingtonpost.com/wp-srv/national/dotmil/arkin02... I do not think the technology involved artificially generated voice though, but simply morphing someone's voice into sounding as the target voice.
- heliumcraft 9y agosource?
- gator-io 9y agoSo much potential for mischief!!
- Nadya 9y agoI see a lot of people claiming that certain things will now be untrustworthy. As if human voice imitators have not existed and could not be paid for prior to this. For $5 you can get Stewie Griffin [0] or Barack Obama [1] to say whatever you want them to say. Any audio-only messages of well known figures should already be considered "compromised" and untrustworthy. Even without the technology to impersonate them. This should be more concerning for "normal people". It isn't that you can no longer trust an audio-only recording of Obama, but that you may not longer be certain an audio recording is from your best friend. (E: Once the technology improves a bit more of course.) [0] https://www.fiverr.com/joe_stevens/talk-like-stewie-griffin-for-you https://www.fiverr.com/joe_stevens/talk-like-stewie-griffin-... [1] https://www.fiverr.com/celebimpression/do-a-custom-barack-obama-impersonation https://www.fiverr.com/celebimpression/do-a-custom-barack-ob...
- simlevesque 9y agoGreat stuff ! Respect from the 514.
- weenkus 9y agoA bit scary thinking someone could do this with ease.
- return0 9y agoWe need a new markup language for intonation and emotion.
- LesZedCB 9y agohttps://xkcd.com/1709/ https://xkcd.com/1709/
- sealthedeal 9y agoThis is super cool, but I know how badly this type of technology could be abused.
- ChairmanPao 9y agoNow people can deny saying things caught on tape. Just show this technology to a jury considering taped evidence, and bring in some experts to testify on how it works. The samples weren't that convincing to me, but could probably be used to switch a word here and there. That may be enough.
- drusepth 9y agoThis is awesome. As someone exploring the fictional storytelling space, this seems like it'd have a lot of fun applications in that space as well. How difficult is it to create/tune voices from parameters rather than training from an audio clip? I build software where people create fictional characters for writing, and having an author "create" voices for each character would be an amazing way to autogenerate audiobooks with their voices, or interact with those characters by voice, or just hear things written from their point of view in their voice for that extra immersion. Having an author upload voice clips of themselves mimicking what they think that character should sound like, but probably would keep traces of their original voice (and feel "fake" to them because they can recognize their own voice), no? Can't wait to see how this pans out. Signed up for the beta and will definitely be pushing it to its limits when it's ready. :)
- hayd 9y agoAnd just as my bank offers a "login via speaking" option. Lovely.
- ksec 9y ago1. Is this company new? 2. Is this better then what Google or Baidu are doing? 3. I remember reading Adobe has something similar. 4. Why ( What happened ) that all of a sudden we have 4 company making voice breakthrough tech like these? 5. What Happen to Voice Acting? Places like Japan where they highly value voice actor. Is Voice even patentable?
- govg 9y ago1) Yes, it is spun off from research at MILA, University of Montreal. 2) Possibly. Google and Baidu have compute resources far beyond a university. This method might be better, and do well with more resources. 3) Adobe's method required far more input data. This apparently requires only 1 minute of your audio to start sounding like you. 4) Deep learning has revolutionized vision, and language processing for a while. It was just a matter of time before people started applying those methods on speech data, with similar surprising results. 5) It will be hard to capture human elements like "emotion" and tone via generative models. Maybe in the future the work will become sophisticated enough to be indistinguishable from human speech, but right now there are some telltale signs that it is artificially generated.
- lordCarbonFiber 9y agoSpecially to point 4, google pushed their wavenet paper a couple months ago. I wouldn't be surprised if some, if not all of these current break throughs are built on that foundation. This sort of application was the first thing that came to my mind after reading the paper. https://deepmind.com/blog/wavenet-generative-model-raw-audio/ https://deepmind.com/blog/wavenet-generative-model-raw-audio...
- kastnerkyle 9y agoThere is an older paper [0] and demo from [1] Alex Graves that inspired a ton of work around handwriting, and then speech. Previous work from Jose Sotelo et. al. (including me) called char2wav [2] is a close neighbor to Graves' approach, though he (Graves) never published the approach for speech so we don't really know. Google's recent Tacotron paper [3] is also a relative to these approaches. WaveNet certainly changed the game in many ways, but approaches to TTS using RNNs have different roots. WaveNet and friends (incl. DeepVoice and NPSS linked elsewhere in this thread) are largely focused on audio modeling, and generally use something closely related to the "classic" TTS pipeline for text in the frontend. The audio modeling results are stellar, and really blew me away personally - basically changing my perspective on what is possible in audio modeling overnight. RNN models try to tackle the whole problem (text + audio modeling) at once, though currently (all?) RNN and attention style models need intermediate / high level hints or pretraining from things like vocoder representations or spectrograms, versus WaveNet's approach using the waveform directly. So they are complimentary in many ways, and I am sure we will see people trying to combine them soon - char2wav has this flavor by using SampleRNN, our lab's take on raw waveform generation though we are still working on the fully end-to-end from scratch training, the inference path is truly end-to-end. Though there are still many details to work out as far as output quality, it seems possible that this will be a productive approach (though I am quite biased). We see similar directions in neural machine translation (NMT) moving from word level representations to word parts or characters directly - one of the big reasons deep learning has come so far, so fast is that a lot of techniques from other subfields can be utilized for new domains, and I think there is a lot more fertile ground for crossover in both directions. Heiga Zen has a great overview talk about how speech synthesis, as a field, overlaps between different approaches and factorizations [4]. His work on parametric synthesis and TTS generally has laid the foundation for a lot of recent advances, and he was also a co-author on WaveNet! [0] https://arxiv.org/abs/1308.0850 https://arxiv.org/abs/1308.0850 [1] https://www.youtube.com/watch?v=-yX1SYeDHbg&t=38m0s https://www.youtube.com/watch?v=-yX1SYeDHbg&t=38m0s [2] http://josesotelo.com/speechsynthesis/ http://josesotelo.com/speechsynthesis/ [3] https://google.github.io/tacotron/ https://google.github.io/tacotron/ [4] https://www.youtube.com/watch?v=nsrSrYtKkT8 https://www.youtube.com/watch?v=nsrSrYtKkT8
- joshmarlow 9y agoFinally, I can have Morgan Freeman narrate my major life events. Update: Reading changelogs before deployment never sounded better!
- anigbrowl 9y agoAs a skilled vocal impersonator I read all my comments aloud in the voice of Morgan Freeman before posting them. Ordinary sentiments such as 'these pickles are quite tasty' are suddenly transformed into profound insights on the human condition.
- keithwhor 9y agoI love this. The business model is too good to be true. 1. Open source voice-copying software 2. At worst, create entire market of voice-fraudsters, at best, very few voice-fraudsters but very high and very real perception of fear of such 3. Become leading security experts in voice fraud detection 4. Sell software / time / services to intelligence agencies, governments, law enforcement, news networks Ethically I'm a bit concerned with (2), but realistically the team is right --- this technology exists, it will certainly be used for good and for bad, and they're positioning themselves as the leading experts. I'm interested to see which VCs and acquirers line up here. Applying a voice to any phrase seems useful for voice assistants (Amazon Alexa, Google Home) but I don't think that's the $B model.
- echelon 9y agoIt sounds like they're training a parametric speech synthesis platform on samples in order to learn the parameters. I wonder if there are are approaches at generating n-phones for concatenative models, or using a hybrid approach. I built a toy concatenative Donald Trump speech system [1], but I don't have an ML background. I've been taking Andrew Ng's online course in addition to Udacity's deep learning program in an attempt to learn the basics. I'm hoping I can use my dataset to build something backed by ML that sounds better. Is anyone in the Atlanta area interested in ML? I'd love to chat over coffee or join local ML interest groups. [1] http://jungle.horse http://jungle.horse
- kastnerkyle 9y agoI tried similar approaches long ago (~2 years now?) with something related to RNN-RBM and it showed some slight glimmer of promise, and still think there might be some clever ways to combine concatenative methods and deep learning to avoid a lot of the noise issues present in parametric models. Then again, maybe it just needs to train longer - it's always hard to tell. I liked jungle.horse, awesome stuff!
- sehugg 9y agoSounds great, I was trying something like this in Keras but didn't get very far: https://github.com/sehugg/kerasspeechcodec https://github.com/sehugg/kerasspeechcodec
- xumx 9y agoBe right back (Black Mirror) let's do it.
- nerfhammer 9y agoHello. My name is Werner Brandes. My voice is my passport. Verify me.
- sna1l 9y agoCharles Schwab uses a voice phrase to authenticate you for access to your account, which is already pretty brittle, but I hope this makes them reconsider more urgently.
- theemathas 9y agoIt's a matter of time before this can compete with Vocaloid.
- mericsson 9y agoRelated Economist article: http://www.economist.com/news/science-and-technology/21721128-you-took-words-right-out-my-mouth-imitating-peoples-speech-patterns http://www.economist.com/news/science-and-technology/2172112...
- cocoa19 9y agoThis technology reminded me of 24 (TV series). The plot of season 2 has Jack Bauer prove a Cyprus recording between a terrorist and high-ranking Middle East officials was forged so the US president would start a war.
- rajacombinator 9y agoWow had no idea something like this was possible. Very impressive.
- lucidrains 9y agoLol, I totally called this.
- jpsim 9y agoCurious choice to name a company & product with a name that sounds like "Liar Bird" when spoken. To me, that looks like they're fully embracing the concept that this can be used for nefarious purposes. If one of their goals is to bring attention that this technology exists and can be misused, the name reinforces that.
- ythn 9y agoI don't see why audio manipulation would be any more nefarious than photo manipulation or video manipulation. Plus, it doesn't even matter. You can write an article with fake quotes and people will believe it without even caring if there is an accompanying sound byte or not.
- coldsmoke 9y agoI guess they've named it that because the lyrebird is an amazing impersonator. The end of this BBC clip blew my mind the first time I saw it. https://youtu.be/VjE0Kdfos4Y https://youtu.be/VjE0Kdfos4Y But you may have a point, and the ethics section makes it clear that they are indeed very aware of that this may be misused.
- anigbrowl 9y agoExcellent work. This will find widespread application in the film/tv/music industry and beyond (and we're not that far away from being able to do the same thing for video). Unfortunately it will also be widely abused, but given the near-inevitability of such technological development I'm already reconciled to that :-/
- LordKano 9y agoThis is impressive. There is now a way for Morgan Freeman and James Earl Jones to be able to narrate movies forever.
- bisRepetita 9y ago1. Buy the rights for "Car Talk" re-broadcast. 2. Record new, current ads using Click and Clack's voices. 3. If the voices sound a little too "mechanic", pretend it's a joke.
- amarant 9y agoThey lost me at "... Consumers are still not lining up to buy EV's" What the fuck are they talking about?
- Tloewald 9y agoThis is very exciting to me because it lets RPGs provide spoken dialog for everything (I'm waiting to see if they can do emotions at all convincingly). Even big budget games suffer from "you can call your character anything as long as it's 'Shepherd'" simply because you can't mention the character's name or any other use-content safely.
- mod 9y agoDoes the API get better results with more training data?
- mzzter 9y agoTrump 6 speaking "... my intonation is always different" sounds very convincingly human.
- stefek99 9y agoI have two domains: - legalscreenshot.com - legalprintscreen.com I also developed a concept of "Reality Check" similar to Touring Test (when VR and AI becomes so convincing >50% people won't distinguish it from base reality)... Too bad I'm on the corporate network and my personal website is blocked: https://genesis.re/wiki https://genesis.re/wiki Aside: do you believe psychedelics should be the part of obligatory astronaut training?
- olleromam91 9y agoSo all my voice commands can be recorded and my voice can be replicated. Cool...i guess
- deavmi 9y agoMy word. Awesome.