42 ms·
This Voice Doesn't Exist – Generative Voice AI
- psychphysic 4y agoMy word those female voices for news and controversy are AWFUL. I only made it 2-3 seconds in. The male narrative voice is silky smooth. In fact I prefer to the classic YouTube male mystery voice that sounds like the narrator had a lobotomy.
- goleary 4y agoImpressive generated voices for TTS
- 7373737373 4y agoStill nothing comes close in terms of "install/usage accessibility" compared to the terribly sounding espeak
- GoldenButthole 4y ago[flagged]
- technick 4y agoWere you listening to it with audiophile level sound drivers?
- GoldenButthole 4y agoNo. The problem is not sound quality. The problem is that the speech sounds unnatural and computer-generated.
- akurilin 4y agoImpressive. Any chance there will be an API version of this product for real-time apps?
- techdragon 4y agoThis is the major issue with the majority of this technology at the moment. Theres a plethora of options available and soon to be unveiled by several startups who are talking up their tech... but they are almost all for "editing"/"after the recording" work. You have to have a complete recorded track you can pass into their software (usually by uploading to their service) and then it will crunch away at the file and work their magic. The current real time options I've found are... lacking, they are mostly fake/toys (not actually using voice cloning, just old school pitch shifting) or tech demo videos, with a scattering of research papers which are highly variable in terms of "how easily can i reproduce this", ranging from "sure if I want to waste money on a google colab instance, to "only works with specific model of video card due to reasons" If you know of any real-time (audio stream in -> audio stream out) voice cloning/transform/replacement tools, feel free to post about them in a reply, this is an area of tech I'm trying to keep on top of and I'm only human so I have no idea what new company or research I might miss.
- matisqe 4y agoHey - ElevenLabs dev here. The quality above works with <1s latency that for some real-time apps is already sufficient. On smaller chunks of text it can be as quick as ~500ms.
- akurilin 4y agoAwesome, thanks for adding that extra color!
- drewbug01 4y agoThe “narrative” example is pretty good, but the “conversational” example is rather unpleasant to listen to. (Especially if you know how well Meryl Streep delivers that monologue in the original: https://youtu.be/Ja2fgquYTCg https://youtu.be/Ja2fgquYTCg)
- xp84 4y agoMaybe it's because I haven't heard the source material, but that Conversational voice really appeals to me. I wish my phone and assistants used that voice. (and also I can't wait for a "real" ChatGPT-era AI to go with it, to put those braindead jokes of an "assistant" Siri, Alexa, and Google Assistant out to pasture)
- tedd4u 4y agoThat's a pretty high bar. Even most Hollywood productions can't afford Meryl Streep, let alone a new site, podcast, or video game. From wikipedia: Mary Louise "Meryl" Streep [is] often described as "the best actress of her generation." Streep is particularly known for her versatility and accent adaptability. She has received numerous accolades throughout her career spanning over five decades, including a record 21 Academy Award nominations, winning three, and a record 32 Golden Globe Award nominations, winning eight. She has also received two British Academy Film Awards, two Screen Actors Guild Awards, and three Primetime Emmy Awards, in addition to nominations for a Tony Award and six Grammy Awards.
- rkagerer 4y agoAgreed, the intro and narrative ones are great. The news one is terrible.
- logicallee 4y agoLet's talk about this "Narrative" example. When I listened to it, my first impression was that it must be the real actor they included for comparison purposes but that they failed to label it correctly. I thought it is not machine-generated. I couldn't tell the slightest artifact except what sounded like a low-bitrate sound encoding (maybe using a codec geared toward speech). Can you tell anything "off" about it? As for the encoding artifact such as a tinny sound or low-bitrate sound, that is the type you hear on an MP3 or low bitrate codec for speech. For example, when I record a message on https://vocaroo.com/ https://vocaroo.com/ the "premier" voice recording service it sounds 10x worse. Here is a sample I just recorded of my own speech: https://voca.ro/18oSJ1sHU5w5 https://voca.ro/18oSJ1sHU5w5 After my first impression that the narrative example might be a real human mislabelled for comparison purposes, I listened to the next two, labelled News and Conversational. I found these very easy to tell as AI-generated. Thinking back to why I found the narrative example so compelling, I thought perhaps the issue is that the first example is in British English which I'm less used to than American English. I grew up in the United States. Perhaps since the accent doesn't match my own, it is harder for me to perceive it as generated. -> Can a native speaker of British English tell us whether listening to the first example you can tell in any way that it is a robot? Maybe it is as obvious to you as the next two are to me. Still, I've listened to a fair amount of British English in my life so perhaps there is an alternative explanation for why the first one was better. For example, it could have been trained on a reader's voice who has narrated thousands of hours in very high studio quality in a fairly consistent way, leaving this type of text much easier to synthesize than the other two examples due to more training data or higher-quality audio. For me, the first one is really indistinguishable from a narrator's true voice, though it does sound a bit tinny which could also happen as an artifact of the recording process. In terms of "how confident are you that this is a real person" the second two examples I would put at 0 - it's totally obvious that it is not a real person, whereas the first one sounds like a 10 to me: obviously a real narrator. (With a bit of artifacting that sounds like an mp3.) [1] The text is here https://www.nytimes.com/2001/11/19/books/chapters/the-lord-of-the-rings-the-fellowship-of-the-ring.html https://www.nytimes.com/2001/11/19/books/chapters/the-lord-o...
- dalmo3 4y agoI found the samples incredibly good. But the samples in their other post about conveying emotions[0] are still far from acceptable. In any case, I'm hoping this can be expanded to other languages as it would be an amazing tool for language learning. [0] https://blog.elevenlabs.io/the_first_ai_that_can_laugh/ https://blog.elevenlabs.io/the_first_ai_that_can_laugh/
- piotr11 4y agoThanks (ElevenLabs dev here), we are constantly working on improving our model, we do out own research and train it completely from scratch. We do support Polish already and the quality is actually better IMO than English as we use a newer generation model: https://www.youtube.com/watch?v=ra8xFG3keSs https://www.youtube.com/watch?v=ra8xFG3keSs Some people think it is fake and we hired a real voice actor to read.
- CobrastanJorji 4y ago> Not only can they be more cost-effective without compromising on quality... That feels dishonest. Even if this AI is just as good at speaking as a professional voice actor (which I'm not sold on), a voice actor does more than just read the line. In ideal circumstances, they have a lot of context for what their character is doing and feeling. Is this potentially a good option for saving money on video game voices? Quite possibly yes. Is there no compromise on quality? No, not yet. Past that, the whole "Ethical AI" section's arguments seem ridiculous. Of COURSE it puts the livelihoods of voice actors at risk. Your product's whole point is that fewer man hours are needed for voice work. Just accept that you're making those jobs obsolete. There's a perfectly good argument that it's okay to do that. Throwing bullshit at us to convince us that "no, the voice actors will still have lots of work, and they won't even have to talk!" just makes you sound like snake oil salesmen.
- deleted 4y ago[deleted]
- underwater 4y ago> Even if this AI is just as good at speaking as a professional voice actor (which I'm not sold on), a voice actor does more than just read the line. In ideal circumstances, they have a lot of context for what their character is doing and feeling. How long do you think that advantage will last for. Years? Months? Weeks?
- sebzim4500 4y ago>Even if this AI is just as good at speaking as a professional voice actor (which I'm not sold on) I think the 'narration' example was voice actor quality. The other two were slightly off.
- TrackerFF 4y agoSounds damn good. Would it be possible to use your own voice for training, and replicate it? Obviously that could come with some serious security risks, but it would also make content presentation much easier for many people. Gone are the days of doing voiceover recordings for videos.
- dmd 4y agoSee https://news.ycombinator.com/item?id=34309306 https://news.ycombinator.com/item?id=34309306
- matisqe 4y agoHey! ElevenLabs dev here - yes exactly! We do rapid voice cloning (just on few seconds samples) that for American accents works really well - which is already available in Beta. We can also do a professional near-identical copy with longer samples too.
- jaapbadlands 4y agoI'm both scared and peeking through my fingers at the thought of the evolution of vocal-tuning plugins like Melodyne. Currently you can basically draw the pitch of a vocal performance, however using AI you could re-render the wavefile and adjust more parameters than simply pitch - such as timbre, inflection, vibrato, dynamics, distortion, openness, softness, breathiness, or a bunch of other vocal attributes.
- spacechild1 4y agoVoice synthesizer plugins, such as Vokaloid or Synthesizer V, can already do that quite convincingly, so it is only a matter of time before it can be applied to existing voice recordings.
- helloworld 4y agoTheir Steve Jobs voice simulation is creepily good: https://www.youtube.com/shorts/34vB41lyQ-A https://www.youtube.com/shorts/34vB41lyQ-A
- jcims 4y agoThis probably breaks HN etiquette but... wow
- abnercoimbre 4y agoWe should be allowed to break etiquette for the rare and the shocking!
- deleted 4y ago[deleted]
- xeonmc 4y agoImagine if in-game voice chat automatically converts player speech into the voice of the character they're playing -- this would resolve a lot of the gender-based harassment problems arising from competitive games requiring vocal communication, since now _everyone's_ default is hiding the actual player's voice, contrasting the "just use a voice changer if you're a girl playing" suggestion which themselves draws attention by being out of the ordinary.
- yieldcrv 4y agoI’m looking forward to NPCs having dynamic responses with real voices Doesn't have to be prerecorded, just trained
- wlesieutre 4y agoGames could have more than three dialogue options again!
- 93po 4y agoI feel like if Bethesda really wants another industry defining game, this is the path they should be taking. AI generated conversation with AI generated voice acting with voice-to-text recognition. You can literally have microphone-voice conversations with NPCs that have rich, AI generated backgrounds and personalities.
- PaulBGD_ 4y agoEven bigger than that (I think at least) is the potential for fully voiced mods. There’s nothing stopping modders at that point from adding content indistinguishable from the base game.
- 93po 4y agoI doubt Bethesda would facilitate this. They'd likely use voice actors to train the voice, and having a famous voice actor saying saucy kink/bdsm/violent things that you tend to see in some mods wouldn't be great PR
- _carbyau_ 4y agoMy voice is my passport, verify me.... aww fuck I have to do a voice activated "I am human" check now?!?
- leeoniya 4y agoRIP Auto-Tune
- abraxas 4y agoIf the music that my grocery store foists upon me is anything to go by then I say it can't come soon enough
- mc32 4y agoThis is awesome for any kind of situation where you need a (human) speaker. No tripping over words, mumbling, mispronouncing --all fluid and audible with perfect enunciation!
- coverband 4y agoThis is cooler than ChatGPT and image generation as far as I'm concerned. If they're able to bring out the emotional connectivity and purposefulness of the human voice, it will be revolutionary...
- belter 4y agoThe laughing examples are pretty impressive. "The first AI that can laugh" - https://blog.elevenlabs.io/the_first_ai_that_can_laugh/ https://blog.elevenlabs.io/the_first_ai_that_can_laugh/
- cheeseface 4y agoThere are so many uses cases for this, even with the current quality. Many game developers dream of having something like this.
- intelVISA 4y agoAwesome, I think a few years we'll hit levels of AI generative media tech where you can produce, as a lone greybeard, a Cyberpunk 2077 tier title. Same # of bugs too ;)
- anigbrowl 4y agoLess than a week ago, I said AI would upend the market for voice actors within the next couple of years: https://news.ycombinator.com/item?id=34271948 https://news.ycombinator.com/item?id=34271948
- bsenftner 4y agoNot only voice actors, include radio hosts, documentary/news content, any voice over for anything, as well as imitation of familiar voices.
- holler 4y agoThis will really open pandoras box for scammers and other bad actors. Grandma won't know she's speaking with an AI.
- elboru 4y agoGrandma already falls for scams. Will I know I’m speaking with an AI?
- bee_rider 4y agoAny interaction that you didn’t kick off is a scam. Whether is is an AI or human is irrelevant.
- yCombLinks 4y agoThat would mean any interaction I initiate is a scam for the other party
- 93po 4y agoI think this easiest question for a turing test to AI: "What would you choose as a turing test for an AI?"
- spaceman_2020 4y ago
- stanislavb 4y agoI'm "waiting" for the time when scammers will start calling us with similar voices.
- telis 4y agoWaiting? Talk to some small business owners, they’re already being bombarded. One common tactic is to ask the A.I. to do maths and see it breakdown and say it’s going to ask it’s “supervisor”
- purplepatrick 4y agoStill sounds pretty fake to me. There’s a hurriedness to the speech and a monotonic uniformity in enunciation that is uncannily machine. Good to know that voice actors will have jobs for a while longer…
- drivers99 4y agoI thought the Narrative one was 100% there. I'd still give the News one 99% and Conversational 98%.
- affgrff2 4y agoYes, for the sake of humanity, I hope the examples are cherry picked and The lord of the rings audiobook is in the train set...
- UncleEntity 4y ago> Good to know that voice actors will have jobs for a while longer… They don’t have to work anymore, just sell their voice and sit at home collecting royalty payments is the future according the TFA. And they’ve been making progress on the roboticness with every new model that comes out. Just a matter of time (and data) for the AIs to figure out how words string together naturally.
- janosdebugs 4y agoThis assumes that legislation/ajudication won't tell AI companies that grabbing any content they can find and not reimburse the original author is "fair use" or something equivalent in other jurisdictions. Here's to hoping.
- Kiro 4y agoYeah right. You would never pass a blind test on this.
- sebzim4500 4y agoI think the narritive one would pass a blind test. The conversational one wouldn't, although it could pass for a bad (human) voice actor.
- idealmedtech 4y agoThis would put companies like Audm out of business, but it seems like they already only employ one voice actor for most gigs (ya gotta respect how much she gets done though!). I wish there was more work for professional voice actors, audiobooks done by the likes of Roy Dotrice are an absolutely fantastic ride
- deleted 4y ago[deleted]
- UncleEntity 4y agoI’ve been reading up on this the last couple of days because…oh, look, squirrel! This seems to me where The Big Guys are going to dominate because it comes down to a big data problem. For example, whisper (admittedly speech to text) was trained on 480,000 hours of speech data scraped from the web. The next ‘contender’ used something like 48,000 hours. Who can compete with that who doesn’t own a whole cloud?
- sudofail 4y agoI think a great use case for this technology could be to preserve dying languages. I'm sure a lot of work has already gone into preserving the written form of these languages, but training models on data sets of native speakers could be a way to preserve pronunciation.
- dr_kretyn 4y agoNice timing as I'm looking for a way to replace espeak. Are there any pretraines text-to-speech models available? Or, some dataset that could be use to train a model?
- mlboss 4y agoTry https://github.com/neonbjb/tortoise-tts https://github.com/neonbjb/tortoise-tts
- UncleEntity 4y agoThat one requires a big GPU and isn’t real-time. If you want to clone a voice and have a shitton of compute to fine tune it’s a good one. If you just want your computer to tell you you need to be out the door in 30 seconds or you’ll miss the bus then not so much.
- dr_kretyn 4y agoI found that it's much easier for me to read and remember when reading with voice assistant for which I need real-time. Ages ago I bought Ivona Text-to-speech and was serving me very well for many years. The last few years I used AWS Polly and espeak (using this https://github.com/laszukdawid/cracker https://github.com/laszukdawid/cracker) but thought that there must be something better.
- UncleEntity 4y agoThere seems to be a fairly wide selection between state of the art and just glueing together a bunch of phonemes it’s just that tortoise-tts is up there with the state of the art. I haven’t looked into the mid range stuff but there’s probably something out there with pretty good quality if you don’t mind doing some coding, end user applications seem to be mostly in the startup SaaS charge by the character domain.
- dr_kretyn 4y agoThanks! That was the first search and has nicely written colab, so will definitely give it a try. However, I've seen in readme that generating a sentence takes quite a long time. > On a K80, expect to generate a medium sized sentence every 2 minutes. Are you aware of other available?
- patientplatypus 4y ago[dead]
- statsstats 4y agoAmazing.
- statsstats 4y agoWhere can I use this? Is it public? Is there an API?
- matisqe 4y agoHey! ElevenLabs here - not public yet, but we will be opening up Beta later this month. API is available directly in the platform.
- didericis 4y agoI can't tell if I'm starting to get that old person "new things are scary" instinct or if my gut level of fear about the implications of these things is warranted. As impressive as a lot of these models are, I can't help but feel like they're going to end up making an incredible amount of sterile soulless content that makes everyone's lives worse. We're already drowning in ad dominated cynical soulless computer generated search results. Are all online forums going to end up being drowned out by cynical pumped out super cheap to produce simulacrums of creative content now too? If I want people to buy more Triscuts next year what's stopping me from writing a bunch of prompts to insert subtle marketing cues to buy Triscuts with entire fake ecosystems of users, fan art, radio call ins, user stories, etc in like every niche community in existence and flooding them with soulless fake interaction? That exists to a certain extent already, but I don't see how this stuff won't make it way easier, way more effective, and way more widespread.
- vouaobrasil 4y agoI agree with this completely. Technology has always made us trade quality for low-quality quantity in exchange for convenience. People now interact more through technology which removes a lot of body language and other enriching experiences. The most dangerous aspect of this is that each step seems relatively harmless: right now, ChatGPT and DALL-E are amusements, but each small step is building a monstrous and as you say, soulless machine that overloads us so much that we will forget what it's like to even be human. I firmly believe (and I have given this a lot of thought) that technology is ultimately evil, and that tech companies are trading short term gain of enormous wealth for the very essence of humanity, preying upon the basic instincts of individuals who are also trading their personal worth for convenience. If I could have one single wish fulfilled in this world, it would be that every single human being gain a natural and instinctual revulsion for advanced technology. If someone asked me what disease was the worst that ever plagued humanity, it would not be smallpox or the flu or COVID, it would be the tech company.
- snek_case 4y ago> If I could have one single wish fulfilled in this world, it would be that every single human being gain a natural and instinctual revulsion for advanced technology. If someone asked me what disease was the worst that ever plagued humanity, it would not be smallpox or the flu or COVID, it would be the tech company. Time to go live in a cabin in the woods and go write your manifesto on a typewriter...
- WheelsAtLarge 4y agoI'm listening to an audiobook whose reader is not as good as some of these voices. At one level, I'm impressed but at an another I'm sadden since we are heading towards uncharted territory. We are looking at a future where we'll have content, video,audio, and text by the truckload. More does not mean better. It just means more blah stuff. I don't think that's the future I'm looking forward to live in.
- Fordec 4y agoThe key will be authenticity and trust. And in the world where the percentage of online content that contains this ends up in the vast minority of content, in person expertise and meetings will have to make a return out of sheer necessity. It's starting to very much feel like we're entering the age of information manipulation outlined in the Ghost in the Shell TV series. Except it isn't a 90's/00's depiction of the future, it's just with far less robots and prosthetics and a lot more mundane. I just keep coming back to the scene where they have satellite video footage of a nuclear submarine preparing for a nuclear attack and the discussion lamenting that it's just video, nobody will believe it as evidence.
- kerpotgh 4y ago[dead]
- sanroot98 4y agoI think you are overestimating the capabilities of ai to create novel content ,high genuine quality content will be always there ,but amount of bs content will increase
- LarsDu88 4y agoI've been using Azure to generate speech audio for my game and it's extremely good. These samples seem even better. I'm wondering how less cherry picked clips will turn out
- rvz 4y ago> At Eleven, we're fully committed both to respecting intellectual property rights and to implementing safeguards against potential misuse of our technology Unlike Stable Diffusion trampling over the copyright of artists without their permission and OpenAI doing the same for code mangled with incompatible licenses and monetizing it and outputting the trained data verbatim whilst opening a pandora's box and then attempting to write detectors and watermarks afterwards. I'm skeptical on Eleven Lab's statement on adding their detectors before release, but we'll see. Should there eventually be an open source version of a competing model by someone, it should be trained on public domain sources. This was the case with Dance Diffusion as Stability AI would have been sued to the ground by the RIAA had they done that. [0] [1] It will only be a matter of time before the legal system catches up with AI generated content and scrutiny over the trained data on copyrighted content without permission and how it was trained. Any output generated by an AI is automatically public domain and un-copyrightable. [2] This AI hype is another VC scam to unload their investments in AI startups to big tech once again and then pretend how AI is making the world better but when they know it is actually doing the opposite with far reaching consequences. Of course it can't be stopped, but it also cannot go unchecked and unregulated forever. [0] https://www.musicbusinessworldwide.com/record-industry-clamps-down-on-ai-based-music-extractors-that-infringe-on-copyrights/ https://www.musicbusinessworldwide.com/record-industry-clamp... [1] https://techcrunch.com/2022/10/07/ai-music-generator-dance-diffusion/ https://techcrunch.com/2022/10/07/ai-music-generator-dance-d... [2] https://www.copyright.gov/rulings-filings/review-board/docs/a-recent-entrance-to-paradise.pdf https://www.copyright.gov/rulings-filings/review-board/docs/...
- hnbad 4y agoAre you really surprised though? The crypto hype was a vehicle for VC backed companies to sell unregulated financial products to retail investors. The "gig economy" was a vehicle for VC backed companies to skirt labor protections and zoning laws. And now AI is a vehicle for VC backed companies to skirt copyright laws. "Disruption" is often just about finding edge cases of existing laws and regulations and exploiting them for profit until legislation catches up.
- rvz 4y ago
- panza 4y agoSay you're an indie game developer. In 2022 you'd pay someone on Fiverr to do a 'trailer' voiceover on your game trailer. This year, you'd use this - and also get a few more languages in there. Next year is gonna be an interesting year.
- hyperific 4y agoShould mention this to the https://thisxdoesnotexist.com/ https://thisxdoesnotexist.com/ dev
- dj_mc_merlin 4y agoThe examples are insanely good. Insanely good. I can barely believe we really live in a world where this is possible. I don't have anything constructive to add.. just wow.
- wand3r 4y agoI work in TTS and i just dont believe this. If these really are random text and not trained on literally the copy they are reading, with no correction I would be surprised. Also, our competitors have good voices but they also take ages to produce. Maybe these really are legit but take like 1 minute to produce or something. So while this is impressive, i doubt that in practice this would be this high quality and could even approach real time
- ThePyCoder 4y agoI want to agree, but I searched on their website and found their narration service with 2 full book examples. I listened to the first one for a while and it's the first time an Ai narrator was good enough to keep me listening: https://www.audiostory.ai/2065785/11707800-alice-s-adventures-in-wonderland-by-lewis-carroll?t=0 https://www.audiostory.ai/2065785/11707800-alice-s-adventure...
- kreddor 4y agoIt's noticable worse than the examples in the blog post. I mean, it's good enough for listening, but no better than the competition.
- kerpotgh 4y ago[dead]
- sebzim4500 4y agoIt's vastly better than any TTS system I have used, but then I've only used a few (mainly phone assistants and the thing built into kindle). What is the competition that you are referring to?
- TarasBob 4y agoThere’s a pretty cool trinity audio bot that converts any Twitter thread into audio: https://twitter.com/trinityaudiobot/status/1613166071690797058?s=46&t=EFOP6iub3yyX5EwnfUYIzw https://twitter.com/trinityaudiobot/status/16131660716907970...
- 93po 4y agoTwitter is bad enough to read much less have to listen to
- aleem 4y agoThey need to take this and similar AI and come up with better dubbing for movies in other languages. Netflix should really lead the way here with the amount of dubbed content that they currently possess.
- piotr11 4y agoExactly (ElevenLabs dev here)! This is actually out mission - make all content available in any language and voice. Dubbing is where we are going!
- barking_biscuit 4y agoIf dubbing is where you are going... does that mean you're also going to pair it with deepfaking the videos to make the facial movements match the new vocalizations? Because that'd be a wild product.
- logicallee 4y agoInterestingly, some of the robot styles take a very obvious and dramatic fake breath. I say "fake" since a robot doesn't need to breathe and it's not exactly considered a phoneme. The fake breaths don't really make the robot sound more convincing. When you listen to the first example labelled "Narrative" you can tell where a human speaker would have inhaled (which is something the AI could have picked up on from copious training data) though the inhale itself could be muted in post-editing, e.g. after the long 24-word first phrase[1] ending in "special magnificence", and then again at the end of the sentence. It could just be the way the AI reads the comma but it is very convincing. The "News" and "Conversational" examples don't include that pause effect. In the cerulean monologue, there is no pause after "for instance" despite it being in the monologue. However, the robot takes a deep dramatic breath after the word "I see"[2]. " Oh, okay. I see, [DEEP LOUD DRAMATIC BREATH BY ROBOT], you think this has nothing to do with you. [LOUD DRAMATIC HALF BREATH BY ROBOT] You go to your closet and you select I don't know that lumpy blue sweater for instance because you're trying to tell the world that you take yourself". There is no pause on the comma around "for instance" though the script has one. I decided to check whether the robot is just copying the original film exactly and that's not it either.[3] Comparison: Robot: "Oh, okay. I see, [DEEP LOUD DRAMATIC BREATH BY ROBOT], you think this has nothing to do with you. [LOUD DRAMATIC HALF BREATH BY ROBOT] You go to your closet [no breath] and you select I don't know that lumpy blue sweater for instance [QUICK HALF BREATH BY ROBOT] because you're trying to tell the world [no breath] that you take yourself too seriously to care about what you put on your back but [no breath] what you don't know is that sweater is not just blue it's not turquoise it's not lapis it's actually cerulean." Original: "Oh, okay. I see [no breath] you think this has nothing to do with you. [loud long breath] You go to your closet [breath] and you select I don't know that lumpy blue sweater for instance [no breath] because you're trying to tell the world that you [breath] take yourself too seriously to care about what you put on your back but [breath] what you don't know is that sweater is not just blue it's not turquoise it's not lapis it's actually cerulean." Text: "Oh, okay. I see, you think this has nothing to do with you. You… go to your closet, and you select… I don’t know, that lumpy blue sweater for instance, because you’re trying to tell the world that you take yourself too seriously to care about what you put on your back, but what you don’t know is that that sweater is not just blue, it’s not turquoise, it’s not lapis, it’s actually cerulean. " I've annotated the breaths in the "conversational" robot sample vs the original film: Robot Original Same/different? I see... [Loud breath] [no breath] Different with you... [Loud quick breath] [loud long breath] Similar your closet... [no breath] [breath] Different for instance... [QUICK half breath] [no breath] Different that you... [no breath] [breath] Different back but... [no breath] [breath] Different The robot's loud dramatic breath is unmistakable, but it's clear it's not copying the source exactly, since it occurs at different places. [1] The text is here: https://www.nytimes.com/2001/11/19/books/chapters/the-lord-of-the-rings-the-fellowship-of-the-ring.html https://www.nytimes.com/2001/11/19/books/chapters/the-lord-o... [1] The text is here: https://artdepartmental.com/blog/devil-wears-prada-cerulean-monologue/ https://artdepartmental.com/blog/devil-wears-prada-cerulean-... [2] https://www.youtube.com/watch?v=us52N76XA28&t=1m24s https://www.youtube.com/watch?v=us52N76XA28&t=1m24s
- logicallee 4y agoBy the way if anyone is in this thread due to working on AI speech synthesis for any company, I am interested in AI as well as audio production and I would love to talk about joining the team as an AI researcher. Just send me some mail, my email is in my profile.
- pronlover723 4y agoWhat are the odds of this kind of thing being open source so I can use it at home. So far, most of the "good" text-to-speech systems are all commercial services https://aws.amazon.com/polly/ https://aws.amazon.com/polly/ https://cloud.google.com/text-to-speech https://cloud.google.com/text-to-speech https://azure.microsoft.com/en-us/products/cognitive-services/text-to-speech/ https://azure.microsoft.com/en-us/products/cognitive-service... And now one is also a service. I tried using tortoise-tts on my M1. Generating a 7 minute speech took 3 days and, while better than the 15 yr old text-to-speech built into the OS it wasn't close to the quality of the services above. Maybe I don't know who to use it but of course it's not as simple as text-to-speech. You need the system to ideally understand the text it can act out parts Of course see my username. I want to generate personal adult content so I'd prefer not to upload it to a service.
- yreg 4y agoAny time I see AI model news on hn nowadays, my first question is whether I can run it locally, and if not, what are the alternatives that I can run locally.
- EarlKing 4y ago> what are the alternatives that I can run locally ...you will be disappointed by the answers to that question for the foreseeable future.
- yreg 4y agoI'm the opposite of disappointed. The amount of public pretrained models that have been popping up recently is crazy.
- deleted 4y ago[deleted]
- bdhcuidbebe 4y agoSame model with random tweaks applied. Just because there is a new toy doesn’t mean capitalism gave up.
- 152334H 4y agoNo mention of any other competitors that've been doing this stuff for several years? Uberduck? Fakeyou? Coqui? 15?
- sebzim4500 4y agoThis appears to give by far the best results that I have seen. Can you link to a speech sample that you think is similar in quality to the article?
- jurassic 4y agoI'd like to see this technology become cheap and ubiquitous enough that everyone can choose for themselves what voice they would like to hear right at the moment of consumption. It's always a huge bummer when there's a book I want to listen to on audible with terrible narration. Somebody must have liked that voice for the person to be hired, but people's tastes differ and sometimes the people they've selected just really grate on my ears. It would also be cool if celebrities / existing voice talent could somehow license the synthesis of their voice. I read something about James Earl Jones doing this with Disney for future Star Wars projects. I'm sure there are people out there who would love to have every work they listen to be in the voice of their favorite narrator/celebrity.
- feoren 4y agoOkay can I ask a question that has been bothering me for a long time? Why do seemingly all these text-to-speech programs attempt to produce spoken voice based solely on raw text? Why don't they consume a MIDI-like text-markup language where you can write phonetic pronunciations along with markup about the emotion, volume, speed, etc.? I feel like this is a huge unnecessary roadblock holding back this kind of technology. It'd be like if every music composition program rendered a wave file not by MIDI or VST, but by trying to visually read sheet music. I totally understand why TTS solutions that have to consume arbitrary content, like screen-readers, need to read purely raw text. But content creators don't need to be limited to raw text! Why is everyone doing it that way? Where is the TTS markup language for content creators?
- bredren 4y agoI don’t know how many of the solutions offer this, but there is a markup language for TTS: https://en.wikipedia.org/wiki/Speech_Synthesis_Markup_Language https://en.wikipedia.org/wiki/Speech_Synthesis_Markup_Langua... Amazon Polly, (which seems kind of ancient with all these new solutions showing up) has supported SSML for some time. AWS Polly SSML docs: https://docs.aws.amazon.com/polly/latest/dg/ssml.html https://docs.aws.amazon.com/polly/latest/dg/ssml.html
- KRAKRISMOTT 4y agoIn practice they are next to useless, the expressions are not very...expressive (just try it in the AWS editor). I suspect a LLM would be able to infer the context or we can use prompt engineering to generate the appropriate tokens encoding emotions for the intermediate neural codecs directly (Mel spectrograms are so passé now post Vall-E).
- slim 4y agomaybe the only way to express speech precisely is the speech itself ?
- TaylorAlexander 4y agoSomething I always noticed is that they get Morgan Freeman to do voiceovers for science shows, but he’s not a scientist so he has a sort of generic inflection when he talks about the various ideas in the script. And then you watch Carl Sagan’s COSMOS, where he co-wrote the material, and there is so much depth and expression to his delivery. There’s a lifetime of public speaking, specifically delivering complex scientific topics to a general audience, that Sagan drew from when recording his show. Sagan would have learned this through conversation with people, and careful updates to his expression and delivery as he matured. I guess an LLM could improve upon previous methods but I would also say there is a gap that even humans struggle with, which requires really complex knowledge both of public speaking and of the material. It may be a long time before we can really master that with AI systems.
- Animats 4y ago"voice owners and their licensors" Is that even a thing? You can't copyright a voice. There can be a personality right under state law, but the main case on that was someone hired to sound like Bette Midler for a commercial.
- Prunkton 4y agoThe text to speech function at the top of the article is the actual product but they are not going the extra mile and record it again for the other speed multiplier like x.7 or x2.0. You can clearly hear the mp3 struggling, especially at 0.7 speed. It would have been interesting how they perform in comparison. The fact you are able to adjust the voices is even one of their selling points. I really wonder why they haven't done that
- antman 4y agoI always wondered why those generative voices dont capture the feeling of the text per segment and incorporate it to the output e.g. news, narration, first person hunted by vampires, whatever. Seems like a kind of low hanging fruit. Disclaimer: I use tons of audiobooks so that might not be what people need in general
- fritzschopen 4y agoRespeecher is doing this for years now. i don't see any major advancement.
- junon 4y agoRespeecher is a completely different product with different goals and usecases...
- ggerganov 4y agoAbout a month ago, I made a toy bot that listens to your voice with OpenAI Whisper, generates a response with GPT-2 and vocalizes the response using the Eleven Labs. The TTS quality produced by the Eleven Labs algorithm was mind-blowing to me. The API that they provided was super easy to use. Good product!
- monk1 4y agoGood to see that authors/maintainers of AI models are beginning to think about attribution. But it seems like this will be a hard problem to solve. For example, say my voice was part of the training data set, to what degree can I lay claim to the newly created voices? Also, will there be some sort of grading/ranking (e.g. it could be argued that some of the voices used in the training set are more desirable than others, and therefore their "owners" deserve better fees etc.)?
- afro88 4y agoHow long before we have a meta "This 'this X doesn't exist' doesn't exist"?
- firechickenbird 4y ago> severely underhyped: voice AI These two words made it all sound like they are just trying to ride the AI wave instead of actually solving a real world problem
- piotr11 4y agoHey - developers behind ElevenLabs here. Thank you so much for the constructive and positive feedback - we’re taking it onboard! We’re currently focused on researching and deploying a different way for speech synthesis that can generate nuanced intonation and emotions by understanding text and taking context into account. Additionally, we provide creators with a way to clone their own voice based on very short samples. With the published blog post, we are now deploying a way to help them design entirely new ones! Anyone will be able to generate that level of quality just with a copy-paste. We are planning to open up Beta later this month. Our goal is to let you convert any written content into high-quality, compelling audio. To address a few questions that frequently came up: - Latency for our streaming TTS is <1s with quality results available above, which is the usual problem with existing good TTS models (like tortoise-tts) - We can clone voices instantly, based just on 5s of speech, without training required - We are working on adding SSML-like support for better control; speed controls will be coming as part of that too - API is directly available as part of Beta; we are preparing the infrastructure to scale easily for the release! We are hiring researchers, frontend and full-stack developers! If you are interested, send over your GitHub account and short message to founders[at]elevenlabs.io.
- diminikolaou 4y agoHey Piotr - just wanted to say congratz for the awesome work so far man. The quality is genuinely unbelievable. I don't know if you guys are ready to take clients at scale, but I don't see any reason why all newsletter creators wouldn't use your tech right now to address whole new markets. I'll be following the journey, excited for what's to come.
- TheMrZZ 4y agoHi! Are your models english only, or do you plan on tackling other languages?
- piotr11 4y agoThey will be multi-lang, the tech scales to any language and we are working to add more (it is relatively easy). Here is the demo in Polish TTS: https://www.youtube.com/watch?v=ra8xFG3keSs https://www.youtube.com/watch?v=ra8xFG3keSs
- smusamashah 4y agoI have only ever listened to one audio book and that was "Hitchhiker's guide to the galaxy" by Stephen Fry. This is nowhere close to that. It does mimic the ups and downs of voice but they don't add up. The don't make sense. They don't really have any connections with what is being spoken. But since it can do expressions, it probably only needs special markers in text to tell it how to really read a sentence.
- singedproxy 4y agoStephen Fry is considered one of the best audiobook readers of all time. This AI voice is still better than 100% of AI audiobooks in the market, and likely better than a good portion of HUMAN readers as well.
- sebzim4500 4y agoThe conversational one doesn't sound like an AI but some of the emphasis is still a bit awkward. If I didn't know better I would have thought it was recorded by a person who was uncomfortable having their voice recorded. Still insanely impressive though.
- 2OEH8eoCRo0 4y agoThis, and tools like it, could revolutionize video game voice acting. Have any video game engines integrated tools like this so developers can use them?
- devops000 4y agoWhy someone should listen a voice when is faster to read the blog post ?
- pixl97 4y agoWhy are all people exactly the same? It's odd that the audio book business went out of business years ago because people only want to read. And yes this was a snarky answer because even if you don't realize it, it was a snarky question.
- glerk 4y agoAbsorbing information through your audio input while the visual input is busy with a mindless task is amazing. You can listen to an article or an audiobook while doing laundry.
- fireant 4y agoEveryone is different, I read rather slow and can ingest the article much quicker with audio (esp. sped up audio) than by reading it.
- angusturner 4y agoAs someone working on singing synthesis, I know how hard it is to get that last 10% quality that makes a human listener instantly recognise if the voice is real or generated. These are really impressive results! For anyone interested, my team’s singing work: https://youtu.be/LPy20zSWhZA https://youtu.be/LPy20zSWhZA)
- meremortals 4y agoVery well done! Any suggestions on where/how one might learn to do something similar? I love the idea of being able to swap singers on a given track
- imtringued 4y agoIf you are going to have such an intensive particle effect in your videos at least bother to upload a 4k version so there is a tiny chance that not every single frame consists of nothing but artifacts. Also don't put gumi and English in the same search query on YouTube. I don't know how they did it but the voices from six years ago sound better than SOTA TTS based on deep learning today...
- singedproxy 4y agoClearly the point of the video is its AUDIO content, not the visuals. The lack of a "4k version" does not make any difference other than saving you bandwith :-)
- lobo_tuerto 4y agoThe "conversational" example should be named "Karen".
- QuantumGood 4y agoMore advanced scams potentiated by technology advancements are an arms race hard to keep ahead of. Despite all the possible positives, this seems almost inherently dystopian.