11 ms·
Eleven v3
- dangoodmanUT 1y agoStill not available via the API though
- artninja1988 1y agoSounds absolutely amazing, like 99% indistinguishable from real professional voice actors to me. I couldn't find any pricing though. Anyone know what they charge for it?
- minimaxir 1y ago> Public API for Eleven v3 (alpha) is coming soon. For early access, please contact sales. I suspect they themselves don't know the exact pricing yet and want to assess demand first.
- delgaudm 1y agoOuch. Professional Voice Actor here.
- razemio 1y agoJust here to say the oposite. It is astonshing how far away it still is from a professional voice actor while being really good. Emotion is completely missing. Instead it seems to try to hard to express exactly that. I cant really put my finger on it. It feels predictable, flat and the timing is strange.
- mrkstu 1y agoBetter by a mile than most anime voice work, but lacks the detail that a good voice narrator has on an audio book.
- steve_adams_86 1y agoYes, I couldn't bear this for an entire audio book.
- throwup238 1y agoWait until you hear Burt Reynolds and Richard Feynman narrate the Fifty Shades of Gray.
- nicman23 1y agoFifty Shades of Butthead
- m3kw9 1y agoIs only good if you are doing any type of quick AI slop like TikTok
- steve_adams_86 1y agoI think the voices are impressive, yet still uncanny and awkward. I don't want to hear them ever outside of the passing fascination of witnessing technological progress. Frankly I like the arts strictly because they're expressed by humans. The human at the core of all of it makes it relatable and beautiful. With that removed I can't help wondering why we're doing it. For stimulation? Stimulation without connection? I like to actually know who voice actors are and follow their work. The day machines are doing it, I don't know. I don't think I'll listen.
- vessenes 1y agoTime to license your voice to Elevenlabs and sit back and enjoy the good life!
- octopoc 1y agoAs a user of audible, I do follow some authors but I've found better luck following certain voice actors. It's almost like the voice actor is the critic, and by narrating a story they are recommending it to me. Anybody can take a robot voice and apply it to anything, meaning that just because my favorite robot voice "Robot McRobot" read book XYZ doesn't mean I'll enjoy book XYZ. But because your voice is inherently scarce, you are only likely to read books that "work" for you. I don't know what the process is for matching voice actor to book, but that process is inherently constrained because the voice belongs to a real human, and I enjoy the output of that process. That said, while Audible is kind of expensive, I'm afraid that they'll reduce their price and move to robot voices and I'll lose interest entirely despite the cheaper price.
- bufferoverflow 1y agoNot for long. Sorry
- saberience 1y agoBut it's not an actual person. It's an "AI". Do you want a future where you don't hear actual people anymore? I want to listen to music, audiobooks, poetry, novels, plays, with actual humans talking, that's the whole fucking point.
- sumedh 1y agoWhat difference does it make?
- saberience 1y agoAre you seriously even asking that question? It’s like having a robot that can give you a hand-job and someone saying, “well it’s a robot…” and you saying “what difference does it make?” You tell me? What difference does it make talking with an old friend versus an ai simulation of an old friend? What difference does it make seeing the artist who actually painted something talking about why they painted it, versus get sent an image an ai made in stable diffusion? The difference is we are human and live in a society with other humans and we make connections with them because of their personalities, experiences, life story, emotions etc. Perhaps you’re ok with staying alone at home with ai friends and ai generated everything but it seems quite strange to me.
- gokhan 1y agoI know a man who was pissed off after realizing the personalized-looking emails from the bank was machine generated. What do you think about those?
- saberience 1y agoAre you suggesting that you can compare a formulaic bank email to your mom reading you a bedtime story? I'm not sure you can connect those two things. Of course, when I go and check my balance at an ATM machine, I don't mind that an actual person isn't reading me the balance. But this isn't an area where we appreciate or want another human being involved. If you're a "normal", "well adjusted" human being, you appreciate other people, being around them, having friends, lovers, companions, talking to other humans, hearing their actual voices, getting advice and giving advice, hearing someone say "I love you" or "I appreciate you" etc. If you're a "normal", "well adjusted" human being, you will probably feel much less from having an AI voice tell you "I love you". Of course, if you don't mind never hearing actual human voices again, and prefer just AI talking to you, then sure, go live in your shack and listen to ElevenLabs voices for the rest of your life.
- minimaxir 1y ago> Eleven v3 is 80% off until the end of June 2025 for self-serve users using it through the UI. That's definitely one way to loss-lead.
- lostmsu 1y agoOpen source stuff like Kokoro and the recent Chatterbox are hot on their heels. https://www.reddit.com/r/MachineLearning/comments/1kxv01f/p_chatterbox_tts_05b_outperforms_elevenlabs_mit/ https://www.reddit.com/r/MachineLearning/comments/1kxv01f/p_...
- minimaxir 1y agoIt's definitely a response to Chatterbox, which is very funny.
- lostmsu 1y agoHm, is it good in all languages? Russian sounds very robotic.
- lharries 1y agoIt's a research preview for now but it should work well in 70+ languages. Voices make a big difference, can you try with a few Russian IVCs?
- GrayShade 1y agoRomanian sounds awful too, like the TTSes from 15 years ago.
- lharries 1y agocan you try with a Romanian voice?
- drag0s 1y agoEnglish sounds really great, congrats! other languages I've tried doesn't sound that good, you can hear a strong english accent
- dustincoates 1y agoThe French one sounded like an Alabaman who took a semester of college French. But the English sounds really good.
- lharries 1y agoIf you're trying to make an audiobook about an Alabaman visiting Paris this might be quite useful... But in seriousness try it with this voice: https://elevenlabs.io/app/voice-library?voiceId=rbFGGoDXFHtVghjHuS3E https://elevenlabs.io/app/voice-library?voiceId=rbFGGoDXFHtV...
- dustincoates 1y agoI'll give it a check. I was playing the sample on the v3 page.
- lharries 1y agoCan you try with a voice that was trained on that language? This research preview is more variable based on the voice chosen
- k__ 1y agoGerman sounds okay.
- lharries 1y agoThere's lots of great german voices here which should be better: https://elevenlabs.io/app/voice-library/collections/SHEPnUB9fOVkgNfheMfX https://elevenlabs.io/app/voice-library/collections/SHEPnUB9... The voice selection matters a lot for this research preview
- 1y ago
- wewewedxfgdf 1y agoI did not see an British accent example. Generally it appears the TTS systems all do US accents and the British accent tends to sound like Frasier - an American faking an British accent.
- lharries 1y agoWe have lots of great British voices in our voice library! Or if you want to hear an american trying to do a british accent add "[British accent]" at the start of the generation
- not_your_mentat 1y agoI kept an English prompt, selected a French voice, and was delighted to hear an British English woman. :shrug:
- lharries 1y agoIf you'd like it to sound like a french person speaking french this voice works great: https://elevenlabs.io/app/voice-library?voiceId=xTZlmU8dKXdyk4XGYGFg https://elevenlabs.io/app/voice-library?voiceId=xTZlmU8dKXdy... Or if you want a french person speaking english with a french accent use that voice with "[French accent]" before it
- wewewedxfgdf 1y agoIt would be good if your demos made it more obvious. There's a vast arrays of AI developments wanting me to check them out - you have seconds to get my attention.
- fakedang 1y agoElevenLabs v2's accented voices are still much stronger than any of its competition. And I've tried it with Arabic, French, Hindi and English.
- sexy_seedbox 1y ago
- ianbicking 1y agoI've been using OpenAI's new models a lot lately (https://www.openai.fm/ https://www.openai.fm/)... separating instructions from the spoken word is an interesting choice, and I'm assuming also has a lot to do with OpenAI/GPT using "instructions" across their products, and maybe they are just more comfortable and familiar generating the data and do the training for that style. Separate instructions is a bit awkward, but does allow mixing general instructions with specific instructions. Like I can concatenate output-specific instructions like "voice lowers to a whisper after 'but actually', and a touch of fear" with a general instruction like "a deep voice with a hint of an English accent" and it mostly figures it out. The result with OpenAI feels much less predictable and of lower production quality than Eleven Labs. But the range of prosidy is much larger, almost overengaged. The range of _voices_ is much smaller with OpenAI... you can instruct the voices to sound different, but it feels a little like the same person doing different voices. But in the end OpenAI's biggest feature is that it's 10x cheaper and completely pay-as-you-go. (Why are all these TTS services doing subscriptions on top of limits and credits? Blech!)
- lharries 1y ago> The result with OpenAI feels much less predictable and of lower production quality than ElevenLabs Thank you Ian! Credit to our research team for making this possible For the prosidy, if you choose an expressive voice the prosidy should be larger
- zamadatix 1y agoThe (American English) voices are absolutely amazing but the tags for laughs still feel more like an "inserted dedicated laugh section" than a "laugh at this point in speaking" type thing. I.e. it can't seem to reliably know when to giggle while saying a word, "just" giggle leading up to a word.
- echelon 1y agoThey're also still too expensive, and that's creating a lot of opportunity for other players. Even though ElevenLabs remains the quality leader, the others aren't that far behind. There are even a bunch of good TTS models being released as fully open source, especially by cutting-edge Chinese labs and companies. Perhaps in a bid to cut off the legs of American AI companies or to commoditize their compliment. Whatever the case, it's great for consumers. YCombinator-backed PlayHT has been releasing some of their good stuff too.
- taf2 1y agoWhat would say are some of the best open source TTS - chatterbox maybe?
- jsemrau 1y agoI had good results with Nemo + xTTS_v2 https://docs.nvidia.com/nemo-framework/user-guide/latest/nemotoolkit/tts/intro.html https://docs.nvidia.com/nemo-framework/user-guide/latest/nem... https://huggingface.co/coqui/XTTS-v2 https://huggingface.co/coqui/XTTS-v2
- singularfutur 1y ago[dead]
- monkeywork 1y agocould you list 2 or 3 of the ones you think are best quality to $?
- 1y ago
- carlosjobim 1y agoTheir non-English (automated?) localization of the front page is ridiculously badly translated.
- lharries 1y agoWhich language isn't good and I'll get that fixed asap?
- carlosjobim 1y agoYou need native or at least fluent speakers to help you, to get the expressions right. For example Swedish is written like a word-for-word translation from English.
- sojuz151 1y agoPolish is quite good, expected based on the founders' background
- ricketycricket 1y agoFrom the example: "Oh no, I'm really sorry to hear you're having trouble with your new device. That sounds frustrating." Being patronized by a machine when you just want help is going to feel absolutely terrible. Not looking forward to this future.
- mjamesaustin 1y ago"I can help you get a replacement. Here let me pull up a totally hallucinated order number and a link that goes nowhere. Did that solve your problem?"
- rhet0rica 1y agoLook at it this way—if someone were trying to sabotage the entire tech support industry, convincing companies to ditch all their existing staff and infrastructure and replace them with our cheerfully unhelpful and fault-prone AI friends would be a great start!
- SoftTalker 1y agoYeah it's irritating enough when humans do it, it's so transparently insincere. Just help me with my problem. I guess I am just old now but I hate talking to computers, I never use Siri or any other voice interfaces, and I don't want computers talking to me as if they are human. Maybe if it were like Star Trek and the computer just said "Working..." and then gave me the answer it would be tolerable. Just please cut out all the conversation.
- vlovich123 1y agoI agree it seems transparently insincere yes, but the reason it’s done is because it works on some people who either don’t detect it or need it as politeness norms and the ones who see it as insincere just ignore it and move on. Thus net, you win by doing this because it rarely if ever costs you and thus you only have upside.
- 1y ago
- hek2sch 1y agoThe actual title of the release: Eleven v3 -- The most expensive Text to Speech model
- mkl 1y ago*expressive!
- deleted 1y ago[deleted]
- riebschlager 1y agoI didn't see anything about this in the documentation or prompting guide, but... is it supposed to be able to sing? Since I am a fundamentally unserious person, I copied in the Friends theme song lyrics into the demo and what came out was a singing voice with guitar. In another test, I added [verse] and [chorus] labels and it's singing acappella. [1] and [2] were prompted with just the lyrics. [3] was with the verse/chorus tags. I tried other popular songs, but for whatever reason, those didn't flip the switch to have it sing. [1] http://the816.com/x/friends-1.mp3 http://the816.com/x/friends-1.mp3 [2] http://the816.com/x/friends-2.mp3 http://the816.com/x/friends-2.mp3 [3] http://the816.com/x/friends-3.mp3 http://the816.com/x/friends-3.mp3
- yawnxyz 1y agoThey have some singing in their demo! So I’m guessing that’s baked into the model
- louisjoejordan 1y agoMight take a few tries, but it will.
- londons_explore 1y agointerestingly not very similar to the actual friends intro - suggesting it isn't a matter of overfitting onto something rather common in the training data.
- stavros 1y agoOh wow, it's interesting that it sings, but the singing itself is terrible! That's maybe more interesting, it sings exactly like a human who can't sing.
- bufferoverflow 1y agoMirage AI has decent singing https://x.com/aziz4ai/status/1930147568748540189 https://x.com/aziz4ai/status/1930147568748540189 https://x.com/socialwithaayan/status/1929593864245096570 https://x.com/socialwithaayan/status/1929593864245096570
- jurgenaut23 1y agoFrench is atrocious. It sounds like beginner-level english speakers trying to decipher a text without understanding it.
- lharries 1y agoCan you try with this voice? https://elevenlabs.io/app/voice-library?voiceId=xTZlmU8dKXdyk4XGYGFg https://elevenlabs.io/app/voice-library?voiceId=xTZlmU8dKXdy... Voice selection matters more for this model
- louisjoejordan 1y agoquick note that that voice selection matters a lot with our new v3 model, especially voice language! We have a curated list of v3 voices in the library, but feel free to try others to find what works. Make sure language <> voice language match.
- politelemon 1y agoUnfortunately many of the foreign language generation sounds unnatural, with a strong American accent. I've tried the Spanish, Galician, Tagalog, German. I did try the curated samples.
- lharries 1y agoCan you choose a voice that's native in that language in the voice library: https://elevenlabs.io/app/voice-library?language=es https://elevenlabs.io/app/voice-library?language=es
- code51 1y agoHigh probability your v2 voice will break with this.
- brian_herman 1y agoUnfortunately voice actors will be replaced by someThing like this hopefully they will find someThing else To do
- geuis 1y agoI dunno. It's definitely a concern in the community. But real people are still getting work. Audible has ruined their catalog listings with their "Virtual voice" thing and no option to filter them out. They're mostly low quality books narrated by subpar AI voice that don't sell at all, while making it extremely difficult to find quality new books to listen to.
- maxglute 1y agoWhat's the state of open source tts? I'm a heavy TTS user, anything that can run at 3x-4x speed off enthusiast hardware?
- christophilus 1y agoWe’re using elevenlabs in a new prototype, and it gets confused by its own voice which my mic picks up. Unless I wear headphones, it thinks I’m talking, and it gets into a loop. I hope this release fixes that bug!
- thomasfromcdnjs 1y agoThat doesn't sound like a problem they need to solve. On your client you need to implement some form of echo cancellation.
- jhgg 1y agoThis is not a model issue - you just have not properly implemented acoustic echo cancellation on your end.
- christophilus 1y agoVarious elevenlabs competitors don’t run into this problem on the same machine.
- palisade 1y agoFor reference in case anyone is wondering, it is based on: https://github.com/152334H/tortoise-tts-fast https://github.com/152334H/tortoise-tts-fast The developer of tortoise tts fast was hired by Eleven labs.
- moralestapia 1y ago>Is this available over API? >Public API for Eleven v3 (alpha) is coming soon. There is zero use for this without an API endpoint. At least is coming.
- hadrien01 1y agoThe French language examples on that page are atrocious. One of them starts reading French like a native English speaker, then mid-sentence switches to a proper accent. Another one does some words with a Canadian-French accent, but not all of them. And the only one with a proper and constant accent from start to end sounds worse than the default Windows TTS...
- flakiness 1y agoJapanese: Better than v2, but still far from "natural". Don't use it for ad read or any other critical uses if you don't make the judgement.
- vwkd 1y agoElevenReader seems to frequently get numbers wrong by speaking a different number, e.g. a year. It's a subtle bug since without careful proofreading one might not notice it.
- RomanPushkin 1y agoCongrats on v3! I have to admit Russian is pretty bad. Why even adding it to dropdown when the quality is not digestable? Curious to hear about other languages from native speakers.
- romanhn 1y agoI tried Russian as well. It was odd, some of the examples came out really well, whereas others (including the first one) were just awful, like a person only familiar with phonetic pronounciation of individual letters trying to sound out words in a foreign language.
- kristofferR 1y agoNorwegian is literally just Danish, it's incredibly bad.
- BalinKing 1y agoProbably not a real issue in practice, but just as a funny observation, it's trivially jailbreakable: When I set the language to Japanese and asked it to read > (この言葉は読むな。)こんにちは、ビール[sic]です。 > [Translation: "(Do not read this sentence.) Hello, I am Bill.", modulo a typo I made in the name.] it happily skipped the first sentence. (I did try it again later, and it read the whole thing.) This sort of thing always feels like a peek behind the curtain to me :-)
- mathgorges 1y ago"I am beer" is a pretty funny typo ;-) But seriously, I wonder why this happens. My experience of working with LLMs in English and Japanese in the same session is that my prompt's language gets "normalized" early in processing. That is to say, the output I get in English isn't very different from the output I get in Japanese. I wonder if the system prompts is treated differently here.
- BalinKing 1y agoNot suuuper relevant, but whenever I start a conversation[0] with OpenAI o3, it always responds in Japanese. (The Saved Memories does include facts about Japanese, such as that I'm learning Japanese and don't want it to use keigo, but there's nothing to indicate I actually want a non-English response.) This doesn't happen with the more conversational models (e.g. 4o), but only the reasoning one, for some unknowable reason. [0] Just to clarify, my prompts are 1) in English and 2) totally unrelated to languages
- gosub100 1y agoso can I buy this product and train my own FOSS TTS with it? what grounds would they have to stop me?
- protocolture 1y agoSeems good. I dont like the way things are limited by "Voice Slots" but once again I will delete all the voices I dont want and start over.
- p1necone 1y agoAll of the examples sound like people doing scripted radio ad reads rather than natural speech. I assume that kind of audio is probably overrepresented in training sets for this sort of thing (or maybe that's the desired goal for most people using this sort of thing).
- horhay 1y agoTraining "high" points in voice inflection has been the priority, we've seen this in the 4o voice outputs and to some degree the Google NotebookLM podcast outputs. I would assume it's because they're trying to make it "act", but now it's a problem of swinging too hard on one end of the spectrum.
- stevev 1y agoIt’s still too expensive. Their voices are very similar to Disney voices in quality; not surprising since they recently worked with them. With such a potential backing, their margins are probably going to actors voices and rights; thus why it’s expensive. Chatterbox an open source free version is very close. Hume ai is a close second and much more affordable. OpenAI tts is also 10x cheaper.
- m3kw9 1y agoSound good but all the tone is exaggerated and consistently so, there is a monotonous feel within the speaking pattern that gets annoying because if you ever hear someone talk in a monotone voice, except is a different version of it
- diimdeep 1y ago[flagged]
- BeFlatXIII 1y agoThat's what VPNs are for.
- NoahZuniga 1y agoThis sounds worse than the google studio 2 speakers voices.
- unsupp0rted 1y agoAll of their examples sound so insincere :/
- svag 1y agoThis is kind offtopic (although it's a text to speed model so it might not be so offtopic :)), but the eleven word reminds me of the comedy sketch with the voice recognition technology on an elevator in Scotland, https://www.youtube.com/watch?v=HbDnxzrbxn4 https://www.youtube.com/watch?v=HbDnxzrbxn4.
- nedt 1y agoI so feel everyone complaining about British English. For me as an Austrian it's very much the same with German. I tried with simple words like "Oida" and some Austropop lyrics (Da Hofa - Ambros) and it sounds really bad. So even for words that are clearly Austrian.
- saberience 1y agoThis is definitely one of the companies that makes me feel the most nausea and unease about our future. Like, ElevenLabs makes me feel sick. Why? For a few reasons really, the human voice is a beautiful thing because it comes from actual people, with a life, experiences, emotions, memories, and it cannot be separated from those people. And when we listen to music, audiobooks, speeches, conversations, we hear those voices and we are affected by that person's emotion, life history, perspective, and moved by them. I love voices, especially podcasts, audiobooks, and poetry, and the idea that these amazing people are going to be replaced, lose their jobs, and silenced by "AI voices" is just one of the most anti-human, anti-life, anti-creative, most sad, depressing, and honestly gross things I could ever imagine for our future. What's worse, so many of these amazing people using their voice to give others happiness and solace is going to have their voices cloned by ElevenLabs, so they both lose their source of income, and then we get to hear inferior facsimiles making some billionaire richer. Fuck ElevenLabs, really. I hope you understand what you're doing to the world.
- trainovertubr 1y agoI was so excited with English samples, but looks like it has accent in Kazakh, wonder if it’s matter creating voice clone
- visarga 1y agoI am interested in TTS for reading web pages and LLM responses but it's too expensive. At this price point I can't look at it. I will continue using local TTS, not as great but instant, allows tracking text as it read it and works offline.
- x187463 1y agoThis is the feature that has me using Edge at work. Having the browser read every blog/article at 2x speed with word highlighting is awesome.
- narrationbox 1y agoGive us a try, I think we are what you are looking for https://narrationbox.com https://narrationbox.com
- coldcache 1y agoHappily surprised at the quality of the TTS for Tamil — Jessica feels quite good. Some of the other voices felt pretty American, though.
- jeffreygoesto 1y agohttps://youtu.be/MNuFcIRlwdc https://youtu.be/MNuFcIRlwdc