10 ms·
Emotionally Expressive Text to Speech
- maxdo 6y agoWow sounds very real
- diminish 6y agoImpressive next step for text-to-speect. Wish there was some simple real demos. I also work on the same thing using DL- and hope to open source the "emotional part" of it. We soon can create emotionally expressive youtube videos with synthetic actors..
- sonantic 6y agoThanks for your comments and nice to hear you're also working on TTS! We have a few more samples (without background music) further down on our homepage and plan to add a full dedicated subpage in time!
- microtherion 6y agoThe prosody sounds nice. But two of the longer samples have a lot of vocal fry, and the third sounds like the voice has a stuffy nose and/or a slight lisp. I wonder whether those mannerisms were chosen to camouflage artifacts inherent in their current implementation.
- rowanG077 6y agoOr the mannerisms where chosen to express that the system can produce real voices and not just perfect ones.
- microtherion 6y agoYeah, but wouldn't you include at least one "clean" voice in the samples to show what the system is capable of?
- sonantic 6y agoYep. Each of our TTS models is based on a real actor's voice that has its own nuanced characteristics. Some voices are naturally rougher / croaky while others smooth. As in real life, our differences are what make us unique. Some voices will work better for certain character profiles / scenes - it's up to the user to decide.
- ArneVogel 6y agoThis site has to have one of the worst cookie choice decision popups: https://imgur.com/a/YLsGadP https://imgur.com/a/YLsGadP
- aantix 6y agoGDPR has reverted the web to Geocities. It’s popup hell.
- martimarkov 6y agoIt’s not GDPR. It’s all the tracking website owners put there. The fault is with the owners of the sites.
- deleted 6y ago[deleted]
- trickstra 6y agoThey don't HAVE to serve cookies on the homepage. It's like complaining that your carbon monoxide alarm is way too loud and beeps too often.
- sonantic 6y agoTouché! It is pretty bad, huh? Our new site is only a few days old (rolled out on Wednesday / around the same time as our demo) so we haven't got eveything optimised just yet. Stay tuned for better cookies in the future!!!
- sergeykish 6y agolike option does not really matter
- ad404b8a372f2b9 6y agoAnd the privacy policy page returns 404.
- 6y ago
- amelius 6y agoSounds nice but difficult to judge with the background music.
- hobofan 6y agoI get the criticism (and "you need to find out for yourself" at ~0:44 sounds somewhat robotic), but given that it's aiming at the entertainment industry, where you will have background music most of the time, it also seems like a fair choice of representing real-world usage (where background music might always hide the imperfections a bit).
- DenisM 6y agoThis is very impressive. I wonder if attaching this to a modern-day Elisa will improve the Turing test scores? Emotional load can reduce the requirement for semantic coherence.
- samcodes 6y ago“Emotional load can reduce the requirement for semantic coherence.” - great insight
- spaceprison 6y agoMy daughter is dyslexic and would love to play things like stardew valley, pokemon or even animal crossing but being text only makes them such a slog for her. The same goes for sub titles, she'd be perfectly fine with a robot voice for the actors if they sounded real enough like this. Game changer.
- londons_explore 6y agoWhy not just read the text to her?
- billme 6y agoFor the rest of her life?
- sonantic 6y agoThank you for your comment. One of the reasons we founded Sonantic was to improve accessibility so we are right there with you! We plan to do this by reducing the barriers (both financial and logistical) of voiced content for everyone from indie developers to big AAA studios. We've already begun to see progress on this through partnership with initial customers during our beta.
- mrec 6y agoIs your TTS process necessarily ahead of time, or can it be done at runtime with all the flexibility (templating, generative text etc) that brings?
- sonantic 6y agoThat's the holy grail right there, isn't it? :) We're definitely working towards runtime but still some work to be done there to account for additional complexities and balance trade-offs re: speed, quality, accuracy etc. of the rendered output.
- disabled 6y ago
- schoolornot 6y agoBetween this and Lyrebird there seem to be a high number of cutting edge TTS solutions being worked on in the private sector. Does anyone know why there haven't been much advancement with the FOSS libraries?
- vianneychevalie 6y agoI’m convinced that the barrier to entry in this field in terms of technologic and financial investment is too high for FOSS projects to compete with the commercial solutions We don’t see FOSS pharmaceutical research for instance, I believe for the same reason. The amount of coordination needed and the impossibility to separate TTS projects into sub-parts could also factors.
- sgk284 6y agohttps://voice.mozilla.org/en https://voice.mozilla.org/en “Common Voice is Mozilla's initiative to help teach machines how real people speak.”
- nmfisher 6y agoThe problem with Common Voice is the same resource problem plaguing open source efforts in general - for TTS data, it's not just the size that matters, it's also the quality. For something like Sonantic, you need clean recordings from professional actors in proper recording environments (not to mention the in-house expertise to then filter these down to curate the training/test datasets). That costs money. A million people with laptop microphones will just never get there.
- ani-ani 6y agoSure, open source TTS seems to be lagging behind recent commercial offerings, but pharmaceutical research is actually an excellent example of a field with massive FOSS software usage. TTS is also "purely virtual" which makes it significantly different, and I would say significantly more approachable to open source collaboration.
- 6y ago
- blattimwind 6y agoI could see this being used for RPG games to fix the choice deficiency that has been caused by going for fully voiced dialogue. Also, making Hitler read copypastas even more convincingly.
- ghaff 6y agoAs good as even not-top-shelf voice actor talent is a really high bar. I keep my eye on this space because there are a number of things I do where having even just decent "radio voice" TTS would be useful (and better than I can do myself). But nothing is really there today. In some respects, it's better than I can do myself but certainly not consistently.
- blattimwind 6y agoThe bar really isn't "has to be good out of the box", if it requires some tweaking on a line to line basis that would probably be ok and still much, much cheaper and much quicker to iterate on than voice actors for these high volumes of speech. In a lot of these games the existing voice acting is often consistently poor (literally everything Bethesda ever released comes to mind); certainly quite a few notches below the average AAA voice acting (which is occasionally bad, but on average good).
- ghaff 6y agoFair enough. I'm not really much of a gamer. One of the other challenges with using outside voice talent is that it can be inconvenient/expensive when you need to add/change something. I've been involved with podcasts using an external host and one of the negatives with that process is that if you discover a minor mistake/glitch in the narration late in the process you can't easily fix it.
- nmfisher 6y agoGood point. The samples still have that slight gravelly/metallic timbre that many TTS synthesis systems have. But then I remembered voice acting fluctuates wildly in quality outside AAA games. I'd take well-acted (with rough audio quality) over poorly-acted (but high-audio-quality) voices for most games.
- jariel 6y agoRecommending editing the video down to 43-60 seconds. It would be nice to try with actual text inputs right on the page, that this doesn't exist is tiny flag. A great choice to work with voice actors, because there isn't any 'pure' TTY that's good enough in the most general sense, having the actual voice actor as a working basis will help. Perhaps for small game houses, they can just use something off the shelf, big houses can use a customized voice, and then not worry if they have to make tweaks or changes, they don't have to do a whole production.
- sonantic 6y agoThanks for your feedback! We felt that this storyline / length was best in order to showcase the two different actor's artificial voices and build up to the actual cry. As you've mentioned, we do work with real actors to create our TTS and take misuse of their (artificial) voices very seriously. Because they sound so lifelike, we've made the decision not to allow public access/personal use at this time. Lastly, your assessment is spot on regarding standard vs custom voices. Lots of interest for both!
- zoomablemind 6y agoTTS=text-to-speech, so it's quite reasonable to showcase that chain instead of an edited video. Not diminishing the quality of your product, just pointing out an obvious expectation of the audience that it's presented to. Perhaps, there could be a way to test-drive it directly, with limited choices or combinations of the input text.
- sonantic 6y agoFair point. We will definitely consider adding something like this to the site. Thanks for the suggestion!
- sonantic 6y agoHey HN - Zeena Qureshi (Co-Founder and CEO at Sonantic) here. Thanks for your thoughts and feedback thus far! I'd be happy to answer questions (within reason) about our latest cry demo / emotional TTS! Feel free to fire away on this thread.
- netcan 6y agoYou have the perfect name for someone making emotive machines. Would you do demos for well known speeches/texts? It'd be easier to put this into context that way.
- sonantic 6y agoThank you! We think our name is pretty great too :) Sonantic only creates AI voice models with the consent of the artist/voice owner. We take misuse and copyright infringement very seriously and therefore never train on data (voice recordings) where the original source is unaware of its repurpose. That said, we aim to include more recognisable voices on our platform in the future.
- sytelus 6y agoIs there really a copyright issue? Mimicry artists and tribute bands have duplicated original voices without a problem. Technically, you are not reproducing any recording owned by anyone. It's brand new synthesis! It would be interesting to see if any courts rules this is not the case. You can own your spoken speech, but can you really own spectrogram of wave patterns?
- Hydraulix989 6y agoI'm sure as a startup, that's not a legal battle that is worth the risk. Judiciary activism is not worth it for them.
- sonantic 6y ago:)
- deleted 6y ago[deleted]
- cemregr 6y agoIs there an actual demo?
- Animats 6y agoCan't hear the voices over the music.
- catblast 6y agoI hear them if I concentrate to isolate the voices out of the music. And you can pick up on quite noticeable flaws, especially in cadence and intonation. What somebody described as vocal fry sounds more like synthesis artifacts. The bare samples further on the page also highlight issues in cadence. This is obviously an early demo, but this isn't yet to the level you could narrate an audio book - those little problems will quickly become noticeable.
- moron4hire 6y agoAny plans to support languages other than English? This would be huge in the foreign language instruction field.
- sonantic 6y agoYes! Supporting additional languages (and dialects) is definitely in our roadmap.
- crazygringo 6y agoThis is fascinating. But I'm very curious what the emotional "parameters" are? There are literally at least a thousand different ways of saying "I love you" (serious-romantic, throwaway to a buddy, reassuring a scared child, sarcastic, choking up, full of gratitude, irritated, self-questioning, dismissive, etc. ad finitum). Anyone who's worked as an actor and done script analysis knows there are 100's of variables that go into a line reading. Just three words, by themselves, can communicate roughly an entire paragraph's worth of meaning solely by the exact way they're said -- which is one of the things that makes acting, and directing actors, such a rewarding challenge. Obviously it's far too complex to infer from text alone. So curious how the team has simplified it? What are emotional dimensions that you can specify? And how did they choose those dimensions over others? Are they geared towards the kind of "everyday" expression in a normal conversation between friends, or towards the more "dramatic" or "high comedy" of intense situations that much of film and TV lean towards?
- sergeykish 6y agoI hear same expression with different "strength". There is no play. No motion. Expression should change after response. It does not. There is no dialog. For me it sounds bald, boring. It'd better not to participate in such dialog. We can express emotions without words: xxx: Distress yyy: Support xxx: Hope It maps on music and we have dictionary to describe it. The one I'm listening to is Sorrow and Hopeful - entire track. May be a good start. Write first (classification). Examples you gave I feel live on same scale but extreme values. So even harder. I'd imagine it work like autotune - enhance human input
- stubish 6y agoI've always imagined that this tech would need a markup language. Instead of a script that an actor needs to interpret, the script writer (or an editor, or a translator) would mark up the text.
- jchb 6y agoThere is Speech Synthesis Markup Language (SSML). Amazon Polly and Google text-to-speech supports it, although the best neural-model based voices only support a small subset.
- voiper1 6y agoIs there any pay-to-use or open source voice for Hebrew? Amazon's Polly English voice, Matthew is pretty nice. But they don't have Hebrew. Also Google doesn't have Hebrew. Bing has some attribution requirement that I haven't fully investigated.
- sarabande 6y agoHere is one, עלמה רידר https://www.almareader.com/ https://www.almareader.com/ (not nearly as emotionally expressive as this one though).
- yorwba 6y agoIf you're okay with very robotic speech, espeak-NG has experimental support for Hebrew since last month: https://github.com/espeak-ng/espeak-ng/issues/732 https://github.com/espeak-ng/espeak-ng/issues/732
- sarabande 6y agoIf this could generate well-done audiobooks instantly from a text, that would be fantastic. All e-books could have an audiobook version overnight.
- woah 6y agoIt's kind of odd that they are not pitching this as the primary use case. Seems much more plug and play than game voicing
- mthoms 6y agoIt needs a human to annotate the text with the desired emotion. Ideally, it would be able to infer the emotion from the text itself, but I think that level of sophistication is a long way off. Edit: Actually, this might be a perfect candidate for some sort of crowdsourcing. Imagine Wikipedia pages containing hidden annotations for the proper text-to-speech "tone/cadence/whatever" of each sentence or paragraph.
- bnj 6y agoThis is a really cool idea— amazing to think about something like goodreads taken to the next level where people are sharing their emotion markups for texts. Imagine how you could try different mappings to see which ones you liked...
- sergeykish 6y ago<curios>how would it know?</curios> <dismissive>how would it know?</dismissive> <sorrow>how would it know?</sorrow> <angry>how would it know?</angry>
- dequalant 6y agoThis is amazing! I was looking something like this to come up for a long time. Finally someone did it!
- tomByrer 6y ago@sonantic Seems you don't do real-time yet? If so, have plans for a Web Speech API plugin? I'm about to release a reader demo based around it. https://developer.mozilla.org/en-US/docs/Web/API/Web_Speech_API https://developer.mozilla.org/en-US/docs/Web/API/Web_Speech_...
- yc-kraln 6y agoI have a comment and a question: The comment: I noticed that your demo video also had "emotional" video layered on top of the dialogue. This could be considered manipulative; perhaps consider sharing a naked version so we could attempt to interpret the emotion based solely on the text to speech engine. The question: You mention you met at EF. I was wondering if, beyond bringing you together, you found EF to be worth the cost of admission?
- jasonlingx 6y ago> The comment: I noticed that your demo video also had "emotional" video layered on top of the dialogue. This could be considered manipulative; perhaps consider sharing a naked version so we could attempt to interpret the emotion based solely on the text to speech engine. Close your eyes
- fossuser 6y agoThe music is still there and has an obvious effect. I thought the demo was impressive, but these things do seem like an effort to distract from (or more accurately bolster the effect of) the core technology. Though maybe the right call since this is less a strict technical demo and more a way to drive interest/marketing. The 'high levels of expressivity' comment was more of a flag to me, it's a meaningless phrase alone but it's suggested as an obvious answer. It feels like a mysterious answer [0]. I recognize though this is a marketing video, the core tech demo is cool, and I'm probably being unfairly critical. Flags like that make me more skeptical than I would otherwise be by default. [0]: https://www.lesswrong.com/s/5uZQHpecjn7955faL/p/6i3zToomS86oj9bS6 https://www.lesswrong.com/s/5uZQHpecjn7955faL/p/6i3zToomS86o...
- chrismorgan 6y agoAnd the music. Can’t remove that as easily as the picture.
- sonantic 6y agoHi there - thanks for both your comment and question! As others have said, this is first and foremost a marketing video aimed at attracting target customers. We've got additional clean samples (without background music) further down on our homepage and we plan on adding even more on their own subpage of the site in the future. We've also done a few technical demos at conferences over the past year and will continue to do so. We did meet at EF and it was totally worth it! There is no cost of admission for EF, they actually pay you to complete the program! Granted the monetary funds they provide could be a heck of a lot less than what you are earning at a full time job, so everyone's opportunity cost is different. EF's biggest selling point is their world-class network of highly ambitious individuals, so if you're interested in founding a company (pre-team / pre-idea) I would absolutely recommend looking into it.
- hyperpallium 6y agothe video https://youtube.com/watch?v=zwYiDraKtSA https://youtube.com/watch?v=zwYiDraKtSA
- vessenes 6y agoHi Zeena, I love this! I just filled out your form. I was just mucking around with Nvidia's latest, called flowtron, and I know from that experience there's a significant amount of work between getting a tech demo out and launching a usable product, whether API-based, or with some visual workflow like your video shows. One thing I think worth considering on the commercialization front is whether or not the core offering is the workflow niceties around your engine, the engine-as-API, or both. I'm just a random person on the internet, so take these thoughts with a large grain of salt, but thinking about it, it seems to me that prioritizing integration with say unity, unreal engine, video compositing tools, blog posting tools are all interesting and viable market paths. The underlying networks are going to keep improving for some time, so you're really trying to buy some long term customers. Some stuff that's obvious, but I can't resist: I could off the top of my head imagine using this for massively reducing the cost to develop games, for script writers pulling comps together, for myself to create audio versions of my own writing, for better IOT applications inside the home... I'd really love to be able to play with this. There still isn't a truly non-annoying virtual assistant voice; when the first tacotron paper came out, I was hopeful I would see more prosody embedded in assistants by now, but the longer we live with siri and google, the more sensitive I think we are to their shortcomings. I have a preference for passive / ambient communication and updates, so I would place a really high value on something that could politely interrupt or say hello with information. At any rate, congratulations, this is cool. :)
- microcolonel 6y agoVery cool demo, but the quality of the vocoding is not state of the art, and it's audibly artificial, which is probably why you covered it up with the obnoxiously loud music. Next time be honest about what you have when presenting it; every human with functioning ears is attuned to the sound of speech. This sort of technology would be amazing for narrative video games even with the less than perfect vocoding.
- aasasd 6y agoAs a non-native-speaker, I understood exactly four words from the monologue in the vid. Which might be on par for some movies, often having actors whisper and breathy-voice through the whole thing (ahem House of Cards cough). However, for actual TTS like webpages and audiobooks, the ‘Dina’ voice works much better.
- terrycody 6y agoApplied the form. Really cool. I want the know the price and when can we use it in production.
- deleted 6y ago[deleted]
- dejongh 6y agoBorh Cool and creepy!
- sonantic 6y agoO_o haha thanks, we'll take that as a compliment!
- diskmuncher 6y agoHistory has shown us that technological advancement of this kind will be adopted first by ... Obvious application: H-anime. Reduced parameters for the "emotion" as well.
- wishinghand 6y agoHey Zeena- will there be options to make the voices more unreal? The use case I imagine is for a character with a damaged vocoder or a broken speaker. Other glitchy affectations could be useful too.