16 ms·
NaturalSpeech: End-to-end text to speech synthesis with human-level quality
- hrdwdmrbl 4y ago[Take my money meme] I want this for articles and books now, please
- lvl102 4y agoMicrosoft/Nuance has been doing great in this area. I am very impressed with TTS on Windows. It makes proofing documents that much easier. I do think there is a need for some type of markup (akin to sheet music) for supervised learning.
- vtts 4y agoThe Text-To-Speech service by https://vtts.xyz https://vtts.xyz is the perfect choice for anyone who needs an instant human sounding voiceover for their commercial or non-commercial projects. Got a product to sell online? Why not transform your boring text into a natural sounding voiceover and impress your customers. What about adding a voiceover to your animation or instructional video? It will make it sound more professional and engaging! Our human sounding voices add inflections in the voice that make them sound natural, and our custom text editor makes it easy to get exactly what you want from both Male & Female voices included over 30 different tones, including: Serious, Joyful & normal
- vtts 4y agoDo you want to check how it works? you can test the operation of standard voices and advanced neural voices, at this url: https://vtts.xyz/home/tryme https://vtts.xyz/home/tryme
- webmaven 4y agoWowsers. This is a step change in quality compared to SOTA. I suspect that without evaluating samples as a correlated group, distinguishing between the generated samples and those recorded from a human will be little better than a coin toss. And even when evaluating these samples as a group, I may be imagining the distinctions I am drawing from a relatively small selection that might be cherry-picked. Nevertheless: The generated samples are more consistent as a group, and more even in quality, with few instances of emphasis that seem (however slightly) out of place. The recorded human samples vary more between samples (by which I mean the sample as a whole may be emphasized with a bit of extra stress or a small raising or lowering of tone compared to the other samples), and within the sample there is a bit more emphasis on a word or two or slight variance in the length of pauses, mostly appropriate in context (as in, it is similar to what I, a non-professional[0], would have emphasized if I were being asked to record these). In general for a non-dramatic voiceover you want to maintain consistency between passages (especially if they may be heard out of order) without completely flattening the in-passage variation, but tastes vary. Conclusion: For many types of voice work, these generated samples are comparable in quality or slightly superior to recordings of an average professional. For semi-dramatic contexts (eg. audiobooks) the generated samples are firmly in the "more than good enough" zone, more or less comparable to a typical narrator who doesn't "act" as part of their reading. [0] Decades ago in Los Angeles I tried my hand at voiceover and voice acting work, but gave up when it quickly became clear that being even slightly prone to stuffy noses, tonsillitis and sore throats was going to pose a major obstacle to being considered reliable unless I was willing to regularly use decongestants, expectorants, and the like.
- axiolite 4y ago> distinguishing between the generated samples and those recorded from a human will be little better than a coin toss. The synthesis still doesn't know where to place emphasis. You may be unable to distinguish between POOR human voice work and TTS, but not GOOD human reading. Not a big improvement over some existing TTS. Try e.g. Michelle at: https://cloudpolly.berkine.space/ https://cloudpolly.berkine.space/
- deleted 4y ago[deleted]
- ebjaas_2022 4y agoYou can hear that the human readers place emphasis based upon an understanding of the meaning of the text that they're reading, and also based upon an understanding of the humans at the receiving end. It seeps through that they're human. The AI generated samples are good, but they're bland in comparison. The human emphasized words are typically not emphasized in the AI generated samples.
- SemanticStrengh 4y agoI'd be nice to be able to tag specific words for emphasis in a sentence, where the tagging process would be made via semantic NLU tasks and the voice alteration by the TTS model
- noir_lord 4y agoThat'd be interesting because it'd split the problem into "parse and highlight what should be emphasised" and "do the TTS".
- gattilorenz 4y agoI think there's alrrady research for "TTS after NLG" that does this, since a NLG system can export meta-info about emphasis, in addition to the text (at least in case of non-end2end NLG systems). Whether that makes a big difference in practice, I don't know.
- DantesKite 4y agoThat is crazy. Any way I can start using this soon? I have a backlog of articles I’d love to listen to.
- VGltZUNvbnN1 4y agoTortoiseTTS might be the closest https://github.com/neonbjb/tortoise-tts https://github.com/neonbjb/tortoise-tts It's a few shot multi speaker model so you need just 3-4 little clips to train new voices.
- DantesKite 4y agoThank you very much. This is what I was looking for.
- nwatab 4y agoDo you convert articles and listen to them?
- DantesKite 4y agoI do, but I don’t like many text to speech voices at the moment because they don’t always enunciate in a way that makes sense. So I’ve been looking forward to the day when I can use custom human voices to read me news articles.
- nwatab 4y agoI'm in the same boat. I have tried GCP and AWS, and they sound too robotic. Do you convert articles you wrote or random articles? The reason I'm asking is because I'm thinking about new project that works like Descript and that is much easy to use.
- qgin 4y agoUnbelievable. This has traversed the uncanny valley and come out the other side.
- est31 4y agoVery subtle differences, can be heard, but I have my headphones on. For example, in the last example, "borne" and "commission" seem to have some kind of artificial noise inside the "b" and "c" sounds. The "th" in "clothing" sounds artificial too. Still, it's extremely amazing, and probably in 90% of settings, people won't be able to find a difference at all. It even does breaths: "scientific certainty <breath> that".
- DonHopkins 4y agoNice pitch envelopes. But it's a bit uncanny that natural human pitch envelopes encode and express what you understand and intend to convey about the meaning of the words you're saying, and what you want to emphasize about each individual word, emotionally. Like how you'll say a word you don't really mean sarcastically. It can figure out it's a question because the sentence ends in a question mark, and it raises the pitch at the end, but it can't figure out what the meaning or point of the question is, and which words to emphasize and stress to convey that meaning. (Not a criticism of this excellent work, just pointing out how hard a problem it is!) For example, compare "rebuke and abash": in the NaturalSpeech, one goes down like she's sure and the other goes up like she's questioning, where in the recording, they are both more balanced and emphasized as equally important words in the sentence. And the pause after insolent in "insolent and daring" sounds uneven compared to the recording, which emphasizes the pair of words more equally and tightly. Jiminy Glick interviews (and does an impression of) Jerry Seinfeld: https://www.youtube.com/watch?v=AE2utktZ92Y https://www.youtube.com/watch?v=AE2utktZ92Y
- stubish 4y agoI had always thought a necessary step along the way to natural speech synthesis would be adding markup to the text. But I guess the use case being chased is reading factual information, where you don't use sarcasm or other emotional color. Trying to convert text to emotive speech requires markup IMO. Even the best actors will read their lines the wrong way and need correction by the director.
- DonHopkins 4y agoThat's a good point. There are various xml markup formats for synthesizing text to speech that let you tag words for emphasis and pitch, but there's not granular or expressive enough to mark up individual syllables of words, and it's not useful for singing, for example. It would get really messy of you had to put tags around individual letters, and letters don't even map directly to syllables, so you'd need to mark up a phonetic transcription. At that point you might as well use a binary file format, not xml. But for that kind of stuff (singing), there are great tools like Vocaloid (which is a HUGE thing in Japan): https://www.vocaloid.com/en/ https://www.vocaloid.com/en/ https://en.wikipedia.org/wiki/Vocaloid https://en.wikipedia.org/wiki/Vocaloid VOCALOID5 - Walkthrough https://www.youtube.com/watch?v=UAtVGHl1AFM https://www.youtube.com/watch?v=UAtVGHl1AFM (Check out the "Cool / Cute" slider at 8:22!) Here's a much simpler and cruder tool I made years ago (when xml was all the rage) for editing and "rotoscoping" speech pitch envelopes, called the "Phoneloper" -- not quite as polished and refined as Vocaloid, but it was sure fun to make and play with: Phoneloper Demo The Phoneloper is a toy/tool for creating and editing expressive speech "Phonelopes" that Don Hopkins developed for Will Wright's Stupid Fun Club in 2003, using Python + Tkiter + CMU's "Flite" open source speech synthesizer. It modified Flite so it could export and import the diphone/pitch/timing as xml "Phonelopes", so you could synthesize a sentence to get an initial stream of diphones and a pitch envelope. Then you could edit them by selecting and dragging them around, add and delete control points from the pitch and amplitude tracks to inflect the speech, stretch the diphones to change their duration, etc. It was not "fully automatic", but you could load an audio file and draw its spectrogram in the background of the pitch track, so you stretch the diphones and "rotoscope" the pitch track to match it. https://www.youtube.com/watch?v=qy5cqV8ypIs https://www.youtube.com/watch?v=qy5cqV8ypIs
- sriku 4y agoI wish for the "naturalspeech versus recording" comparisons they'd used a different voice for the synthesized speech. Otherwise the fact that we may not be able to tell them apart by ear (in a blindfold test) doesn't tell us much about how good it is as a speech synth engine with that evidence alone.
- jcims 4y agoThere's no way I'll find it but somewhere along the way there was a collection of samples in which one of these contemporary model-based speech synthesizers (possibly wavenet or tacotron) was forced to output data with no useful text (can't remember if it was just noise or literally zero input). The synthesizer just started creating weird breathy pops and purrs and gibberish utterances. Some of them sounded like panic breathing and it was one of the more jarring things I've heard in quite some time. This isn't exactly it but it's very close - https://www.deepmind.com/blog/wavenet-a-generative-model-for-raw-audio https://www.deepmind.com/blog/wavenet-a-generative-model-for... CTRL+F 'babbling'
- guerrilla 4y agoMost of those sounds hilariously close to Danish, probably because of the glottal stop :P
- deleted 4y ago[deleted]
- DonHopkins 4y agoSounds a lot like Simlish! Katy Perry - Last Friday Night in Simlish: https://www.youtube.com/watch?v=sxyW6AJ-yIk https://www.youtube.com/watch?v=sxyW6AJ-yIk How the Language From the Sims Was Created: https://www.youtube.com/watch?v=FGsbeTV76YI https://www.youtube.com/watch?v=FGsbeTV76YI Simlish Voice Video (Gerri Lawlor and Stephen Kearin, inventors of Simlish): https://www.youtube.com/watch?v=Y_E6026i9tA https://www.youtube.com/watch?v=Y_E6026i9tA Steve and Gerri improvising together in English while playing The Sims: https://donhopkins.com/home/catalog/sounds/Steve_And_Gerri.wav https://donhopkins.com/home/catalog/sounds/Steve_And_Gerri.w... Gerri Lawlor: https://en.wikipedia.org/wiki/Gerri_Lawlor https://en.wikipedia.org/wiki/Gerri_Lawlor At the University of Maryland VAX Lab in the 80's, we had a DECTalk attached to the VAX over a serial line that we'd play around with, but I think the protocol must have used two byte tokens, because some times it would get one byte out of sync and start going "BLLEEGH YAAUGH RAWGH BRAGHK SPROP BLOP BLOP GUKGUK BWAUGHK GYAAUGHT BLOBBLE SPLOP BLAP BLAP BEAUGH GUWK SPLAPPLE PLAP SPLORPLE BLAPPLE"! (*) Just like it was channeling the Don Martin Sound Effects from random Mad Magazines. https://www.madcoversite.com/dmd-alphabetical.html https://www.madcoversite.com/dmd-alphabetical.html (*) Bulemia Meeting Attendees Vomiting, MAD #266 1987, Page 45, On Thursday Evening on West 12th Street.
- sebringj 4y agoYou could tell the difference in that the AI pronounced "Hussars" correctly where as the human reader did not. Without adding in our human error, our AI-trained version will be the more educated one for certain going forward.
- wrycoder 4y agoIs this available in other languages yet?
- microtherion 4y agoGood quality overall, though it's difficult to tell from a small, hand picked set of examples (which appear to come from the training data, too — have the corresponding recordings been included in the voice build or held out?). There is a rather obvious problem with the stress on "warehouses", and a more subtle problem with "warrants on them", where it's difficult to get the stress pattern just right.
- hooloovoo_zoo 4y agoI think they’re held out but all data comes from the same speaker and since they have like 10k samples times say 10 words per sample, practically every word will have been in the training data.
- justinlloyd 4y agoThis is pretty impressive work, except for this one: "who had borne the Queen's commission, first as cornet, and then lieutenant, in the 10th Hussars" Both the NaturalSpeech and the human said pretty much every word in that sentence completely incorrectly for the context of the words. It is the difference between "the car Seat" and "the car seat". "It's pronounced Ore-garh-no" to paraphrase the insufferable Hermione Granger.
- petesergeant 4y agoalso had he been in the 10th Hussars he'd have been "lef-tenant" not "loo-tenant", and it would have been the "hoo-zaaz"
- skykooler 4y agoWow, this is the first speech synthesis I've seen on here where I thought I was listening to a human at first.
- infinitone 4y agoThis is definitely human-level quality. In fact, the synthesized versions pronounce some words better than human. Kudos to MSFT! I think they've been longest in the game too... edit; is the Nuance acquisition compounding yet?
- colordrops 4y agoI actually think the TTS voice is better sounding than the human's voice.
- a9h74j 4y agoTo me the human voice samples have a bit of NPR in them, compared to the TTS.
- hooloovoo_zoo 4y agoCadence still seems way off for the AI. Maybe it’s going word by word?
- themodelplumber 4y agoReminds me of how good the choir instruments are these days. https://youtu.be/ulK3_o7OyEk?t=392 https://youtu.be/ulK3_o7OyEk?t=392
- zigzag312 4y agoThat is just playing audio samples. For true singing TTS check this: https://news.ycombinator.com/item?id=31420072 https://news.ycombinator.com/item?id=31420072
- warning26 4y agoAs far as I can tell the choir instrument in that video is just playing back samples, which are admittedly very high quality samples, but still nothing particularly groundbreaking. I’d love to see something that could actually do choir synthesis using a method like the one in the article.
- themodelplumber 4y agoOh, I thought it was the one where you type in words. Can't remember where to find the word-typing choir plug if that's not the one. But here's I guess a more pure-synth version (depending on what that means, like is any real voice data OK at any point in the algo?), not too bad. https://youtu.be/9vEA4iSjajg?t=517 https://youtu.be/9vEA4iSjajg?t=517 Edit: Ah, here's a type & speak version...who knows what cheats there are, could be somebody under his desk for all I know: https://www.youtube.com/watch?v=NNyQ7FWV2E8 https://www.youtube.com/watch?v=NNyQ7FWV2E8 Less drama, though different software: https://www.youtube.com/watch?v=UAtVGHl1AFM&t=177s https://www.youtube.com/watch?v=UAtVGHl1AFM&t=177s
- itronitron 4y agoreminds me of this CD from the mid `90's >> https://youtu.be/YNM-BJ9JDKU https://youtu.be/YNM-BJ9JDKU
- SemanticStrengh 4y agothx for sharing
- causality0 4y agoIt's interesting that TTS is getting better and better while consumer access to it is more and more restricted. A decade ago there were a half dozen totally separate TTS engines I could install on my phone and my Kindle came with its own that worked on any book.
- mrec 4y agoTo the Kindle thing, I wonder if the rise of audiobooks as a product killed the latter off to remove competition.
- causality0 4y agoOh absolutely, the same way Amazon removed physical buttons so they could use them as leverage for their overpriced Oasis reader. My twelve year old Kindle 3 has TTS, page turn buttons, and a headphone jack, and it cost a hundred bucks.
- exhilaration 4y agoIt was a rights issue. The Authors' Guild argued that TTS required audio rights. Here's an article from 2009: https://www.theguardian.com/technology/blog/2009/mar/01/authors-guild-blocks-kindle-voice https://www.theguardian.com/technology/blog/2009/mar/01/auth...
- autoexec 4y agoJust one more example of copyright stifling progress, innovation, and accessibility. The Authors Guild doesn't screw over the public by exploiting our insane copyright laws as often as the MPA or RIAA, but they've had their moments. https://www.eff.org/deeplinks/2005/09/authors-guild-sues-google https://www.eff.org/deeplinks/2005/09/authors-guild-sues-goo... https://www.cnn.com/2011/09/13/tech/web/authors-guild-book-digitization/index.html https://www.cnn.com/2011/09/13/tech/web/authors-guild-book-d... https://popula.com/2022/01/22/what-kind-of-writer-accuses-libraries-of-stealing/ https://popula.com/2022/01/22/what-kind-of-writer-accuses-li...
- big_fan 4y agoIndustry research lab claims human parity on end-to-end text-to-speech and releases a web page with five samples as proof? Microsoft, you're a little late to the party - Google has been using this playbook for 5 years!
- karmasimida 4y agoI think we have reached the stage of development of AI, I am no longer surprised/excited by this results by any means.
- Loeffelmann 4y agoWhat's a good TTS cloud service that has anything even close to these voices. I looked at the Google and Amazon ones and was pretty disappointed.
- taspeotis 4y agoDid you ... read ... what was written in the linked page? > Authors > ... > Microsoft Research Asia & Microsoft Azure Speech https://azure.microsoft.com/en-us/blog/announcing-new-voices-and-emotions-to-azure-neural-text-to-speech/ https://azure.microsoft.com/en-us/blog/announcing-new-voices...
- lordofgibbons 4y agoJust because they release a research paper doesn't mean it's out in production or if this model will even be economical to run in production for them.
- Loeffelmann 4y agoThank you! Thought it was just a research paper and they weren't in production yet. Samples on that blog post sound really good
- ridgered4 4y agoI played with demos of various cloud providers, a lot of it comes down to difference of opinion it seems as I've had people exclaim on here something was fantastic that sounded like a robot. My opinion is Microsoft's Azure TTS is the best and IBM's Watson TTS was a pretty close second. I remember finding Google's disappointing as well. https://azure.microsoft.com/en-us/services/cognitive-services/text-to-speech/#overview https://azure.microsoft.com/en-us/services/cognitive-service...
- thorum 4y agoSee also the recently published Tortoise TTS, which IMO sounds even better: https://github.com/neonbjb/tortoise-tts https://github.com/neonbjb/tortoise-tts
- heldtec 4y agojust had a quick play with this ! great !!
- ducktective 4y agoWOW! I'm flabbergasted! Checkout `Compared to Tacotron2 (with the LJSpeech voice)` or `Prompt Engineering` section!
- SemanticStrengh 4y agoIt sounds better and yet sounds like an overcompressed mp3
- ahaferburg 4y ago> If I, a tinkerer with a BS in computer science with a ~$15k computer can build this, then any motivated corporation or state can as well. Huh.
- sandreas 4y agoFor german users, I can recommend to take a look at https://www.thorsten-voice.de/ https://www.thorsten-voice.de/ https://github.com/thorstenMueller/Thorsten-Voice https://github.com/thorstenMueller/Thorsten-Voice where someone contributed a huge set of his voice samples and a tutorial / script collection to build a pretty decent TTS model LOCALLY. Quality-wise it is not as good as the samples in the article, but its free and pretty easy to follow for a tech enthusiast.
- gyulai 4y agoThanks for the pointer. Good non-English TTS is hard to come by, and this one sounds truly amazing!
- midjji 4y agoWhile every sample they provide is suspiciously similar to the human version,(indicating overtraining, either on the samples or on a single voice), where I would have expected a different if still human quality voice from a fully functional system, this tech is coming, and soon. And when it does, voice acting will no longer prevent videogames from having complex stories, and we will find out if the industry is still capable of making them. Looking forward to it :)
- SemanticStrengh 4y agoI'm looking forward to it for makind audio books which have no audio or plain reading me selections on webpages at desired speed
- UncleEntity 4y agoOne of the samples even has a breath intake at the same point as the recording. Not sure how they do that (didn’t read the paper) but I first thought it was the recording and compared the two to find the breathing wasn’t quite as natural as an actual person with lungs.
- SemanticStrengh 4y agoCan someone please upload the results? On https://paperswithcode.com/sota/text-to-speech-synthesis-on-ljspeech https://paperswithcode.com/sota/text-to-speech-synthesis-on-...
- baxuz 4y agoIs there any TTS engine which isn't based on English? I'd love to be able to use an assistant device in Croatian in my lifetime.
- rjzzleep 4y agoI have the same with other languages. One thing that I find fascinating is that while all the western TTS and language understanding frameworks allow only one at the time. The Chinese ones happily do Chinese and English at the same time. I recently tried out the open source android TTS engines and they don't seem all that great even though there's been years of development on them. Can anyone that knows comment on what the complexity here is?
- numpad0 4y agoAs a native Japanese and sort-of English speaker, I need to pause flush audio pipeline to switch between languages. Konnichiwa spoken in the middle of an English sentence and こんにちは in native Japanese, are separate expressions and have to be pronounced differently. From my experience there likely isn't a single unified language model inside a human head, and if so it's not a surprise to me that one cannot be effortlessly made for computers.
- follower 4y agoWhile these Free/Open Source licensed TTS engines have a "relatively large" number of non-English voice options neither seems to include Croatian. * https://rhasspy.github.io/larynx/ https://rhasspy.github.io/larynx/ * https://mycroft.ai/blog/mimic-3-preview/ https://mycroft.ai/blog/mimic-3-preview/ (forthcoming, currently in beta, actively looking for feedback on non-English language quality) From a quick look here, there didn't seem to be any Croatian open source data options either: * https://openslr.org/resources.php https://openslr.org/resources.php I have a vague recollection that there was at least one (I think) Eastern European country that managed to get government funding in order to support creation of local language assistive device text to speech. So, it doesn't seem like an impossible task but certainly a non-zero amount of work to collect & process appropriate audio data. Hope you get your dream at some point in future. (As everyone deserves assistive devices in their own language.)
- exebook 4y agoOne thing I've noticed is that I can hear human inhale before they continue speaking. Got curious if tts of the future should have this feature too.
- imdsm 4y agoThat's a really valid point, like making robots blink.
- addandsubtract 4y agoGoogle had a presentation a while back with an AI voice adding "uhm" and breathing pauses in their responses.
- maa5444 4y ago
- p1necone 4y agoThis kind of stuff is going to be amazing for indie gamedevs. I want a model trained for "powerful narrator voice" and villain speeches.
- follower 4y agoI imagine that our concept of what a villain sounds like tends to be extremely personally biased but here's a couple of options [Advisory: Contains threatening language.]: * http://www.sndup.net/p33q http://www.sndup.net/p33q * http://www.sndup.net/sppn http://www.sndup.net/sppn I created these samples in a relatively short time using the Free/Open Source (which I think is an important factor for indies) text-to-speech project Larynx & an narrative editor I finally released the other weekend: * https://github.com/rhasspy/larynx/ https://github.com/rhasspy/larynx/ * https://rancidbacon.itch.io/dialogue-tool-for-larynx-text-to-speech https://rancidbacon.itch.io/dialogue-tool-for-larynx-text-to... Now, I would really like to link you directly to audio of the next two but considering it's currently in beta behind an (automated response) email address, I think that may not be appropriate, so, instead... * Visit & get access to the beta here: https://mycroft.ai/blog/mimic-3-preview/ https://mycroft.ai/blog/mimic-3-preview/ * Copy & paste this SSML into the form: https://pastebin.com/Bwd7LCbj https://pastebin.com/Bwd7LCbj It's definitely a noticeable step up again in quality. There's an alternate pair of voices if you move the "_" from one "name" attribute to the other in each "voice" element. I intentionally didn't edit the text to remove some of the artifacts both to give a realistic impression of the current state & because sometimes they add interesting texture. :) Note the beta voices are "low" quality.
- Quequau 4y agoI don't suppose anyone could recommend a good text-to-speech for Linux? Command line is fine but it would be much better if it could trivially take clipboard content for input. The last time I looked I found stuff that wasn't that great and was pretty inconvenient.
- follower 4y agoI can recommend taking a look at: * Larynx: https://github.com/rhasspy/larynx/ https://github.com/rhasspy/larynx/ * OpenTTS: https://github.com/synesthesiam/opentts https://github.com/synesthesiam/opentts * Likely Mimic3 in the near future: https://mycroft.ai/blog/mimic-3-preview/ https://mycroft.ai/blog/mimic-3-preview/
- Quequau 4y agoThanks!
- addandsubtract 4y agoCoqui: https://github.com/coqui-ai/TTS https://github.com/coqui-ai/TTS
- blueflow 4y agoWhat the fuck is end-to-end text? My Bullshit-O-Meter is off the charts. End of what? I only know end-to-end encryption.
- qayxc 4y agoNothing BS about it. In ML, end-to-end means feeding raw data (e.g. text) to the model and getting raw data (e.g. waveform audio) out. This is on contrast to approaches that involve pre- and postprocessing (e.g. sending pronunciation tokens to the model, or models returning FFT packets or TTS parameters instead of raw waveforms). It's a common and well understood technical term in this context.
- jollybean 4y agoIt'd be nice if we could input our own text because otherwise these things are subject to a lot of training corpus and other biases. Sounds really good though.
- PedroBatista 4y agoVery cool, but.. What's the end game here? because I cannot use it, I cannot buy it and this seems more than just a scientific paper. So what's the objective here?
- nwatab 4y agoSounds nice. I'm interested in making business based on TTS
- msluyter 4y agoTotally tangential comment: You can click play on any/all of the samples simultaneously, resulting in a neat sonic effect vaguely reminiscent of Steve Reich's famous "Come out." [1] [1] https://www.youtube.com/watch?v=g0WVh1D0N50 https://www.youtube.com/watch?v=g0WVh1D0N50 (skip to like 7 minutes in to get the idea)
- explorigin 4y agoIt's clear that their dataset contains a lot of newscasts. I wouldn't call this "natural" speech. But it certainly has an application for replacing newscasters/announcers.
- coolspot 4y ago> We train our proposed system on 8 NVIDIA V100 GPUs with 32G memory Sounds like openly reproducing this result is within independent researchers’ reach.
- IYasha 4y agoAs a TTS daily user, sometimes I'm even fine with espeak quality for system messages. But one thing concerns me more than beauty of the voice - the ability to process mixed language text and abbreviations. And I don't see these problems addressed in this project. (
- UltraViolence 4y agoI'm not sure what to make of this. The TTS output seems identical to the recording. Why don't they use this tech to recreate some dead actor's speech, for example?