11 ms·
Deep-learning text-to-speech tool for generating voices of various characters
- SommaRaikkonen 6y agoWelp, after messing around with a few voices I was completely impressed with Glados's. This is really cool because I have no idea how the character's voice was synthesized, but apparently ML can do it for me so props to that.
- smrq 6y agoI'm pretty sure the real Glados voice effect is mostly pitch correction and formant shifting. You can do it with Melodyne at least (which, to be fair, is also computer magic-- just a different kind than this one!) I just found a video on YT with an example of recreating this in Melodyne: https://youtu.be/1oQn66gvwKA https://youtu.be/1oQn66gvwKA
- deleted 6y ago[deleted]
- aksss 6y agoMy favorite is Carl Butananadilewski, but I just ended up making him say actual phrases from ATHF in the end. Was hoping to see Meatwad as a character option.
- giantrobot 6y agoGLaDOS was voiced by a real person [0]. Her voice had some effects added but mostly just her trying to sound like a computer. [0] http://ellenmclain.net/ http://ellenmclain.net/
- pure-struggle 6y agowill this be open source eventually?
- pure-struggle 6y agohttps://twitter.com/fifteenai/status/1342304487474606081 https://twitter.com/fifteenai/status/1342304487474606081 found an answer. "There's no point in releasing a poorly done model, and to do so for the sake of popularity would be despicable. My goal is to achieve indistinguishability, which I certainly know is possible. Anything short of near-perfection is unacceptable. "
- whatshisface 6y agoReleasing a poorly done intermediate result would give either competitors or colleagues a leg up in the race, depending on whether one sees them as competitors or colleagues.
- scrollaway 6y agoMegalomania, always a great excuse. AI and ML users are massively benefiting from open source but too often refuse to release their data. It's like we're back in the middle ages and alchemy is back in style.
- hooloovoo_zoo 6y agoJudging by how the model and site are put together, I think this is some software engineer's hobby project. Not wanting to spill their secrets doesn't make them a megalomaniac for the same reason being a magician doesn't make one a megalomaniac.
- scrollaway 6y agoExcept magicians do actually share their secrets; there is an active trade around it, conferences, discussions and lots of reading material available. The barrier of entry is higher than any old open source project but it's not inaccessible and comparable to alchemy. I was talking about ML in general, not just this project. See OpenAI and their latest release for example: no public product, no trained model. Just alchemy.
- 6y ago
- bailey1541 6y agoIf it doesn’t work on mobile why bother sharing?
- clxxx 6y agoIt works on mobile for me. Tried it on both safari and chrome on an iPhone running iOS 14.
- nmfisher 6y agoFrom the about section: > How much does maintaining the servers cost? > It depends on the amount of traffic, but the minimum baseline is around several thousands of US dollars every month. This is expected as inference is very GPU intensive and a sufficient number of instances need to be spun up to handle thousands of requests coming in every minute. Everything is paid out of pocket. Wow, impressive commitment for something that's free.
- vsupalov 6y agoYeah, running anything related to AI involves GPU instances. An alternative is to point people to using Google Colab where you can get access to a GPU for free, but that's not a smooth end user experience for most folks.
- aisofteng 6y ago> running anything related to AI involves GPU instances This is not true. A _lot_ of AI applications use algorithms such as logistic regression or random forests and don’t need GPUs - partly, of course, because GPUs are so expensive and these approaches are good enough (or more than good enough) for many applications.
- vsupalov 6y agoWhoops, sloppy generalization on my part. You're completely right of course, thanks! I've been focusing on deep learning a lot lately, to the point where AI has become an alias for those exciting new GPU-heavy techniques.
- mickof 6y agoYou just sort of assume that this is correct? The person[1] running this comes across as a severely unstable character, that number is probably hyperbole. [1] https://twitter.com/fifteenai https://twitter.com/fifteenai
- 15ai 6y ago
- atum47 6y agoWhile you're typing the word the text box don't show it, when you complete the word then it shows on the text box. Brave, Android. Besides that, amazing results. Congratulations.
- uberman 6y agoThis was amazing!
- vsupalov 6y agoThe results are really impressive. At the moment I'm considering spending a low 3-figure amount for a professionally spoken intro for a new podcast. Some of the lines I generated are in my top 5 easily, human speakers don't have a lot of edge for short generic blurbs of text anymore it seems.
- dnsiseuzb 6y agoHow does this compare to wellsaid labs?
- centimeter 6y agoThis is extremely impressive. I wonder if this will lead to a resurgence of "moon man" style videos with well-known characters rapping extremely offensive lyrics.
- deleted 6y ago[deleted]
- suyash 6y agofun but what are the legal implications of using these voices for projects? Does the license cover the use of these voices?
- Meph504 6y agoseriously fuck anyone that is putting in forced time delays on their terms, how about you let me read what it is you are doing before requiring shit like this.
- duckmysick 6y agoIf you don't agree with the terms, including how they are presented to you, you can always reject them and leave the site.
- danShumway 6y agoI don't usually expect much from demos like this, but I'm kind of surprised how impressive the results currently are. They're definitely not perfect, you're definitely getting some odd clipping and noise, but this shows a large amount of promise. Being able to generate voices for games would enable a lot of interesting indie projects. IMO people should be paying more attention the market implications of products like this than to the social implications. There are a lot of projects that just aren't really feasible right now that could be if this kind of technology was more polished and generally available for commercial/self-hosted use. And in those cases, you don't even need to do inference, makers will likely be willing to mark up their scripts themselves. Anyway I digress. Congrats, this is really cool!
- Pfhreak 6y ago> people should be paying more attention the market implications of products like this than to the social implications. People will absolutely suffer harm from this tech, but hey, think about the dollars that could be made! No, we should absolutely be paying more attention to the social implications.
- danShumway 6y agoEh, this technology currently falls very squarely into the category of "almost good enough that I could use it for a creative project, but not nearly good enough that you're going to be able to convince me that the results aren't generated." I'm not primarily interested about the dollars, I'm interested in allowing communities to do creative things. I think people are looking at this tech like it's only going to be used for deepfakes, and they're underestimating the extent it's going to be used to create voice-acted game mods, animations, anonymization tools, and other creative/helpful projects. If you're really worried about this stuff though, you can take some comfort in the fact that by far the worst examples on the site are of real-world voices. This is currently technology that as far as I can see is far more suited for generating new voices or voicing cartoon characters with well-defined patterns/inflections than it is for imitating the president.
- 6y ago
- SV_BubbleTime 6y agoIs the author being cute putting Chell from Portal and Freeman from Half-life in there, and then there is no audio? It would be a weird oversight if not intentional because the author is clearly familiar with Valve games.
- twangist 6y agoI get nothing but "Error code 422: Server error", even on input "Hello", in FF, Safari and Chrome.
- Roritharr 6y agoI'm pretty happy with the results I get. I've toyed around with a similar goal, but with the idea of approaching voice actors to give them a powerful tool to sell a "low quality" version of their voice in bulk. That way an up and coming author could use a tool like this and some elbow grease to create an Audiobook with famous voices.
- code51 6y agoI'm fearing this will end up with a massive debt on their part.
- trowngon 6y agoAre there open source projects like this?
- CookieAnon 6y agoI have CookieTTS where I reseach lots of experimental stuff. (You can see my credits on the 'Thanks' section of 15.ai) I can get about 90% of the quality of 15.ai currently. I think I could surpass 15.ai but not without some help.
- nmstoker 6y agoThere's Mozilla TTS https://github.com/mozilla/TTS https://github.com/mozilla/TTS Here's a sample from a TTS model + vocoder I released for it. I've no wish to deter the motivated, but it'd take a bit of figuring out how to set things up and you'd need to read the docs and code to get oriented :) https://m.soundcloud.com/user-726556259/sherlock-wavegrad-sample-260k https://m.soundcloud.com/user-726556259/sherlock-wavegrad-sa... Links to the models are here: https://discourse.mozilla.org/t/creating-a-github-page-for-hosting-community-trained-models/70889/11?u=nmstoker https://discourse.mozilla.org/t/creating-a-github-page-for-h... Is originally trained on two novels read by the same narrator on LibriVox (ie in public domain)
- terrycody 6y agoIs there a simple interface like the example in this thread to use the tool for a non developer regard Mozila TTS yet? I can't find one...
- nmstoker 6y agoThere's the demo server which has a simple web UI where you can input text to be spoken, but in regards to setting it up locally it's not that suited for a non developer https://github.com/mozilla/TTS/tree/master/TTS/server https://github.com/mozilla/TTS/tree/master/TTS/server https://github.com/mozilla/TTS/wiki/Build-instructions-for-server https://github.com/mozilla/TTS/wiki/Build-instructions-for-s... There's also a version in docker: https://github.com/synesthesiam/docker-mozillatts https://github.com/synesthesiam/docker-mozillatts And various Colabs too, which are fairly easy to get going with: https://github.com/mozilla/TTS/wiki/TTS-Notebooks-and-Tutorials https://github.com/mozilla/TTS/wiki/TTS-Notebooks-and-Tutori...
- MartinoPalmitos 6y agoHalf-Life's Gordon Freeman voice is really spot-on!
- kebman 6y agoPretty cool! I tried it with this small dialogue, and then edited together two voices in Reaper from the downloads: Bob: “Hello, John.” John: “Oh, hello there, Bob.” Bob: “Yes, hello. It's what I said. Why do you keep repeating what I say, John?” John: “I didn't repeat you! I merely said hello, you dimwit!” Bob: “There you go, being condescending again. Fuck you!” John: “What? You're the one who started it!” Try it yourself, or write something different. Either way, good fun!
- mvts 6y agoNice work on the Gordon Freeman Voice :D
- demonictoaster 6y agoThe security implications of this kind of tech are scary. Going forward it will become really easy to reproduce the voice of anyone! It seems not a lot of training data is required to achieve reasonable results (e.g. Spong Bob is just 27min of voice, Half Life Black Mesa Announcer is just 1.9min!!). This stuff could be easily leveraged for scams and deep fakes (along with deep learning models that could also tweak lip movements to match the voice for example). Thankfully, there is also a very active area of research that leverages similar tech to detect deep fakes.
- dschooh 6y agoThese kinds of discussions are common with articles about deep fake video and audio. While I do not disagree with your point, here are two quick thoughts: - We have had perfect image manipulation capabilities for quite some time now. We have had written text manipulation capabilities for hundreds of years. - People will continue to believe what they believe, whether there is deep fake video and audio or not.
- demonictoaster 6y agoAgree with you. Hopefully people are more and more aware that they cannot trust anything out there. We are soon reaching a point where we can make anyone say anything we want, including in audio and video format.
- spyder 6y agoIt's already happening: A Voice Deepfake Was Used To Scam A CEO Out Of $243,000: https://www.forbes.com/sites/jessedamiani/2019/09/03/a-voice-deepfake-was-used-to-scam-a-ceo-out-of-243000/ https://www.forbes.com/sites/jessedamiani/2019/09/03/a-voice...
- superasn 6y agoReally impressive. Do you plan to implement an API like Amazon, Google that lets you generate TTS for price?
- wongarsu 6y agoI too think that this has potential as a cloud TTS service. However that does open up all the moral and legal cans of worms around this. I could imagine some of the voice actresses not being very happy about somebody else commercializing their voice without their consent. The obvious way to get around this is to keep this as the showcase and to pay some people to add their voices to the paid version. I imagine this would sell just based on being decent TTS with a wide range of voices, even when people don't know the voices offered.
- bravura 6y ago15ai, do you mind talking a bit about the methods you are using?
- Baeocystin 6y agoThe fact that you included Chell as a voice choice (and 'generated' a null audio clip to boot) earns a chuckle. The quality of the voices across the board earns wide eyes and an eyebrow raise. Thanks for sharing this, it's remarkable work.
- high_byte 6y agoGLaDOS hahahaha this is just... perfect. Stanley Parable Narrator funny you should mention this.
- st1x7 6y agoYou should really see what happens when you click reject on their terms and conditions prompt.
- junon 6y agoI'm rarely impressed by demos like this. This is a clear exception. Not only that, but the creator seems cool and down to earth. Thanks for sharing, this is incredible work.
- mensetmanusman 6y agoAs Alexa and Siri have improved over the last couple years and gotten a more human voice, it has been interesting observing my young children (1-4) interact with such devices. There is definitely a sense of ‘who is that’ coming from their little minds that they are sometimes quite perplexed about. ‘It’s a computer’ is starting to feel like a cop-out answer as these things improve...
- hmate9 6y agoIt’s incredible how little data is required for amazing output! Only a couple of minutes of talking needed. You can find a couple of minutes of taking of anyone, so the security implications are huge!
- EugeneOZ 6y agoPlease give me a hint how to control the speed - Portal:Wheatly is too fast for me. Amazing toy! Thanks for "download" link, I'm creating a collection of GlaDOS phrases now.
- rkagerer 6y agoOne of the voice actors is John de Lancie! https://soundcloud.com/user-860705643/q-pandemic-rant-no-music/s-IKeCiovdTJ8 https://soundcloud.com/user-860705643/q-pandemic-rant-no-mus...