10 ms·
AI Voice Generator: Text to Speech Software
- oefnak 4y agoThe music accompanying most of the samples is so loud you can barely hear the voices. This makes it difficult to get a good idea of how life-like the voices sound.
- jldugger 4y agoSeems intentional. It's better than whatever version of `say` ships with macOS but I can still hear a lot of artifacts and giveaways.
- bemmu 4y agoNo API it seems? I was looking for better TTS for my AI video presentation generator side project. Which one has the best voices out of those offering an API?
- throwaway675309 4y agoI'd be curious to know how this differs from Vocode.ai, which has been around for over a year now, and has voices from Sir Mix-A-Lot to Bender from Futurama. https://fakeyou.com https://fakeyou.com
- verdenti 4y ago[dead]
- buf 4y agoMurf.ai has also been around for a few years. I own a creator platform (with 500k or so Voice Actors) and have been very interested in AI Voices so I've been watching this develop for a bit. IMO, Murf's marketing page has better results than their product. I think the VAs on my platform are in trouble, but they still have a little ways to go.
- deleted 4y ago[deleted]
- consta 4y agoDo you mind sharing a link to your creator platform too?
- buf 4y agohttps://www.castingcall.club/ https://www.castingcall.club/ I'm an indie entrepreneur, so it's just me on this project, but it's been great fun.
- EZ-Cheeze 4y agoI'd work on that as a intern for free You can try some really really really interesting things with half a million users
- nmfisher 4y agoJust curious - have the worlds collided yet, with people asking voice actors to record data for training TTS models with their voice?
- buf 4y agoCurrently, projects asking for TTS model training are banned on the platform, but only because there was outrage amongst the users.
- fxtentacle 4y agoMy company has hired some people for TTS voice training on Upwork. About 90% of the voice actors resented the implications that someone else could make their voice say stuff that they disagree with. But some of them also found the idea of becoming digitally immortal very attractive. The same way some people like to put up a marble statue of their heroic deeds, others like to record themselves for the internet. In my opinion, both types of people want to avoid being forgotten and surely if you become a famous TTS voice, you'll have a Wikipedia entry...
- ms7892 4y ago[dead]
- tkgally 4y agoI tried it with several paragraphs of text. The many options offered—various voices with adjustable pitch, speed, and pauses, and customizable pronunciations for specific words and names—are attractive, and I can imagine a lot of potential uses. Like other voice synthesis software, though, it does not seem able to adjust the pauses and intonation to indicate emphasis and contrast the way a skilled human narrator does. I wonder if that will be coming as the AI becomes more meaning-aware.
- anonytrary 4y agoI entered a paragraph from the beginning of this article https://en.wikipedia.org/wiki/Hilbert_space https://en.wikipedia.org/wiki/Hilbert_space I selected several different voices, but it only generated between 2 and 11 seconds. Only got up to the first sentence...
- dannyw 4y agoIs there an open source version, that's as good as stable diffusion is when it comes to AI art?
- seydor 4y agoI expected software to be downloadable
- causality0 4y agoThis is trending into the uncanny valley where instead of sounding like a really good TTS system it sounds like an absent-minded half-illiterate cretin reading a script. Not sure that's a step in the right direction.
- logicallee 4y agoAbsolutely. But nothing a simple "in an enthusiastic style" in the prompt won't fix, that's how AI works right? :P
- thom 4y agoYou perhaps joke, but I suspect the combination of large text models to work out the subtext, and still fairly simple TTS models to render audio with a variety of emotional tones, is going to be very powerful in the future.
- stavros 4y agoIsn't it? You have to enter the uncanny valley before you cross it. The fact that this is in it means that it no longer triggers our "this isn't a person" response, and instead triggers our "this is a lazy person" one.
- mcbits 4y agoI already use TTS to listen to e-pulp fiction books, which is something a lot of people wouldn't put up with, but I could easily listen to books narrated by any of those sample voices, especially if it switched between distinct styles for each character without changing the voice. But that would still probably require human work to tag the dialogue with the right characters.
- deleted 4y ago[deleted]
- fxtentacle 4y agoDoes anyone know how the business model can work for such a product? I would expect that anyone working on scripts with voice-overs professionally would want to use their favorite movie/audio editor. That means from a user perspective, a "AI Voice VST/AAX Plugin" is strictly superior to whatever cloud GUI anybody builds. (EDIT: Also, running AI as a SaaS means murf.ai needs to pay for pricey datacenter GPUs. Any user-downloadable software will have much lower operating costs.) And the big elephant in the room with speech AI is that it's so easy to copy the tech. Just like Stable Diffusion did with images, TTS developers just train on public audio from the internet, so there is no dataset moat. And arXiv is full with papers that produce pretty good results, if implemented correctly. And NVIDIA has a collection of freely downloadable TTS models with good/usable quality. To me, it seems like it's only a matter of time until someone builds a high-quality open source TTS VST plugin and then all those SaaS offerings are basically worthless. In effect, what I'm asking is: What is the competitive moat here? How can murf.ai defend against a motivated high school kid with $100k in EC2 credits?
- Rastonbury 4y agoFor the segment of mom and pop store who need an explainer video or Facebook ad made in Canva and don't want to pay someone to record, they want easy of use, realism and editability/speed. My friend who runs a Shopify store asked for this. They are not going to fiddle with VST plugins or local/cloud GPUs.
- fxtentacle 4y agoAren't they better off hiring cheap on Fiverr for someone else to do the entire video? The traditional reason against this was that you'd want your narrator to sound like a native speaker. But if AI fixes that, is there any downside to outsourcing video voice-overs to cheap labor countries?
- Rastonbury 4y agoHow is that better? The AI should be cheaper and the with less hassle (creating a job, reviewing freelancers, negotiating) with less risk of poor quality/reworks and disputes and yes accent is a big one. The ideal TTS product for such a person would be something like: sign up and pay > choose voice > paste text > download audio
- O__________O 4y agoAudio samples to me feel lower quality than other samples I have seen from competitors, but been awhile seen I looked into text-to-speech so unable to quickly post example of a competitor’s less glitchy samples. EDIT: Here just one example, recall others, but unable to find them: - https://www.resemble.ai/ https://www.resemble.ai/ Anyone know why/how this company appears to be growing quickly?
- fxtentacle 4y ago[1] suggests that Murf.ai received $1.5 mio in Seed funding in 2021, has 12 employees and only made $78k in 2022 in revenue. Even if we price their developers at only $100k annually including all benefits, that suggest that they are close to being bankrupt cough close to raising the next funding round. [1] https://getlatka.com/companies/murfai https://getlatka.com/companies/murfai
- mritchie712 4y agoYeah, I feel spoiled by all the AI products that have come out in the past year. This one is underwhelming.
- jacooper 4y agoResemble feels a bit better, but still suffers from some weirdness, like some uncanny valley for voices
- stevehiehn 4y agoLove to use this for my procedural music experiments. I wonder if the EUA has any issue with that. It'd be awesome if the pitch and tempo were mapped to music pitch and tempo i.e pitch A440hz and 60bpm. I just tested text like: "one, two, three, four" and it looks like you could manually map it to pitch and bpm in a DAW.
- dejobaan 4y agoYMMV, but I recently picked up Synthesizer V with the Natalie voice—I think it's pretty incredible for a singing synth. You could potentially procgen out a source file (IIRC, it's plaintext), have Synthesizer V render it, and thereby skip the autotune/beat matching.
- stevehiehn 4y agoNoted, Thanks
- karmasimida 4y agoI listened some of the samples ... the artificialness of AI voice is very much present there. Unless the pricing is aggressively cheaper, can't say I am that impressed with the product.
- AltruisticGapHN 4y agoHuman voice is a carrier of emotion, it helps co-regulate our nervous system. It is extremely rich in signals. It is known in modern trauma therapy for example that people who are emotionally disconnected or in a state of shock have less "prosody" in their voice - the voice becomes more monotonous. In my opinion this tech is bad - and the more we spend time listening to artificial voices I would bet it can have a disregulating effect on the listener's nervous system. There is also a unhealthy trend on YouTube where creators actually voice their content, but they speak really fast and they cut all the pauses. It's really stressful to listen to in my experience, and I believe also unhealthy for listeners on the long run. It's no wonder that some creators who are just chill in their videos, sometime attract a wide audience, become a father-like figure almost - they could talk about anything - because younger people nowadays are just starving for this co-regulation effect. Like I'm watching a certain "Dwayne" and I don't need to agree to everything he says.. but the delivery is so calm and grounded , and there's none of that speeding up / cutting pauses non-sense, that it genuinely helps me as I am recovering from trauma. It calms me down. It's kinda unfortunate that at same time modern trauma models are gaining ground on YouTube, all about vagus nerve, fight/flight/freeze etc, the concept of capacity in the nervous system... at the same time you have an increasing assault from this really disregulating content... I guess all I can say s more than ever you have to be really aware of what you consume.
- thedorkknight 4y agoYeah the sped-up and micro-edited content is hard to miss after you start spotting it. I've definitely stopped watching certain channels just because of how grating it is
- prox 4y agoInteresting point you raise. I enjoy listening to Sovietwomble on Twitch, he just speaks relaxed like a radio host (he often mentions this as his inspiration) and he verbalizes what he is doing or thinking.
- croon 4y agoYou mentioning youtube fast cuts is sort of the reaction I felt too, but at least in those cases they're usually filmed around the same time by a human. Outside of the generated glitching in the sound here, my main complaint is that sentence umpteen sounds the same as sentence one. When we speak regularly, our intonation and cadence moves over time and the subject matter. A sentence here sounds okayish, but all the sentences in a row sounds like they're generated discretely (which I assume they technically are), and all the cohesion is gone.
- lovelearning 4y agoI really liked the way they've implemented their user interfaces and interactions. And the overall user experience too to a large extent, though I wish the actual TTS felt faster and responsive. As for its core functionality, sounded good enough for my modest needs.
- aszantu 4y agotime to get rid of smartphones, phones in general and only talk to ppl in person xD
- miki123211 4y agoIf you have any coding knowledge, you can get similar-quality voices for much cheaper from Azure, Google, AWS and IBM Watson. Azure gives you 500k characters for free per month, and then it's $16 per one million characters, paid per character. If you're using this to generate voice overs / videos, these rates are so low that you can basically forget about them existing. You have to use the API, but if that's fine with you, it's definitely worth it.
- vlugorilla 4y agoCan you elaborate on how could one achieve this? I have knowledge about Python and Golang.
- zackkatz 4y agoHere are the docs: https://learn.microsoft.com/en-us/azure/cognitive-services/speech-service/ https://learn.microsoft.com/en-us/azure/cognitive-services/s...
- liminalsunset 4y agoYou don't need coding to use Azure. Look up "Azure Audio Content Creation", which is a web-based application hosted by Microsoft that can be used to generate audio using a GUI. There is some rather annoying process to setup an account with resource groups etc required but it does give you 500k free characters, or you can just abuse the free demo applet on the website without signing in (may need to clear cookies and reload once in a while), and just tape the audio that comes out. I had very good luck with some of the Azure voices to create a YouTube video. My favourite right now is Sara (US English), because in testing she sounds the most emotionally natural. Interestingly, if you choose a voice from another language, and ask it to speak English, sometimes it will replicate a non-native accent, which I found somewhat amusing
- fancymcpoopoo 4y agowhere is the intelligence here? can people stop using the term AI for everything computer generated?
- ben_w 4y agoI've heard these voices a few times on youtube recently. I close those videos within seconds of recognising that the voice is synthetic. I'm not sure why my reaction is so strongly negative (I don't have this for GPT or SD). My first thought was "Infinite free generation means infinite A/B testing, and I don't want to be part of that", but that should exclude those other AI also.
- zulban 4y agoIt also means no human thought it was worth their time to voice the video. Not a good so sign.
- deleted 4y ago[deleted]
- exodust 4y agoIf you write naughty words you receive a telling off via email... "Our system has detected content that might be inappropriate...we request you to remove such content." I was sent this moments after signing up and entering one single word starting with F.
- Dowwie 4y agoThe black and white artwork is everywhere these days. Does the style have a name?
- amelius 4y agoWrapping-paper style?
- welshwelsh 4y agoI don't understand why text to speech approaches are so common. It's really hard to specify exactly what you want with text. It seems to me like speech-to-speech would be much better: start with your best attempt to produce the audio yourself, with the emotion, rhythm and timing you want. Then let the AI do the "last mile" transformation, taking your voice and making it sound like someone else, like how neural style transfer can change a picture to another style.
- schroeding 4y agoMy guess would be that text-to-speech scales very well for arbitrary data, for e.g. automatic audiobook generation, speech-to-speech does not. But yeah, fully agreed, for individual projects speech-to-speech appears to be a better idea, much more data to work with in there. Otherwise it will be a Vocaloid-like experience, where you have to tinker with the intonation of individual words. There is significant work in this area, too, e.g. Zero-Shot Voice Style Transfer: https://auspicious3000.github.io/autovc-demo/ https://auspicious3000.github.io/autovc-demo/
- staindk 4y agoI'm late to this but IIRC this is kind of the tech that LTT has started making use of for spanish audio - I don't know any of the nitty gritty details but I think they feed the english track as well as the script into the AI and get a much more natural-sounding translation out of it. For sections that don't come out "right" you can help it along by re-training just that section etc. See vid for some discussion around it - https://www.youtube.com/watch?v=_5uCvcyD0Eo https://www.youtube.com/watch?v=_5uCvcyD0Eo
- photoGrant 4y agoIt wouldn't accept my password. I used an asterisks and then it complained not to include whitespace. There was no whitespace. I gave up.
- takyon 4y agoCheck out Wellsaid Labs (https://wellsaidlabs.com/ https://wellsaidlabs.com/). Much better quality for longer text.
- techload 4y agoSuppose you would like to create an audio version of a long text to listen while commuting. What free tools would you use to acomplish that?
- deleted 4y ago[deleted]