5 ms·
This is really impressive. Is anyone else really freaked out by that? I'm having an uncanny-valley type feeling, because the audio is 99.9% convincing, and the
by maxverse 3y ago
This is really impressive. Is anyone else really freaked out by that? I'm having an uncanny-valley type feeling, because the audio is 99.9% convincing, and the only thing that gives it away is inconsistent, not-always-correct pronunciation. But I'm not picking at the technical details. I'm freaked out by how good it is, and how easily it could pass for some person reading a human-written podcast. I know speech generation has been getting progressively better, but I guess I haven't heard it in a while (compare to the stock TikTok voice, for example.) Coupled with an LLM, this is too close for comfort.
- skaushik92 3y agoFor me the intonation of words and the pauses between them seem quite off from natural speech as well.
- williamstein 3y agoIt sounds like unnatural "podcast speech".
- dalmo3 3y agoAs a non-native speaker, I very much prefer this more steady voice than the unbelievably fake intonation a lot of American podcasters/YouTubers use when reading from a script. Mark Rober comes to mind as having a particularly unnatural and annoying cadence, intonation and pitch shifting during sentences.
- noduerme 3y agoIt does. It's irritatingly stilted in exactly the same way as The Daily.
- chaxor 3y agoIt's good - but the even more crazy thing is that this can be done by a script kiddy in a few hours - not an expert who spends months or years trying to whittle at some part of the process as it was several years ago.
- wondercraft 3y agoYou're right that it's a script, but it does have some intricacies that require a lot of testing to get it right. Examples: Using LLMS: - formulate the right prompts for the intro and outro generation - pass the content of a post in segments while maintaining history, as if you do in one go you will exceed token limit - figure out how to integrate comments properly - turn the summary into spoken format, not condensed written Using TTS: - train the right voice, one that fits the content. Not all voices of a TTS engine have the same characteristics. - understand the bugs of the TTS engine. For example Elevenlabs that we're using (and its beyond amazing overall and the team fantastic), is struggling when given this "$2.5". It will read it out "dollar 2(long pause) 5". - a few more things Overall: - Figure out how to connect all of the different segments, music intros, outros etc
- m00dy 3y agoIt is still a kid's play and more importantly having a low barrier to get into this is scary as hell
- jasonjmcghee 3y agoOh man, these edge cases are frustrating. I ran into “it’s a 50…………50 chance”, apparently it reads a hyphen as (long pause) too. I’m bullish on being able to give cues which are not read like Bark is doing. Their audio quality isn’t quite as polished as eleven labs, but it’s convincing / uncanny valley in other ways - laughs, throat clears, stutters, pauses
- wondercraft 3y agoThanks! Yeah the quality of this audio was what made us automate this process. With LLMs doing the leg work on the content curation as well, it's fairly straightforward. We built a UI around it on https://app.wondercraft.ai/ https://app.wondercraft.ai/ if you wanna check it out. The TTS engine used is elevenlabs btw.