4 ms·
> I wonder if a style-transfer style algorithm could be used to map the intent of a sentence to a simulated voice. There's definitely research/proprietary soft
by follower 4y ago
> I wonder if a style-transfer style algorithm could be used to map the intent of a sentence to a simulated voice.
There's definitely research/proprietary software that can enable a person speaking in desired manner to have their voice control the expression of the generated speech.
Here's a related issue on a Open Source text to speech project which I only learned of today: https://github.com/neonbjb/tortoise-tts/issues/34#issue-1229995622 https://github.com/neonbjb/tortoise-tts/issues/34#issue-1229...
> I tend to view most of these things through the perspective of what would help mod-maker's for video games
Yeah, I think there's some really cool potential for indie creatives to have access to (even lower quality) voice simulation--for use in everything from the initial writing process (I find it quite interesting how engaging it is to hear one's words if that's going to be the final form--and even synthesis artifacts can prompt an emotion or thought to develop); to placeholder audio; and, even final audio in some cases.
> (and I suspect various open source voice sample sets would become pretty popular).
That's definitely a powerful enabler for Free/Open Source speech systems. There's a list of current data sets for speech at the "Open Speech and Language Resources" site: https://openslr.org/resources.php https://openslr.org/resources.php
Encouraging people to provide their voice for Public Domain/Open Source use does come with some ethical aspects that I think people need to be made aware of so they can make informed decisions about it.
Given your interest in this topic you might be interested in this (rough) tool I finally released last week: https://rancidbacon.itch.io/dialogue-tool-for-larynx-text-to-speech https://rancidbacon.itch.io/dialogue-tool-for-larynx-text-to...