4 ms·
By the time we have generative neural networks capable of replicating human voice with emotion and nuance (way more difficult than neutrally reading a text on W
by blixt 11y ago
By the time we have generative neural networks capable of replicating human voice with emotion and nuance (way more difficult than neutrally reading a text on Wikipedia), I think it's fair to assume we'll also have decent "thought vector" networks that – much like how neural networks can turn words or even sentences into vectors and back (translation) – can turn the meaning a character wants to convey into multiple sentences arranged in unique ways. Basically taking your example a bit further.
- aphelion 11y agoAn algorithm understanding enough about the text to infer the correct emotional inflection to give a speech may edge into the category of strong AI, but I would have guessed that it would be easier to create a neural network that, given a text spoken in one voice, could transform it into another with the correct stress, intonation, etc. Perhaps even that's a more difficult task than I assumed, although speech generation seems to receive a lot less academic and industrial attention than speech recognition and understanding.
- michael_h 11y agoYeah, it's a bit more difficult of a task than you've assumed. Speech synthesis receives a lot of attention, but it's hard, so you rarely hear any news about it. People are throwing DNNs at it at the moment, but nothing earth shattering has come of it (yet). I have a couple of 'naturalness' filters that use DNNs and about 30% of the time, they drop all of their tones and I end up with an angry whisper as output. I don't work late too often.
- TuringTest 11y agoFor people interested in how hard it is, I recently read this [1] NYT article providing a comparison of synthetic speech that IBM experts tested for Watson in the Jeopardy competition. [1] http://www.nytimes.com/2016/02/15/technology/creating-a-computer-voice-that-people-like.html?_r=0 http://www.nytimes.com/2016/02/15/technology/creating-a-comp...
- hacker_9 11y agoI'm not sure those two are related. A 'thought vector' seems far more advanced than emulating emotion in voice no? Though you raise a good point about text itself not carrying any weight behind it. Solution here could be to also be able to tag words as 'emotional' or 'angry' in a visual editor, which the NN uses as hints when synthesizing.
- contingencies 11y agoMany games approximate a reasonable thought vector already. For example, Mount & Blade: Warband is a fantastic but relatively dated game (2010) which creates a world of hundreds of characters, each of which dynamically alter their relationship with the player along a single-axis based upon complex events. When pressed for information on subjects such as who within an alliance should be allocated newly conquered lands, they will make a decision and explain themselves by virtue of their relationship with others and the player. Whilst simplistic, it's actually extremely immersive already... and that was 2010. This year a 2016 remake Mount & Blade: Bannerlord is due for release... looking forward to wasting hundreds of hours more examining improvements!
- mrec 11y ago> replicating human voice with emotion and nuance (way more difficult than neutrally reading a text on Wikipedia) How much of that difficulty could be skipped with a bit of basic markup, to indicate things like emotion (<sad>, <happy>), tone (<sarcastic>, <forceful>) etc? That is, how much of the difficulty is inferring the emotion/tone/etc from the text, as opposed to expressing it in the generated speech?