3 ms·
I work in this space, so I'd love to give a bit of detail on the ML if you're interested: Style transfer is trickier to do with speech than images! One signif
by whuffman 7y ago
I work in this space, so I'd love to give a bit of detail on the ML if you're interested:
Style transfer is trickier to do with speech than images! One significant issue is the lack of a good "content" versus "style" distinction. In images you can get great results by calling the higher-level features of an object classifier network "content" and holding that constant. Some people have tried this for audio with e.g. a phoneme classifier, but there are additional characteristics (such as inflection) that relate the emotional content of speech which wouldn't be held constant.
Another issue is that much of the speech classification work is done in spectrogram (or with further processing MFCC) space, which lets you treat audio similar to images and leverage a bunch of technology that we have for classifying those. But for synthesizing speech, spectrograms aren't a fantastic representation, because small errors in spectrogram space can translate into large errors in the waveform which are very clearly audible, and humans in general are pretty sensitive to audio errors. There are cool neural spectrogram inversion methods out there which can help, but those should still be trained to be robust to the kinds of errors that a style transfer algorithm would make, so it's still pretty tricky.
My company, Modulate, is building speech style transfer tech; and we've found a lot more success with adversarial methods on raw audio synthesis, where the adversary forces the generator to produce plausible speech from the target speaker!
One of the coolest parts of the kind of BMI research in this article, to me, is the potential to buy back some latency margin for speech conversion! If you're working on already-produced speech, there are super tight latency requirements if you want to hear your own speech in the converted voice - over 20-30ms for the entire audio loop, and you start to get echo-like feedback that makes speaking difficult. Even without looping back, you don't want more than 100-200ms of latency in a conversation before it starts impeding the flow of dialogue. This means your style transfer algorithm gets almost no future context, and limits the kinds of manipulations that you can do (not to mention the size of the network that you can do them with, depending on available compute power!).