3 ms·
Yes, pop in a normalizing flow that removes tone and recreates it with a small audio sample as context. https://arxiv.org/abs/2312.01479 https://arxiv.org/abs/
by programjames 2y ago
Yes, pop in a normalizing flow that removes tone and recreates it with a small audio sample as context.
https://arxiv.org/abs/2312.01479 https://arxiv.org/abs/2312.01479
- woodson 2y agoWhat I found is that, for cross-language use-cases, this often just applies the intonation of the “context” sample to the created sample, which, if they are from different languages, usually gives the wrong result (in the sense that it sounds off).