3 ms·
This is pretty hard - we use raw data for speech [1, talked about in comment above] but it still needs some work to do really good synthesis. FFT is not really
by kastnerkyle 11y ago
This is pretty hard - we use raw data for speech [1, talked about in comment above] but it still needs some work to do really good synthesis. FFT is not really the way to go either - then you still need to deal with the problems of complex data which is very, very unpleasant. Most people use FFT -> IDCT (cepstrum) or a filtered version (mel-frequency cepstral coefficients, MFCC). This can work but it is a lot of domain knowledge.
One thing we tried in early testing, but did not pursue farther was vector quantized X (where X is MFCCs, LPC, LSF, FFT, cepstrum). Basically you use K-means to find clusters for some large number of K, then simply assign every real value (or real-valued vector) to the closest cluster. The cluster mapping becomes a codebook, and your problem goes from input vectors like [0.2, 0.7, 0.111, ...] to [0, 1, 0, ...] where the length of the vector of 0s and 1s is the number clusters K.
This is a much easier learning problem, and closely corresponds to most "bag-of-words" or word-level models. The quantization is lossy but for large enough K I do not think it would be noticeable. After all, we listen to discrete audio every day, all the time in wav format :)
To synthesize, you can either map codebook points back to the corresponding cluster center, or as most people do, map it to the cluster center with some small variance so you have a little bit of interesting variation.
[1] http://arxiv.org/abs/1506.02216 http://arxiv.org/abs/1506.02216
- jerf 11y agoThank you for expanding on my uninformed, off-the-cuff comment like that.