3 ms·
There is a trick to do this :) If you think about it, you don't need high-resolution information about low frequencies, they don't change that fast. The lower t
by avaku 6y ago
There is a trick to do this :) If you think about it, you don't need high-resolution information about low frequencies, they don't change that fast. The lower the frequency, the fewer data points (complex) you need to reconstruct the original sound, because the lower frequencies don't change as fast throughout time. You wouldn't be able to do this with FFT, which I have started out with, but was unsuccessful. So I had to "invent" the new algorithm, which I've figured out later is equivalent to Wavelet Transform. The only new thing I've done is to make it "real-time", and I have some extra math for "high-precision" frequency detection, which deals with the problem of frequency leakage post-factum (I am willing to reveal this has something to do with rotating complex numbers :)).
- ssfrr 6y agoI'd definitely love to hear more about your technique once you're willing to say more about it. I'm really interested in time-frequency representations in general. One way to get high frequency resolution from the STFT or Wavelet transform is to use the phase derivative within each frequency band, which is usually called "instantaneous frequency", and is closely-related to the phase-vocoder (different from the daft-punk-famous vocoder). The main issue with this kind of instantaneous phase estimate is that it assumes there's only one dominant frequency within each band - does your method improve on that?
- avaku 6y agoThe core algo is basically like a vocoder (I think), tracking certain frequencies, and as others pointed out might be very very similar to https://en.m.wikipedia.org/wiki/Constant-Q_transform https://en.m.wikipedia.org/wiki/Constant-Q_transform (I guess another question is how to implement this effectively in "real-time"). About the question of "assuming there is only one frequency within each band" -- yes, it's a problem, and that's what's causing the "frequency leakage" that you see in the app with the "high-precision: off" (and FFT). To solve this, I've added some extra math on top, which is available via "high-precision: on" in the app. I will release more details when I open source the algo. Please subscribe to email list if you want to receive an update, I am not super active on social media otherwise :).