31 ms·
A question comes to mind: What would be preserved if converting the visuals back into audio? This may help answer another question: What type of visual would a
by undershirt 6y ago
A question comes to mind: What would be preserved if converting the visuals back into audio?
This may help answer another question: What type of visual would allow feeling the music without hearing?
- avaku 6y agoAs a matter of fact, I've converted the frequency representation that the core algo extracts back to audio without noticeable quality loss. Although in the app I am just showing the amplitude, but the phases are also important (people usually say that the phase of the note is not important for perception, but actually in complex sounds it's important to preserve relative phases of different frequencies). I think this intermediate representation that I've extracted using this algo would be much better for machine learning on the sound data than either 1) raw sound or 2) frequencies extracted with FFT. But there is an engineering difficulty to overcome: the frequency data that the algo extracts doubles in its amount (more frequent frequency samples) with each octave... It's like wavelet data... It's a challenge to feed this data to standard ML algorithms, need to think of how to configure the inputs, it would have to be highly hierarchical. My initial goal was to "learn" different instruments from raw sound data, and this intermediate representation is good, because it allows "translation invariance" across frequencies. I've described it here: https://vsound.app/high-precision-in-frequency-domain.html https://vsound.app/high-precision-in-frequency-domain.html It's a work in progress... This app is just for me to see if people are generally interested in this area, I don't want to do something for months only to discover nobody wants it (although it's been a few months I've worked on this app LOL).
- undershirt 6y agoI don’t really understand your descriptions. Maybe you can ground it for people better by comparing visuals produced by normal FFT and your novel method. I might go further by letting people hear the difference between them (by transposing the visuals back into audio, to ground our sense of quality loss in the original medium, sound).
- avaku 6y agoGreat ideas, will do! Thanks for suggestions!
- undershirt 6y agolooking forward to it! great work so far, congrats on publishing
- avaku 6y agoThank you! Subscribe to the email list on the front page, if you want to get updated. I'll probably release a desktop app soon, also free.
- ssfrr 6y agoThe difficulty he's describing also applies to wavelet transforms: higher frequency bands have wider bandwidth, so they need to be sampled more often. This means the resolution is different in different frequency bands. With the short-time fourier transform (STFT), which is widely used in Music Information Retrieval, each frequency band has the same bandwidth and is sampled at the same rate, so the output of the transform is a rectangular matrix. With a wavelet transform the output is sort of a trapezoid, because you have more samples in higher frequencies. In both cases though it's possible to losslessly round-trip the audio within numerical precision (neither the STFT nor the Wavelet transform lose information).
- avaku 6y agoExactly!
- klodolph 6y agoIn general, one of the problems with the spectrogram is that pure frequencies will get smeared across a few buckets, due to the finite window size. Similarly, transients will get smeared across a few windows. These two are at odds—no free lunch. If you increase the frequency resolution, you decrease the time resolution, and vice versa. If you are using a normal FFT this is no problem. You can reconstruct the original signal with a very small amount of error. However, this works because the FFT preserves phase data for each of the buckets. The spectrogram does not preserve phase data, so it will really mangle things. (I’ve tried this, but it’s been a while. You get a sound which is recognizeable but total garbage otherwise.)
- avaku 6y agoThere is a trick to do this :) If you think about it, you don't need high-resolution information about low frequencies, they don't change that fast. The lower the frequency, the fewer data points (complex) you need to reconstruct the original sound, because the lower frequencies don't change as fast throughout time. You wouldn't be able to do this with FFT, which I have started out with, but was unsuccessful. So I had to "invent" the new algorithm, which I've figured out later is equivalent to Wavelet Transform. The only new thing I've done is to make it "real-time", and I have some extra math for "high-precision" frequency detection, which deals with the problem of frequency leakage post-factum (I am willing to reveal this has something to do with rotating complex numbers :)).
- ssfrr 6y agoI'd definitely love to hear more about your technique once you're willing to say more about it. I'm really interested in time-frequency representations in general. One way to get high frequency resolution from the STFT or Wavelet transform is to use the phase derivative within each frequency band, which is usually called "instantaneous frequency", and is closely-related to the phase-vocoder (different from the daft-punk-famous vocoder). The main issue with this kind of instantaneous phase estimate is that it assumes there's only one dominant frequency within each band - does your method improve on that?