4 ms·
As a matter of fact, I've converted the frequency representation that the core algo extracts back to audio without noticeable quality loss. Although in the app
by avaku 6y ago
As a matter of fact, I've converted the frequency representation that the core algo extracts back to audio without noticeable quality loss. Although in the app I am just showing the amplitude, but the phases are also important (people usually say that the phase of the note is not important for perception, but actually in complex sounds it's important to preserve relative phases of different frequencies).
I think this intermediate representation that I've extracted using this algo would be much better for machine learning on the sound data than either 1) raw sound or 2) frequencies extracted with FFT. But there is an engineering difficulty to overcome: the frequency data that the algo extracts doubles in its amount (more frequent frequency samples) with each octave... It's like wavelet data... It's a challenge to feed this data to standard ML algorithms, need to think of how to configure the inputs, it would have to be highly hierarchical.
My initial goal was to "learn" different instruments from raw sound data, and this intermediate representation is good, because it allows "translation invariance" across frequencies. I've described it here: https://vsound.app/high-precision-in-frequency-domain.html https://vsound.app/high-precision-in-frequency-domain.html
It's a work in progress... This app is just for me to see if people are generally interested in this area, I don't want to do something for months only to discover nobody wants it (although it's been a few months I've worked on this app LOL).
- undershirt 6y agoI don’t really understand your descriptions. Maybe you can ground it for people better by comparing visuals produced by normal FFT and your novel method. I might go further by letting people hear the difference between them (by transposing the visuals back into audio, to ground our sense of quality loss in the original medium, sound).
- avaku 6y agoGreat ideas, will do! Thanks for suggestions!
- undershirt 6y agolooking forward to it! great work so far, congrats on publishing
- avaku 6y agoThank you! Subscribe to the email list on the front page, if you want to get updated. I'll probably release a desktop app soon, also free.
- ssfrr 6y agoThe difficulty he's describing also applies to wavelet transforms: higher frequency bands have wider bandwidth, so they need to be sampled more often. This means the resolution is different in different frequency bands. With the short-time fourier transform (STFT), which is widely used in Music Information Retrieval, each frequency band has the same bandwidth and is sampled at the same rate, so the output of the transform is a rectangular matrix. With a wavelet transform the output is sort of a trapezoid, because you have more samples in higher frequencies. In both cases though it's possible to losslessly round-trip the audio within numerical precision (neither the STFT nor the Wavelet transform lose information).
- avaku 6y agoExactly!