27 ms·
Don't want to spoil it for you if you really don't want to know but I want to share to others in case they do because I found it so interesting when I first lea
by turkeygizzard 4y ago
Don't want to spoil it for you if you really don't want to know but I want to share to others in case they do because I found it so interesting when I first learned!
It looks like others shared the paper: https://www.ee.columbia.edu/~dpwe/papers/Wang03-shazam.pdf https://www.ee.columbia.edu/~dpwe/papers/Wang03-shazam.pdf
It's short but very cool. I read it a while ago and honestly can't pretend I fully grokked everything, but my understanding was that you can't just use a Fourier transformation alone. Noise would basically make this impossible.
So what I'd consider the key insight is that they compressed songs down to "fingerprints". IIRC they noticed that songs, even in noisy environments, preserved certain bits of information. Particularly, they could look at the spectrogram and see peaks of amplitude in the tapestry. They essentially set some radius and scanned the spectrogram. In a given radius, only the largest amplitude value in time and frequency would be preserved. So you've reduce a 3MB song to several bits.
This would be good enough for small databases (I think). But it's intractable for anything practical. So they built hashes out of these fingerprints using pairs of the preserved peak bits. They would choose a certain peak (called the anchor point), record its time offset from the start of the song, and then form pairs with other nearby peaks, saving the pairs of frequencies (but discarding e.g their amplitudes). So for each of these anchor points, you would get a 64 bit value: 32 bits for the time offset and track ID and 32 bits of frequency-pairs.
When you wanted to look up a song, they would fingerprint your snippet into multiple 32bit hashes and compare them against the frequency-pair hashes in the database. If a song was a good match, then you would see that your snippet matched against multiple hashes from that song, and specifically they matched linearly over time (I'm struggling to explain this bit but it's visually obvious if you look at Figure 3 in the paper).
I probably got some of this wrong, but I hope it's a helpful summary of the paper. I remember struggling to understand parts of it, so please let me know if anything I said is egregiously wrong!