3 ms·
Good question. The answer is that time is an invariant here. Dejavu uses a locally sensitive hash (LSH) just like any other approximate search hash might. The
by muzakthings 10y ago
Good question. The answer is that time is an invariant here.
Dejavu uses a locally sensitive hash (LSH) just like any other approximate search hash might. The key to note is that we're binning both the frequency and time units of the spectrogram, giving us room for noise/error. In fact, you can tune the granularity to which this happens by adjusting the Dejavu FFT window (DEFAULT_WINDOW_SIZE). This will create a spectrogram with few (and therefore larger) cells.
The trade-off with smaller cells is that with too much granularity (or if the audio is even a tiny bit stretched or we have small Doplar shift effects), we may miss the fingerprints we want (false negative). On the other hand, with too low granularity, we risk having our frequency bins too large and having perhaps another song/query audio match when it shouldn't (false positive).
Luckily in either case, we don't need to see all the fingerprints, just enough to align properly in time.
So to answer your question, Dejavu is fairly resistant to noise (and can be tuned with FFT settings) as long as the audio's original timing is unchanged.