5 ms·
I think all (so simple) you have to do is parse all the tracks ever made, and say generate a sequence of snapshots of what the tune sounds like and the delta. e
by goldcd 4y ago
I think all (so simple) you have to do is parse all the tracks ever made, and say generate a sequence of snapshots of what the tune sounds like and the delta.
e.g. if it was notes (for simplicity) E,D,C,D,E,E,E,D,D,D,E,E,E is the start of "Mary had a little Lamb"
Millions of tracks contain the note E. Many hundreds of thousands probably have the note D next - and as you work through the sequence, you're pruning down that list until you who what it is.
Bit that makes my mind hurt though, is the data-structure you put those sequences into to make it quickly searchable.
Users can start recording at any point in the song - so you can't just prune a tree down from a known starting point.
There's going be be background nose - so you need some way of "when you have no choice left", I presume sticking wild-cards into the previous decisions, to see if you end up back on a known track.
Yeah - I think it's magic as well.
Other thoughts:
I used it back in the UK when it launched, and the first track I ever used it on dialling (2580 - the numbers down the middle of your keypad) was also a French track (MC Solaar – La Vie Est Belle)
I always felt they missed a trick, just identifying music (and then trying to sell you stuff).
Surely they could have used the same tech to seamlessly mix all music together.
(i.e. take the sequences within tracks they find hard to differentiate, and then use these points to allow two tracks to be mixed together).
What's the minimum number of tracks it would say take to seamlessly mix from Megadeth to Mozart?
- zelos 4y agoThey used to have a paper on their website describing their algorithm in simplified form but I can't find it any more. Wikipedia has some details: https://en.wikipedia.org/wiki/Acoustic_fingerprint https://en.wikipedia.org/wiki/Acoustic_fingerprint I believe it's very sensitive to changes in timing, so it doesn't work on live performances etc. (based on reading I did 13 years ago before an interview at Shazam, which to this day still remains my worst interview performance)
- m-p-3 4y agoI'll also plug AcoustID from MusicBrainz https://musicbrainz.org/doc/AcoustID https://musicbrainz.org/doc/AcoustID
- senko 4y agoWe use AcoustID in MusicBox[0] to identify and deduplicate content, and it works great for us. What we do is calculate the acoustic fingerprint of every uploaded content and compare/check for duplicates (only authorized staff can upload, but this still helps a bunch with user errors and in cases where you need to reupload a track). Then we compare the fingerprints, using this[1] approach, so we can fine-tune the similarity based on our needs. In our case it's been very effective. Yes, live versions are treated as different ones (which is exactly what we need in our case, so it's a feature for us), but mechanical differences between tracks (volume, slight distortions from codec, different compression levels or remasters, or track being cut differently) are just ignored. If you ever want/need audio fingerprinting, I can warmly recommend it. [0] Music streaming service optimized for cafes, restaurants and other venues - https://musicbox.com.hr/ https://musicbox.com.hr/ [1] https://groups.google.com/forum/#!msg/acoustid/Uq_ASjaq3bw/kLreyQgxKmgJ https://groups.google.com/forum/#!msg/acoustid/Uq_ASjaq3bw/k...
- jefftk 4y ago> live versions are treated as different ones I think you're talking about a live recording vs a studio recording? But what I think zelos was talking about was "someone is currently playing music live, what is it?", which is a lot harder because you need to recognize the essence of a song and not the essence of a recording of a song.
- senko 4y agoYeah, agreed, that's way harder and not something AcoustID can do.
- muizelaar 4y agoThis paper perhaps? https://www.ee.columbia.edu/~dpwe/papers/Wang03-shazam.pdf https://www.ee.columbia.edu/~dpwe/papers/Wang03-shazam.pdf
- neon_electro 4y agoThe "doesn't work with live performances" bit is borne out by my consistent experience failing to identify some songs at live performances, but with the "DJ Set" form of live performance, tempo shifting music without pitch shifting it still appears to get the goods more often than not.
- saghm 4y agoMy instinct is that it probably isn't as simple as you describe because not only are there multiple notes at a time in a given track (i.e. chords), but there are also several tracks playing at once! It's possible that they're literally generating data like {guitar 1: C chord, guitar 2: single note E, bass: single note E} for every point in time, but even then each instrument isn't playing the exact same rhythm most of the time, so the notes won't exactly line up. I guess I don't think it's completely computationally infeasible to do it this way, but it seems more likely that they're just trying to separate the music from the background noise and then try to find the closest match to the music audio as a whole rather than trying to separate it into component.
- goldcd 4y agoSorry - I wasn't clear. I don't mean they're listening for notes. They're just analyzing the wave-form/fingerprint/whatever-you-want-to-call-it that's being generated at a moment, and then one form the next moment, then the next. One of these might match random points in many songs, but a far smaller subset of these will have the same three in the same sequence.
- saghm 4y agoFair enough! I imagine that having many instruments at once would improve the ability to diff the waveform/etc. rather than hindering it then.
- xhevahir 4y ago>if it was notes (for simplicity) E,D,C,D,E,E,E,D,D,D,E,E,E is the start of "Mary had a little Lamb" As far as I can tell these operate on audio, not symbolic music.
- plussed_reader 4y agoFFT data tends to get quantized, normalized, and counted for analysis purposes.
- goldcd 4y agoThey do (and I said so) - but I couldn't think of an easy way to write that in a post here.
- nibbleshifter 4y ago> Surely they could have used the same tech to seamlessly mix all music together. (i.e. take the sequences within tracks they find hard to differentiate, and then use these points to allow two tracks to be mixed together). What's the minimum number of tracks it would say take to seamlessly mix from Megadeth to Mozart? I noodled around with this idea in my free time a few years ago, got absolutely nowhere really usable with it (I probably put in a couple hundred hours). I knew I was limited by my dataset (small), code quality (terrible) and understanding of musical theory (virtually nil). Maybe I'll pick up that idea again - even doing beat matching would be kind of neat.
- bambataa 4y agoShazam as a product feels a bit odd. Almost as if they’ve never quite outgrown their slightly sketchy “advertised on MTV2 alongside the Crazy Frog” origins. They must have loads of data on songs people actually want to know yet never really managed to turn themselves into anything more sophisticated.