8 ms·
Show HN: Shazam-like acoustic fingerprinting of continuous audio streams
- Xeoncross 9y agoThanks for sharing. Processing PCM audio signals is something that is actually useful for more things that people realize.
- dest 9y agoHope it will be useful! This lib is a brick in an adblock for radio broadcasts I have been developing for a while and that I am progressively open sourcing.
- slig 9y agoCare to share more details of how it works? Thanks!
- dest 9y agoThis, I will publish at a later time ;)
- vitovito 9y agoHow closely can it correlate audio broadcasts of the same audio that were captured at different offsets? e.g. two independent streams, identifying the same 30-second commercial, but the audio streams are offset from each other by half a sample length?
- dest 9y agoIt correlates quite well. Maybe some fingerprints will be present in only one of the two streams, but most of them will be present in both.
- snissn 9y agoHave you considered building a podcast player app that can automatically skip ads?
- dest 9y agoYes, I more or less have, but isn't the fast-forward option in podcasts players good enough? BTW a few months ago, I talked to an Australian dev that did podblocker.com, but the project does not seem active anymore https://news.ycombinator.com/item?id=13799700 https://news.ycombinator.com/item?id=13799700
- icebraining 9y agoI don't actually mind the ads, but fast-forward is a pain if you're listening when doing other stuff at the same time (running, biking, driving, cooking, etc).
- StavrosK 9y agoI hacked something in an hour once, and made a program that would recognize the song that was playing and played the video clip of that song from YouTube in sync: https://www.youtube.com/watch?v=K6FxfZH_ZK4 https://www.youtube.com/watch?v=K6FxfZH_ZK4 The phone in that video is just playing a song, it doesn't have any connection to the computer at all.
- dest 9y agoNice. How did you recognize the song? Cross correlation or fingerprinting? How big was your song database?
- sunsetMurk 9y agoyou have public repo for this? awesome stuff.
- StavrosK 9y agoIt was a really dirty 40 lines of code, so I don't have it publicly anywhere, but I can upload it somewhere when I get home if you want.
- maephisto 9y agoAwesome share!
- dest 9y agothank you!
- megamindbrian2 9y agoCan it fingerprint other streams?
- dest 9y agoYou mean audio streams? Of course. Just change the URL next to curl and that's it.
- dest 9y agoOP here. This lib is a brick in an adblock for radio broadcasts I have been developing for a while and that I am progressively open sourcing.
- throwmenow_0140 9y agoCool concept. Do you silence/skip the sound when you don't recognize a song? Nice idea for Spotify for free without ads - piping the sound into a virtual device that gets silenced when ads play. Edit: I don't want to support the notion that we should avoid 10$/month for such a great service, I was just curious about the technical implementation.
- dest 9y agoThere's speech and original musical content, like bootlegs, mixes, that cannot be detected with fingerprinting. For a free Spotify without ads, have a look at http://www.stationripper.com/ http://www.stationripper.com/ (it's old software) Edit: Station Ripper had caused a law/politics debate in 2005 in France about the right to do private copies of legal media http://www.assemblee-nationale.fr/12/amendements/1206/120600154.asp http://www.assemblee-nationale.fr/12/amendements/1206/120600...
- hammock 9y agoGoogle Pixel 2 phones are doing this now as an out-of-the-box feature. It's continuously listening and the song name appears on your lock screen. https://venturebeat.com/2017/10/19/how-googles-pixel-2-now-playing-song-identification-works/ https://venturebeat.com/2017/10/19/how-googles-pixel-2-now-p...
- dest 9y agoI wonder how much battery it drains.
- goldenkey 9y agoTheres a whole bunch of crap Android phones do now in thr background like the "Ok/Hey Google" assistant voice shortcuts. I turn it all off. But supposedly it only uses a special lower power chip to do these passive listening actions. Not sure about music recognization though -- seems like it would involve a decent amount of memory even if performing a convolution..depends how much buffer time.
- 0x00000000 9y agoA new one for me the other day was "Google Nearby". Enabled by default and some company in the airport using it to push ads to your notifications. Disgusting and maybe the final nail in the coffin for Android for me. As a long time diehard android user, iPhone sounds better and better every day.
- dmitrygr 9y agoCould you please tell me which airport this was and if you remember the location and the ad? Thanks a lot
- 0x00000000 9y agoIt was either BWI or LAX and it said "Your ad here" with some url. I had no idea what it was so I long pressed on the notification and did some searching
- throwmenow_0140 9y agoVery cool stuff! It seems that all those solutions are based on the analysis of visual representations of spectrograms. Is this common or could you just use 2d arrays which encode the same information - would this be more performant? Nice blog post about this stuff: http://willdrevo.com/fingerprinting-and-audio-recognition-with-python/ http://willdrevo.com/fingerprinting-and-audio-recognition-wi... - https://github.com/worldveil/dejavu https://github.com/worldveil/dejavu
- dest 9y agoYou mean 2d arrays containing the raw audio signal? No, this would not work because you do not know the phase along the y dimension when you want to compare to another signal. Another method to detect an audio pattern is cross correlation on the raw audio signal. But it is very expensive in computation power and memory. The longest operation with fingerprinting is often the DB query that is associated. Lots of work to do there. In that space, Will Drevo's work is really good. I will share my DB implementation later.
- throwmenow_0140 9y agoI meant the spectrogram encoded as a 2d array, but I guess there isn't a big difference when the db query is the most expensive part. I've always wondered: Is there a way to compare fingerprints with humming sounds or live recordings? Those fingerprinting techniques don't seem to be suitable for those tasks, do you know of any methods to accomplish this?
- dest 9y agoYou have special fingerprint algorithms that are suited for sound modifications like pitch https://biblio.ugent.be/publication/5754913 https://biblio.ugent.be/publication/5754913 but it's not going to work with humming or live audio. I don't know if such a thing exists. If you want to do some research, here is a short review paper on the topic http://www.cs.toronto.edu/~dross/ChandrasekharSharifiRoss_ISMIR2011.pdf http://www.cs.toronto.edu/~dross/ChandrasekharSharifiRoss_IS... As for 2d array spectrogram, it is not needed in my lib (expect when plotting is activated). I only care about maxima in the spectrum of each data window. In other words, 1d spectra are enough.
- ww520 9y agoIsn't Shazam patented?
- dest 9y agoMaybe, but I don't know. I'm in France and this lib is software only, so probably Shazam patents are not enforcable here. Anyway, IANAL and cheers to Shazam people
- disappearance 9y agoIf it is, it hasn't stopped DubSet [1] from licensing some tech to Apple and Spotify [2]. [1] http://www.dubset.com/mixscan/#intro-2 http://www.dubset.com/mixscan/#intro-2 [2] http://variety.com/2016/digital/news/spotify-apple-music-remixes-dubset-1201881686/ http://variety.com/2016/digital/news/spotify-apple-music-rem...
- peterburkimsher 9y agoThat's great! I was just thinking about rewriting Shazam as a machine learning project. I'm wondering how to use my Chord Progression data to make a different audio fingerprinting algorithm. https://peterburk.github.io/chordProgressions/index.html https://peterburk.github.io/chordProgressions/index.html
- durkie 9y agodo you think this could be useful for detecting changes in songs? like if i'm listing to a big mix of songs and they don't have timestamps of when the song changes, but that is info i would like to have...
- dest 9y agoYes it could be. You need a song database to detect changes, and that is hard and/or expensive to gather. Commercial services are available in that field. ACRCloud was mentioned in another comment.
- joren- 9y agoAnother implementation of this algorithm can be found at [1]. It also includes several other algorithms for acoustic fingerprinting that can serve as a baseline. See [2] for a paper on one of the other implemented algorithms and a comparison. [1] https://github.com/JorenSix/Panako https://github.com/JorenSix/Panako [2] http://www.terasoft.com.tw/conf/ismir2014/proceedings/T048_122_Paper.pdf http://www.terasoft.com.tw/conf/ismir2014/proceedings/T048_1...
- dest 9y agoThank you for having released Panako. Note that I gave the link to the relevant paper in a previous comment https://news.ycombinator.com/item?id=15811221 https://news.ycombinator.com/item?id=15811221
- toomuchtodo 9y agoThanks so much for your work on this. Interested in running it against the Internet Archive’s audio collection.