3 ms·
The user can paste a youtube url which will then be analysed, fingerprinted and matched against a database of 7+ million audio fingerprints. It does not only id
by tk42 12y ago
The user can paste a youtube url which will then be analysed, fingerprinted and matched against a database of 7+ million audio fingerprints. It does not only identify a single song but is able to identify multiple songs contained in a single file or video and generates a timeline listing which tracks it contains at which time.
Our matching algorithm is based on the open source echoprint-codegen fingerprinting method, which we have built our own stack around:
- Replaced Solr/Tokyo Tyrant with Elasticsearch
- Reimplemented matching-logic
- Crawlers search multiple sources for audio files to be indexed (mp3s arent stored long term, only fingerprinted then deleted)
- Indexing about 1 new track per second
- Found method to verify unrealiable ID3 tags (in progress, current database also includes unferified)
- mogilefs as primary data store for fingerprints
- perl everything
We also provide a free music identification API.
Any feedback would be much appreciated!
- tomtoise 12y agoUntil you posted this clarification, my first thought was - What makes it different from Shazam? Thanks for clearing that up, good luck with your site!
- corobo 12y agoDoes it include music from oft-used music sources in YouTube videos such as AudioMicro and Incompetech[1]? I'd guess it's not really possible on the AudioMicro front as they're paid-for music that'd cost you a fortune to index but may be worth adding the latter [1] http://incompetech.com/music/royalty-free/ http://incompetech.com/music/royalty-free/
- unltd 12y agoI've worked on the echoprint-codegen algorithm for my current project ( trak.rocks ) and I'm curious about how you reimplemented the matching logic ? Do you plan to document/opensource you work ?
- mo0 12y agoFirst we rewrote echonests truescore logic in perl and then altered slightly and implemented some extra checks to further try to exclude false positives. We also believe what they used in the late song/identify API might have been different from what is open sourced in https://github.com/echonest/echoprint-server https://github.com/echonest/echoprint-server Also we pack each individual hash before storing in Elasticsearch and gained at least 50% storage space this way. Our Fingerprint data is quite different from theirs(unreliable ID3 tags, N versions of same track) which is why we needed some tweaks. So far the matching is still far from perfect... Whether we will open source the whole thing at some point we don't know yet.
- caractacus 12y agoWhen you say the matching is far from perfect, is that at your end or on the part of the echoprint / echonest code? You made tweaks because you found issues with what they were doing....?
- tk42 12y agoThe reason for it being far from perfect is likely a combination of both. If the correct song is indexed there is a high probabiliy for us to find the right match. However if its not, with a bit of bad luck a false positive can happen easily with the default solution (and ours too). Also when analysing a youtube video it can happen that in a 30sec snippet only 10 secs are a matching song and 20 are unrealated or 15 are one matching song the other 15 match a different one in which case 2 tracks or multiple versions of 2 different tracks will have relatively OK scores. Deciding what to consider a match (or whether to try different queries for the same or slightly altered timespan prior to deciding) is not trivial in these cases and our changes are mostly concerning when a match will be considered a match by altering thresholds and how matching truescores will be looked at in relation to other fingerprints true scores. Due to issues like these, specifying a timeframe for analysis will often produce better results. http://static.echonest.com/echoprint_ismir.pdf http://static.echonest.com/echoprint_ismir.pdf
- brianwhitman 12y agosong/identify supported both ENMFP and Echoprint, and AFAIK the Echoprint matching path was exactly the same as is published on Github. I know at some point we did adapt the Solr end (for example, we removed the N most occurring codes) for speed optimizations. Many users of Echoprint in the wild have adapted the python matching logic for their use case as well as changed the hash update rate on the codegen. A great modification to watch was Sunlight Labs' "Ad Hawk", which ID'd commercials: https://github.com/sunlightlabs/adhawk https://github.com/sunlightlabs/adhawk
- samcrawford 12y agoDo you have any intuition for whether the echhoprint-codegen algorithm would be suitable for saying whether two voice recordings match? One would be a little lossy, the other pretty much perfect.
- unltd 12y agoEchonest can work with voice but is optimized for music so you might encounter a lot of false positive with it. Check out the echonest board on google. It's a recurring topic.
- lucaspiller 12y agoWhere do you get the MP3s from in the first place, and how long did it take to index 7 million?
- mo0 12y agoCrawling the internet for mp3s. It took us a couple of months to get to 7m.
- tixocloud 12y agoThat's an absolutely creative way of doing it! Congrats. I'm interested in how you did it. For example, how do you exactly "crawl the internet"? Did you have a bunch of sites that you've pre-selected and then just crawl those or did you actually follow through on links? Thanks.
- tk42 12y agoOne-click hosters and cloud hosters with sharing options(such as docs.google.com), several music streaming sites, usenet. Also the amount of mp3s hosted on regular webservers which are indexed by and easily found using the usual search engines is mindblowing
- rubicon33 12y agoI am also interested in knowing more specifics about your crawling process (if you can divulge). Did you just crawl random sites and search for .mp3 content on them? Or did you have a set of pre-defined search sites to craw?
- GarethX 12y agoIt'd be good to have an option to provide an email address which results could be sent to once it has them, rather than keep checking an open tab.
- mo0 12y agoGood idea we'll keep it in mind. Bookmarking the ident page and coming back later to check on results will work too