3 ms·
I would construct a list of invariances that a classifier should be robust to, then construct (either empirically, analytically or via a machine learning algori
by highd 10y ago
I would construct a list of invariances that a classifier should be robust to, then construct (either empirically, analytically or via a machine learning algorithm using a synthetic dataset) a series of features that are invariant under these transformations. This is sometimes called "fingerprinting" and Shazam uses something like it to look up audio. I'm sure there are many specific approaches described in the literature.
I would then use these features to identify videos that have the same content as these already reported - pick your favorite clustering algorithm. This would ensure that for a single copyrighted video being reported all plausible manipulations that fall under the model are also discovered.
Point being, this is something that could be implemented in less than a month or two by an individual or small team. If Google hasn't implemented it it's because they don't want to, not because it can't be done.
- contingencies 10y agoThat's a particularly robust model of solution. Even simpler solutions could match subsets, eg. normalized/simplified audio waveforms (cheap) or image stills (eg. using imagemagick) at probable change-of-scene points (eg. MP4 keyframes). Many people here could hack together something to do this in under one day.
- highd 10y agoTrue! And one should always KISS to start, at least.
- rifung 10y agoWow thank you for the detailed response. Seems like you are much smarter than I am! How do you know this would perform any better than what they are currently doing though? Do you think this would catch the cases which currently pass their filtering like videos with mirrored images? And would this produce less false positives? According to this post, it seems they already do something similar to what you've described: https://www.quora.com/Does-Content-ID-look-for-a-match-of-the-actual-code-or-a-match-of-the-audio-produced-by-the-code/answer/Dwayne-Charrington?srid=iHLF https://www.quora.com/Does-Content-ID-look-for-a-match-of-th... I suppose in the end this is all going to just be speculation since they haven't disclosed what they actually do to try to find copyrighted content. Can you even predict how well a machine learning or classification algorithm will do without just trying it empirically?
- highd 10y agoWhat you can say is if there are N videos and k types of transforms that are used to "hide" content this would be hypothetically perfect with O(n + k) train examples - naive "reporting only" would be O(nk). This means that you should just need one example of each transform and one known instance of a particular copyright infringing media in order to get high performance. Mirroring etc. would just be one example of the transforms an approach like this would be robust to. It's hard to say what the error rates will be, but you should be on the last 5-10% Pd/Pfa pretty much instantly and then if you want more you can just acquire more true labelled data (if a user can find a copyrighted video on your site, I would think a contractor could, too!).
- rifung 10y agoThis well beyond what I understand so excuse me if I'm misunderstanding but doesn't this imply that they'd need to guess what k transforms might be used? This just seems like a cat and mouse type ordeal which I can only assume is something they would like to avoid as the complexity of and costs of running the system only gets worse, while the complexity does not add up on the side of those making pirated videos.