3 ms·
One of the most incredible parts is that they've already run feature detection on all 100M images/videos and extracted 50TB of: "SIFT, GIST, Auto Color Correlo
by GrantS 12y ago
One of the most incredible parts is that they've already run feature detection on all 100M images/videos and extracted 50TB of:
"SIFT, GIST, Auto Color Correlogram, Gabor Features, CEDD, Color Layout, Edge Histogram, FCTH, Fuzzy Opponent Histogram, Joint Histogram, Kaldi Features, MFCC, SACC_Pitch, and Tonality"
The good part about this for researchers is not only that this saves dozens of CPU-years of computation (back of the envelope, it would take 15 years for my laptop to extract those SIFT features alone), but that any differences in learning/recognition performance on the dataset can be attributed to the algorithms in question, uncomplicated by which researcher engineered the best features for the dataset. On the other hand, it's a challenging dataset to work with because you can't just download it and process it locally as has been traditionally done. I'll be interested to see how many take advantage of it.
- kastnerkyle 12y agoWhy the heck would you do MFCC on images? Mel filters try to replicate the perception of human ears on audio. This looks like buzzword soup to me (SACC_Pitch, Tonality? What the heck?!? These seem like audio features - where are the formulas!). I also don't know about your other conclusion, there is no reason you couldn't download this dataset given enough time/bandwidth/storage to process locally. Most people who will work on this could reasonably store a large chunk locally, if not all (~10 TB). This also assumes that you can't reduce/compress the info any further than what flickr provides and that you require access to the entire dataset - if any of the images are 1024x1024 or larger most feature extractions do not need that kind of fidelity. Heck, you could probably make use of grayscale only to reduce the size by a factor of 3 - ~ 17 TB is feasible (though still pretyt insane) to store locally. ImageNet (~1.2 TB) only took me 45 days on a residential (<20 MB connection), and I wold assume that this dataset would be downloaded by entities with much higher download b/w. I would also assume that many algorithms, like the type that attack CIFAR10 et. al., would also be willing to reduce the dimensionality and recompress, further reducing storage overhead. How big is each image? Also, where are the hyperparameters they used to calculate all of these features? Extracted features aren't really that useful without context/reproducibility. All that said, I think most of these features are decent and the dataset is amazing, but I would rather see them release the raw data set and its PCA/ZCA/other transform - maybe Gabor filtered etc. as well. Lower level preprocessing is more useful for doing representation learning IMO - these higher level features are not that useful for ML algorithm developers. SIFT is patented for heaven's sake! How are we supposed to build algorithms on top of things like that... I am excited about the dataset but feel that there could be more done to truly enable researchers. This feels like a "look how much data we have/look how awesome and used flickr is" thing to me.
- reitanqild 12y ago> Why the heck would you do MFCC on images? Maybe it was the videos? > SIFT is patented All the better then that Yahoo acquired a licence somehow, processed the data and made the results available. You don't need a license to compare keypoints I hope?
- kastnerkyle 12y agoThat is the point - if SIFT is patented (though it is free for research it is not clear that Yahoo got commercial licensing) it is prohibited to generate their keypoints, even if it is to compare to your own! Just because they don't prosecute, doesn't mean it is OK. I guess in this case, Yahoo might have taken the heat for you. Hopefully they got a proper license agreement, as that would be great for everyone.
- reitanqild 12y agoI make an educated guess that Yahoo has a valid license for a public project like this ;-) > it is prohibited to generate their keypoints, even if it is to compare to your own But you can still compare them against each other and against other data sets.
- DanBC 12y ago> Why the heck would you do MFCC on images? Mel filters try to replicate the perception of human ears on audio. First page of Google hits. MFCC Based Face Identification http://www.img.cs.titech.ac.jp/~akbari/pmwiki/uploads/Site/Sangeeta-rep.pdf http://www.img.cs.titech.ac.jp/~akbari/pmwiki/uploads/Site/S... Identification of satellite images based on mel frequency cepstral coefficients http://ieeexplore.ieee.org/xpl/login.jsp?tp=&arnumber=5383270&url=http%3A%2F%2Fieeexplore.ieee.org%2Fxpls%2Fabs_all.jsp%3Farnumber%3D5383270 http://ieeexplore.ieee.org/xpl/login.jsp?tp=&arnumber=538327...
- kastnerkyle 12y agoFrankly, both papers you linked give no justification for using the mel filterbank instead of any other filtering/frequency reduction before doing the IDCT to get the cepstrum. This is what I mean! Why take a filterbank designed for audio and apply it to images except to use the buzzwords? In any case, at least I know it is used by researchers (even if I don't agree with the why). Thanks!
- joshvm 12y agoYou wouldn't download the lot, methinks. Not unless you have a big ole cluster to handle it. The index is only 12GB and contains enough metadata that you can whittle it down to a subset, pull the comments and filter based on those, and ultimately produce a list of photo IDs to grab from the collection. That's a couple of day's work for a grad student, it's not even Big Data.