4 ms·
First, my earlier comments on a "competing" approach from Facebook may help give relevant context for how to think about these numbers: https://news.ycombinator
by apu 12y ago
First, my earlier comments on a "competing" approach from Facebook may help give relevant context for how to think about these numbers: https://news.ycombinator.com/item?id=7393378 https://news.ycombinator.com/item?id=7393378
Briefly skimming through this paper, it appears that these numbers are not a fair comparison, as this paper uses the unrestricted protocol of LFW[1], whereas the other methods in the ROC curve shown in the paper are using the restricted protocol. As you might imagine, the latter is more restrictive -- specifically in terms of amount of training data allowed. And as I mentioned in my previous comment, training data is king in these kind of systems -- more is always better.
To go slightly out on a limb, I think more significant than the new theoretical model proposed in this paper is probably the use of lots of different types of datasets for training. (Significantly more data >> more complicated models, most of the time.) But I'd have to read the paper much more carefully to be sure about this.
[1] http://vis-www.cs.umass.edu/lfw/results.html http://vis-www.cs.umass.edu/lfw/results.html
- tormeh 12y agoWhat's the point in limiting yourself to small datasets? It forces you to be "clever" about preprocessing, because the levels of freedom in your learning algorithm must be limited to match the size of the dataset. Being clever like this is precisely what we're trying to avoid with machine learning algorithms. It's better to just shove the raw data into a very general algorithm like neural networks and let the data do the configuration. And do that you need _lots_ of data. There is a conflict between an algorithm's performance and its learning speed. Restricting ourselves to low amounts of data means we get algorithms with lower maximum performance than we otherwise would. These are then claimed to be better than algorithms which need more data to generalize well.
- gcr 12y agoThat's a great point, and it's exactly why the LFW results page makes a clear, distinct separation between algorithms that use outside training data and algorithms that do not: http://vis-www.cs.umass.edu/lfw/results.html#notesoutside http://vis-www.cs.umass.edu/lfw/results.html#notesoutside You can probably guess which ones generally do better. (Note that this is separate from the "unrestricted" vs "restricted" issue)
- apu 12y agoThe goal is to allow for "fair" comparisons. But I agree with you in general. I suspect even LFW's creators might; I don't know if they expected the benchmark to still be tested on this many years after creation. I've always felt that all benchmarks in vision should come with expiration dates, because eventually everyone starts implicitly overfitting to them.
- gcr 12y agoAlso, one huge reason why we might "limit ourselves to small datasets" is because that's the only way we can compare many algorithms. Outside training data has a huge influence on algorithm accuracy. For example, since Facebook has access to bajillions of face images that they could use to train their classifier, and since they can't share those images with the rest of the community, it's unclear whether their excellent performance is because they have gobs of data or whether it's really a better algorithm. I bet that a simple algorithm like a naive SVM might do leagues better if we could train it on Facebook's (hidden!) dataset and test on LFW. It just isn't reproducible---a measuring stick is only meaningful if everyone uses the same measuring stick.
- tormeh 12y agoIsn't small datasets the reason for using SVMs? You are, after all, operating in linear space preprocessed with a kernel trick. Far less expressive than a neural network.
- gcr 12y agoWell, isn't it a bit awkward when dataset size is what forces you to select a certain algorithm that you otherwise wouldn't use? "Gee, I would love to train a neural network, but it's just so much data and I'm on a deadline; maybe I should just use an SVM and hope for the best ... ..." That's one of the big surprises behind deep learning these days: it's now feasible to do things like "train a big neural network on bunches of images" in a sensible time. It's an optimization thing as much as it is a machine learning thing, in my opinion.
- tormeh 12y agoI didn't mean that training time is the constraint. It's that when training a more general algorithms (ANN) your hypothesis space has more dimensions than when training a more specialized one (SVM). You therefore need more data to train an ANN than a SVM. The reason for choosing SVM is not that your dataset is too big for ANN, it's that it's too small for ANN.
- gcr 12y agoI think they used the correct ROC curve. Actually, it looks like they cherry-picked a few results from the Commercial and the Academic categories, but everything is from the unrestricted class. We should wait until Erik Learned-Miller lists this on the LFW results page. He'll know how to interpret their results.