10 ms·
> It's a contest Yes, who cares? The bigger picture is advancing the field, not scoring some bigger number in some artificial environment. > This was about as
by king_of_nouns 11y ago
> It's a contest
Yes, who cares? The bigger picture is advancing the field, not scoring some bigger number in some artificial environment.
> This was about as cheating as it gets.
Like I said, so was data flow analysis originally.
--
The point is to not just dismiss this as "cheating" but take a closer look at how the current benchmark is flawed and how this sort of shortcut might be useful.
- rsfern 11y agoThe blog post explains why the behavior in question is cheating, not a "useful shortcut". Overfitting a model isn't a real advancement to the field. The point is to take a closer look at the current benchmarks and potentially come up with better ones.
- daveguy 11y agoThis sort of "shortcut" exploits a well known flaw in the training any machine learning or regression model using any algorithm. It is the problem of overfitting. The problem is that given enough parameters you can tweak those parameters so that your model exactly fits a target data set (including the random noise). The reason for excluding and limiting access to the test set is so that parameters can't be modified such that the model is adjusted to the specific test data set. It is a well known problem and exploit that the rules are designed to guard against. Here is the wiki page on overfitting: https://en.wikipedia.org/wiki/Overfitting https://en.wikipedia.org/wiki/Overfitting Say the test data set, by chance, has 1% more dogs than the training set. By tweaking your algorithm to guess "dog" an extra 1% of the time you may be able to get .1% increase in your success rate. Also, that tweaking isn't a manual process it's part of the training process based on feedback from the results of your test set score. The improvement is enough to "win", but it's not because your algorithm is better. That goes along with the argument that the data set is at its end of life. Maybe we're at the point that gaming the system is the only way to eek out the .1% needed to "win". In that case it's time to move on to a tougher test. EDIT: I'd like to point out that the ImageNet competition is continually on top of the "time to make it more challenging" aspect. They introduced localization in 2011 (identifying not just what, but where items are). The 2015 competition includes, for the first time, recognition and localization tests in video clips.
- scott_s 11y agoTo support what the others have said, but put in different words: the point of the rule is so that it will be more likely that breakthroughs in contest results translate to breakthroughs in the field itself.
- tokipin 11y agoTo give an extreme example of overfitting, let's say you have a test set which has these examples: B -> 11, H -> 5, M -> 11. Well, you can solve these easily. Just make a function like this: function solve(letter) { return {B: 11, H: 5, M: 11}[letter]; } machine learning 4Head . Of course, the point is to solve for examples you haven't seen yet, so this "solution" isn't amusing anyone. That said, overfitting commonly happens even without trying to overfit, and people use techniques to minimize it.