4 ms·
Meh.. I'm not so sure about this. Didn't people claim "cheating" back when the first compilers started doing data flow analysis too?
by king_of_nouns 11y ago
Meh.. I'm not so sure about this.
Didn't people claim "cheating" back when the first compilers started doing data flow analysis too?
- daveguy 11y agoIt's a contest. It has specific rules to prevent a very specific type of synthetic advantage (training to the test set). This was about as cheating as it gets. That's why the company fired the guy in charge of the group that cheated.
- king_of_nouns 11y ago> It's a contest Yes, who cares? The bigger picture is advancing the field, not scoring some bigger number in some artificial environment. > This was about as cheating as it gets. Like I said, so was data flow analysis originally. -- The point is to not just dismiss this as "cheating" but take a closer look at how the current benchmark is flawed and how this sort of shortcut might be useful.
- rsfern 11y agoThe blog post explains why the behavior in question is cheating, not a "useful shortcut". Overfitting a model isn't a real advancement to the field. The point is to take a closer look at the current benchmarks and potentially come up with better ones.
- daveguy 11y agoThis sort of "shortcut" exploits a well known flaw in the training any machine learning or regression model using any algorithm. It is the problem of overfitting. The problem is that given enough parameters you can tweak those parameters so that your model exactly fits a target data set (including the random noise). The reason for excluding and limiting access to the test set is so that parameters can't be modified such that the model is adjusted to the specific test data set. It is a well known problem and exploit that the rules are designed to guard against. Here is the wiki page on overfitting: https://en.wikipedia.org/wiki/Overfitting https://en.wikipedia.org/wiki/Overfitting Say the test data set, by chance, has 1% more dogs than the training set. By tweaking your algorithm to guess "dog" an extra 1% of the time you may be able to get .1% increase in your success rate. Also, that tweaking isn't a manual process it's part of the training process based on feedback from the results of your test set score. The improvement is enough to "win", but it's not because your algorithm is better. That goes along with the argument that the data set is at its end of life. Maybe we're at the point that gaming the system is the only way to eek out the .1% needed to "win". In that case it's time to move on to a tougher test. EDIT: I'd like to point out that the ImageNet competition is continually on top of the "time to make it more challenging" aspect. They introduced localization in 2011 (identifying not just what, but where items are). The 2015 competition includes, for the first time, recognition and localization tests in video clips.
- scott_s 11y agoTo support what the others have said, but put in different words: the point of the rule is so that it will be more likely that breakthroughs in contest results translate to breakthroughs in the field itself.
- tokipin 11y agoTo give an extreme example of overfitting, let's say you have a test set which has these examples: B -> 11, H -> 5, M -> 11. Well, you can solve these easily. Just make a function like this: function solve(letter) { return {B: 11, H: 5, M: 11}[letter]; } machine learning 4Head . Of course, the point is to solve for examples you haven't seen yet, so this "solution" isn't amusing anyone. That said, overfitting commonly happens even without trying to overfit, and people use techniques to minimize it.
- ska 11y agoYou might have a point if there was any sort of parallel between what they are doing and data flow analysis, but there really isn't. The parallel would be more like adding a detector for certain benchmarks in your compiler, and outputting hand tuned assembly for that case ... except even worse than that, because you'd have to implement the detector and assembly generation in such a way that it made your compiler behave worse on general input.
- titanomachy 11y agoI'm not sure if the analogy was intentional but your comment made me immediately think of the Volkswagen emissions scandal.
- ska 11y agoThere is a parallel, but only to the first part. What VW did was cheat on benchmarks, much like certain driver vendors and compilers have been known to do. But what we're talking about here is much, much worse from a design point of view. It specializing your system for the benchmark, in such a way that you are actually making it worse in general.