3 ms·
Background: The technique is actually pretty fascinating. This is something that's been well understood by the cryptography community for decades, but is someh
by solve 11y ago
Background:
The technique is actually pretty fascinating. This is something that's been well understood by the cryptography community for decades, but is somehow just recently being fully appreciated by the ML community. See here:
http://blog.mrtz.org/2015/03/09/competition.html http://blog.mrtz.org/2015/03/09/competition.html
https://www.kaggle.com/c/restaurant-revenue-prediction/forums/t/13950/our-perfect-submission https://www.kaggle.com/c/restaurant-revenue-prediction/forum...
Summary -
Submitting guesses to a system that gives you back scores for your guesses, will quickly leak out enough information that you can reverse engineer a huge number of hidden numbers/labels in surprisingly few iterations, e.g. 700 iterations to covertly extract 10,000+ real numbers with high precision. This surprisingly rapid convergence is a bit reminiscent of the birthday paradox.
Further, this not only lets you win the against the "test" dataset, as apposed to the final "validation" set, but this allows you to significantly increase the data available to you to train your model on, since now you can train your model against both the "test" and "training" datasets.
Layman summary -
ML breaks datasets into 3 partitions "test", "train", and "validation". In cases where they're evenly split, this technique can double the training data you have access to, which is a massive advantage in ML competitions where scores differ by tiny amounts.
Moral judgement -
My opinion, this moral argument is misdirecting the attention from where it needs to be. Yes, it's bad what occurred here. But at this point, in 2015, and with tools readily available to crack this problem effortlessly, it's inexcusable for contests to allow so many scoring reports against their validation sets anymore. It's no longer a question of whether contestants will do it, but how many of them will. We'd might as well just let people self-report their scores on an honor system, if we're going to be this overly trusting.
Try creating a contest system like this in the cryptography field any time in the past 3 decades and you'd be insulted and laughed out. Allowing so many scoring reports against the validation set is fundamentally flawed. The only solution is to globally limit calls to the scoring api.
Another proposed solution -
Allowing everyone to see everyone else's guesses & resulting scores against the "test" set, so that everyone is on equal ground for reverse engineering the "test" set, and then globally limiting the number of scoring attempts so that the test set isn't reverse engineered too significantly.
Overfitting the "validation" set actually is not a problem either way, because none of these contests are dumb enough to let anyone score against the validation set at all until the contest submission deadline is over.
- grayclhn 11y agoDo you have a link explaining how the equivalent of these contests are run in the cryptography community? I'd love to read more.
- mcguire 11y agoWhy not keep the validation data secret, give the teams both the training and test data, and make sure they know that their pre-validation submissions are running against data they already have---they're really just testing the submission process?
- kastnerkyle 11y agoAcademia is basically "self-reporting on the honor system". It works generally but there are lots of holes. Ultimately, "trust but verify" is necessary to avoid getting caught in a wave of hype, or at least having someone in your own "circle of trust" say it works. This system leads naturally to elitism and a bunch of other problems which are seen in academia, but it seems better than the current alternatives to me given the current rabid focus on exact percentage score instead of quality/utility of an idea. The "right way" to do it is test once only per model/paper. If you are interested there are a huge number of sneaky ways overfitting can happen in ML [1]. Also interesting that you too see crypto and ML as related - I see them as opposites of the same coin. One tries to pull signal out of noise, the other tries to bury the signal in noise... but special noise. [1] http://hunch.net/?p=22 http://hunch.net/?p=22