5 ms·
I should have read the comments, I just asked the same thing. I always wondered how that worked, though - you don't submit an algorithm, do you? Just a spreadsh
by dasboth 11y ago
I should have read the comments, I just asked the same thing. I always wondered how that worked, though - you don't submit an algorithm, do you? Just a spreadsheet with your predictions, in which case how can they evaluate your algorithm on another test set?
- deleted 11y ago[deleted]
- qznc 11y agoPublish the final test set right before the deadline. Alternatively, require the participants to provide an API and call them, instead of them submitting things.
- p4wnc6 11y agoYou get to see the feature vectors of the test set, or at least submit your algorithm to execute upon them. If you can see the features of the test set but not the labels or target variable for each feature vector, then you can't use it to help you (e.g. the problem of modeling the distribution over outcomes for a given feature vector (whose target outcome you don't know) is the problem you're solving). But if you can repeatedly run different models against the test set and you do get to see a score, you could do something like random parameter searching or other optimization ideas to tune your algorithm to be highly overfitted to the test set. Another way to avoid this would be to develop several test sets that are roughly "equivalent" in terms of the distributional properties, and then randomly change the test set periodically, or change the test set right after the final submission deadline, to discourage people from pursuing overfitting.
- dasboth 11y agoThat makes sense. I was thinking of the multiple test set approach, but rather than forcing people to re-submit on a new test set near the end (if that's what you meant), the organisers could just return the score based on a random fraction of the test set. 'Cheaters' would then overfit to just half of the actual test set and this would be apparent when the final scores (on the other half) are revealed. I suspect this is what Kaggle do, as they don't make you re-submit at the end (AFAIK).