2 ms·
Kaggle hosts data science competitions. There is a public leaderboard and a private leaderboard. During the competition only the public leaderboard is visible.
by compbio 11y ago
Kaggle hosts data science competitions. There is a public leaderboard and a private leaderboard. During the competition only the public leaderboard is visible. After the competition ends the private leaderboard is revealed and your standing on this private leaderboard decides your final score.
The public leaderboard gives some feedback on your model performance. But when a human is in the feedback loop, then there is a risk of overfitting. Overfitting can be explained basically as: "memorizing the data" or "learning from noise, not signal".
When you overfit, you do well in cross-validation and may do well on the public leaderboard, but your predictions do not generalize well to new data.
What this team did was to submit a lot of predictions and only take the predictions that improved public leaderboard score. Then they'd add slightly random noise and try to submit again. They repeated this until they ranked nr. 1 and left quite a few other competitors scratching their heads: How did they do this? Did they find a perfect ML algorithm? Is there data leakage? Are they cheating?
When the private leaderboard was revealed, this team dropped around 2000 spots. Their good performance on the public leaderboard was purely artificial. The contest was valid and well-organized (this could happen on any Kaggle competition with little data). They did not receive a prize.
There are other benefits to ranking well on the public leaderboard: It helps with teaming up with other high-ranking competitors and you can market yourself (recruiters are pretty interested in the top 10, eventhough the competition has not ended yet.)