4 ms·
It depends on the competition. For this competition it was to be expected as: There was little training data and there were a lot of outliers (the problem was p
by compbio 11y ago
It depends on the competition. For this competition it was to be expected as: There was little training data and there were a lot of outliers (the problem was predicting restaurant profits). Luck played a huge role here. The 9th place team likely had a mediocre highly-variant model that got a very lucky leaderboard split.
This forum post discusses competition variance and poses a metric to quantify "leaderboard shake-up": https://www.kaggle.com/c/liberty-mutual-fire-peril/forums/t/10187/quantifying-leaderboard-shake-up/52919 https://www.kaggle.com/c/liberty-mutual-fire-peril/forums/t/...
The Public Leaderboards are very helpful though! When you have setup a solid local cross-validation pipeline, and the public leaderboard agrees with your local evaluation, then you can try a lot more algorithms and parameters, without using any submission. Especially when working in teams this is important as you may have only 1 submission every 2 days.
Also, the more advanced Kagglers can use leaderboard feedback to increase model accuracy: Cluster the data sets with objective measures. Apply a modifier (restaurants from this region get 0.95 x previous prediction) and look at the result. If the split between public and private is random, and your clustering is objective, then an improvement on public leaderboard should reflect in private leaderboard.