3 ms·
It's quite sad that this post is even necessary. That said, having a proper training/cross-validation/validation setup is sometimes not that obvious, as you hav
by pyduan 13y ago
It's quite sad that this post is even necessary. That said, having a proper training/cross-validation/validation setup is sometimes not that obvious, as you have to stop and think about possible sources of contamination -- some sampling biases, for instance, can be quite tricky to detect, or your algorithm design might be flawed in some subtle way.
Personally, I wish people emphasized more the importance of a general understanding of econometrics when doing machine learning. In most of the introductory courses I've seen, the link between both field is never made explicit, despite the obvious analogies (coincidentally, there was an article by Hal Varian on the front page two days ago that discussed how both fields could benefit from sharing insights [1]).
Understanding the idea behind minimizing generalization error is one thing, but I find that thinking in terms of internal/external validity and experiment design often gives people a more intuitive understanding of validation procedures, both regarding why and how we should do it.
The same goes for understanding effect size, confidence intervals, causality (and causality inference), and so on.
[1] https://news.ycombinator.com/item?id=6870387 https://news.ycombinator.com/item?id=6870387
- zmjjmz 13y ago>stop and think about possible sources of contamination One great one from my Machine Learning professor was an assignment where we were required to normalize our data to [0,1]. After doing this and then going through the typical cross-validation cycle, he had us try and figure out where we contaminated our validation sets. As it turns out, we all normalized our data before splitting it up, which meant that training data influenced testing data. It's a simple fix, but if you've done that and gone to run a large convolutional neural network for a week only to find that you made a stupid error like that, it can be pretty painful. (Especially since the bad generalization error might not be obvious until you use it the model in production)
- im3w1l 13y agoMaybe one could benefit from a sort of blinding procedure, where the person designing the learner is never allowed to even look at the validation data.
- jfim 13y agoIf both your training and testing datasets are representative of actual data, wouldn't the normalization function be nearly equivalent in both datasets?
- azmenthe 13y agoI'm a bit late to the conversation but I agree with you and just wanted to add my quick two cents. I used to work in algorithmic trading (the kind which aims build consistent viable portfolios, not the HFT arms race). This of course relies heavily on building your model, which can be anything from some simple linear regressions to more advanced techniques more commonly associated with the buzz word of machine learning, this applies to all predictive methods. You begin searching the training data to find optimal model parameters and then verifying performance on the validation set. The number ONE mistake I saw most was that when you get bad results on the CV set, going back to step 1.5 instead of just throwing the whole model out. To take your same core idea, tweak it slightly, add/remove a few parameters and restart the process. Unfortunately doing this enough times and your CV set starts to become the training set. Thus leaving your true validation set the day you turn it on live in production with real money. It's never a good feeling to see your positively skewed returns in your training, testing and "CV" set morph into essentially a mean zero random distribution in production. This was quite an important lesson to learn for me.