3 ms·
"One of the ML mistakes that you will hurt you the most: not collecting the right features. If the information you need isn't in your data, you won't recover it
by mxwsn 4y ago
"One of the ML mistakes that you will hurt you the most: not collecting the right features. If the information you need isn't in your data, you won't recover it with a better model."
The word 'information' is doing some heavy lifting here. There is nuance - what information could be learned or computed from the data? We can only really guess. When we're wrong about this, then we will design models imperfectly.
Is there enough information in internet text scrapes to solve unseen high school mathematics word problems? The feature engineers of olde would have said no. But there is, to a surprising degree.
- kazinator 4y agoThe way of speaking about "data" being an undifferentiated mass of bits, from which you get "useful information" is pretty well-worn, widely observed piece of rhetoric. (If it's wrong, the scope of the mistake reaches far beyond this article.)
- npalli 4y agoThis approach of data maximalism seems to work well primarily for NLP and CV. However, there are lot of domains where throwing every feature without some consideration of information will not get you good results. To take a simple case, you are asked to model (forecast) the sales of an item at a retail store. Jamming in features like - color of people shirts, average length of their nails, the duration of traffic signals in front of the mall, the length of the store name etc.. will not lead to good results. In contrast, seasonality, prices, promotions, weather etc. will most likely lead to good results. Why does the second set of features work while the first doesn't? I like to think of the second set having a model of causality, to figure that out you need some domain expertise and experience. Without that expertise, the feature set of non-causal factors is almost infinite.
- bloaf 4y agoI think the real biggest "gotcha" is not understanding where the control loops are in your data. The classic example is the "air conditioned room" thought experiment. Imagine feeding an AI model a dataset with a bunch of datapoints like this Room hot | AC on full blast Room cold | AC off Room warm | AC on 50% Then asking the AI how to make a hot room cold. The AI is very likely to say "turn off the AC" because "AC off" is a feature of cold rooms. Maybe we think that if we give the AI more information, it could do better. Lets imagine that our room has a window and an oversized AC. We'll now be feeding it info about the source of heat as well as the source of cold, so surely it will get the right answer this time: Window open | AC on full blast | Room cold Window closed | AC off | Room cold Window 50% | AC on 50% | Room cold This dataset is going to have the AI thinking that neither the window nor the AC impact room temperature, and that you might be able to close a window by turning off the AC. The key thing to understand is that control loops change the "physics" of your system and that knowledge about the control-physics does not transfer to the uncontrolled-physics.
- nl 4y agoMachine learning can deal with this very easily by incorporating time as a feature.
- cwilkes 4y agoThat’s the point: without knowing that time is an important feature your ML model will come to the wrong conclusion.
- nl 4y agoI don't think that was the point the OP was making. I think they were trying to argue that ML cannot learn control loops or maybe that the system needs to "understand" physics (Not sure - it wasn't really clear to me). The fact there is a control loop here doesn't actually seem important if you represent the features correctly. I'd argue that data representation is actually what you need to get right (ie, in this case the behavior in the time domain).
- skybrian 4y agoThe correlation exists. If you see that the AC is running a lot then you can predict that the outside temperature is warmer. What machine learning algorithms even try to determine causation?