5 ms·
This is interesting: A Kaggle competition to predict loan defaults gives an extreme example. This competition had 100s of raw features. For privacy reasons, th
by valgaze 8y ago
This is interesting:
A Kaggle competition to predict loan defaults gives an extreme example. This competition had 100s of raw features. For privacy reasons, the features had names like f1, f2, f3 rather than common English names. This simulated a scenario where you have little intuition about the raw data.
One competitor found that the difference between two of the features, specifically f527 — f528, created a very powerful new feature. Models including that difference as a feature were far better than models without it. But how might you think of creating this variable when you start with hundreds of variables?
The techniques you’ll learn in [OP's cours] would make it transparent that f527 and f528 are important features, and that their role is tightly entangled. This will direct you to consider transformations of these two variables, and likely find the “golden feature” of f527 — f528.
- deleted 8y ago[deleted]
- mhuffman 8y agoAn interesting thought experiment is what if it turns out that the "golden feature" of f527 -- f528 is race or gender? Here is a talk about times when that happens exactly: https://www.chrisstucchio.com/pubs/slides/crunchconf_2018/slides.pdf https://www.chrisstucchio.com/pubs/slides/crunchconf_2018/sl...
- tomnipotent 8y agoThis. This so hard. Data captures existing biases, so models will likely lead to more of the same outcomes. We see this already with crime statistics in the U.S., not to mention college acceptance rates, home & auto interest rates, even job offers. We may have gone to the moon, but by and large scientists of today are just as susceptible to mistakes as prior generations that had "solid evidence" like how European ancestry correlated with higher IQ. As good as it feels to wrap ourselves in the warm fuzzy blankets of "being woke" and thinking we're somehow different, I'm going to keep remaining skeptical of the results until I know exactly how they were reached (and the black box of deep learning makes that a little tricky).
- Iv 8y agoThanks that was a really interesting presentation. And to people who did not get to the end, I'd like to point out the last slide which is a nice tl;dr: "Early on I said I wouldn’t be giving any ethical prescriptions.I will, however, give one meta-ethical prescription: formalize your ethical principles as terms in your utility function or as constraints.It is nearly certain that tradeoffs between these principles exist, and if we don’t acknowledge this, we run the risk of unknowingly engaging in bad actions."
- drilldrive 8y agoThis won't ever take off once the media finds that the machine learning is biased towards some specific sex or race (particularly if such demographics are normative historically for the country at large). Just look at the contentions of prison reform automated systems and the like.
- theothermkn 8y ago> This won't ever take off once the media finds that the machine learning is biased towards some specific sex or race (particularly if such demographics are normative historically for the country at large). I'm not sure I understand. If it is discovered, or "found" (by the media, I guess), that "these systems" are biased, I suppose that should result in them, being fixed, even at the expense of them "tak[ing] off." In your view, is that a good thing or a bad thing? Is it better or worse that "the media" find it? Also, by "normative historically," did you mean "representative?" In other words, do you mean a mathematical norm, as in an average? Or an ethical norm, as in a rule? (Hetero-normative, normative ethics, and so on.) Surely we're not talking about "correct biases?" It's just that my impression is that the language of "normative" in this context, a context which reads as sociological, is most often associated with societal, cultural, or ethical "norms," which are things that are viewed as correct in an "ought" sense. > Just look at the contentions of prison reform automated systems and the like. Can you elaborate? Because I think very few of us have looked at these examples. Which "prison reform automated systems and the like" have failed to "take off" due to a discovery of bias by "the media?" Sorry for all the scare quotes. I'm just trying to piece it together and am having no luck.
- mhuffman 8y agoThis entire document [0] is about exactly what you are trying to piece together. And it talks about a criminal justice algorithm that was "discovered" to be racist. And it explains that these types of things "being fixed" is not really possible in a way that doesn't screw some group over ... you just kind of have to decide who you want to screw over. Have a read it is very interesting. [0] https://www.chrisstucchio.com/pubs/slides/crunchconf_2018/slides.pdf https://www.chrisstucchio.com/pubs/slides/crunchconf_2018/sl...
- AznHisoka 8y agoIf these people were concerned about accuracy and getting things right, why make the variables private at all? Make them public and have your internal team work on the problem (instead of strangers). That way someone with domain expertise can see those variables really are important fundamentally, and add them as features. It’s akin to someone giving you a bunch of anonymous variables to predict stock prices. You would throw out irrelevant variables like the fog level that day, or moon cycles of course. But you would never know if they were anonymous variables. thus you would blindly feed in garbage in your models. Are you concerned about accuracy or privacy? You really can’t have it both ways, imo.
- eanzenberg 8y agoI have a hard time believing this, as a linear combination of features doesn’t add any information for linear models. c(f527 - f528) = cf527 - cf528 which is a constrained version of a linear model af527 + b*f528. If they saw their std err of the feature shoot way down thats very different than improving model performance. For explanations, key is having the input features be as intuitive as possible. So if a subtraction of features is more intuitive than the two feature separately then the explanations will be more intuitive as a whole
- rm999 8y agoNot to disagree because you're totally correct, but the winning model used GBMs (a non-linear model) and a bunch of subtracted correlated features to great success. When using GBMs, subtracting two features can add a lot of value. https://storage.googleapis.com/kaggle-forum-message-attachments/41007/1148/Josef_Feigl_Documentation.pdf https://storage.googleapis.com/kaggle-forum-message-attachme...
- amrrs 8y agoWorth a watch - Winning with Simple linear models https://youtu.be/68ABAU_V8qI https://youtu.be/68ABAU_V8qI
- rm999 8y ago>One competitor found that the difference between two of the features, specifically f527 — f528 I read through the thread where f527 — f528 was discovered, and mostly the finding was that those two features alone were very predictive. Further research found that the two features are highly correlated, but negative of each other. See the write-up of the winner here: https://storage.googleapis.com/kaggle-forum-message-attachments/41007/1148/Josef_Feigl_Documentation.pdf https://storage.googleapis.com/kaggle-forum-message-attachme... >Yasser Tabandeh posted a pair of features, f527 and f528, which can be used to achieve a very high classification accuracy ([1]). Furthermore f527 and f528 are very highly correlated. However, it is not needed to keep both features. Their difference f527 − f528 contains the same amount of information He also explains his method of finding other pairs of "golden" features in a 2-step iterative process. Original thread: https://www.kaggle.com/c/loan-default-prediction/discussion/7115 https://www.kaggle.com/c/loan-default-prediction/discussion/...
- bigger_cheese 8y agoOne thing I have often found is that data context matters for a lot of modelling work. It's all well and good to throw 1000 inputs named F_1...F_1000 into a model and see what gets spat out but I question how valid that is going to be. To give a recent example from my workplace: There is an online gas chromatograph hooked up to a reactor - external modelling (of the throw 1000+ variables at the wall and see what sticks variety) showed the percentage of nitrogen, as measured by chromatograph was strongly correlated to something we'd asked them to investigate. The people doing the modelling who had no experience with basic chemistry latched onto this and started to get really excited they wrote a large report full of recommendations about nitrogen. When my manager saw this he burst out laughing - nitrogen is an inert gas it plays no role in the reaction what the data was actually showing is that nitrogen content varies in response to the concentration of other gasses (H2, Oxygen and CO/CO2) changing - if you recall air is typically 78% nitrogen. We don't control Nitrogen at all it's value is determined based on the other gasses in the reactor. I think this demonstrates pretty well how easily people can see something in the data and jump to some conclusions that while well meaning lack a fundamental understanding. Model explainability is good but input explainability is just as important - important question to ask are often. -What am I feeding into this model? -What is the expected behavior of this input? (I don't know is a valid answer here) -Did the input behave as I expected it to in the model? -If Not why? (This is the interesting and potentially valuable part)
- adrianratnapala 8y agoSo if I am understanding this story, some people found a correlation and mistook it for causation. An obvious no-no, albeit a easy trap to fall into. That's assuming what you were looking for was a handle on how you could get more or less of whatever it was you wanted. But if you just wanted a diagnostic, then the data is telling you that the nitrogen content is (in practice) such a diagnostic. Depending on what you already knew, that might or might not be interesting or useful, but it is hardly nonsense.
- bigger_cheese 8y agoYes that's essentially correct the N2 means something just not what the analysts thought it meant. Their recommendations were misleading and could have saved some time if they'd discussed it with someone familiar with process.
- adrianratnapala 8y agoIs this step one in re-inventing principle component analysis?
- Iv 8y agoI find it is a failure of deep learning that a model is unable to spot these by itself. A dense layer would have a neuron connected to both f527 and f528. A positive weight on the first and a negative one on the second would give f527-f528. The fact that stochastic search is incapable of making these things emerge is problematic.