5 ms·
I hear this a lot. In my opinion, people overestimate their ability to “understand” non-neural net models. For instance, take the go-to classification model:
by agentofoblivion 8y ago
I hear this a lot. In my opinion, people overestimate their ability to “understand” non-neural net models.
For instance, take the go-to classification model: Logistic Regression. Many people think they can draw insight by looking at the coefficients on the variables. If it’s 2.0 for variable A and 1.0 for variable B, then A must move the needle twice as much.
But not so fast. B, for instance, might be correlated with A. In this case, the coefficients are also correlated and interpretability becomes much more nuanced. And this isn’t the exception, it’s the rule. If you have a lot of features, chances are many of them are correlated.
In addition, your variables likely operate at different scales, so you’ll have needed to normalize and scale everything, which makes another layer of abstraction between you and interpretation. This becomes even more complicated when you consider encoded categorical variables. Are you trying to interpret each category independently, or assess their importance as a group? Not obvious how to make these aggregations. The story only gets more complicated for e.g. Random Forests.
I think it’s best to accept that you can’t interpret these models very well in general. At least in the case of some models (like neural nets), they approximate a Bayesian posterior, which has some nice properties.
- disgruntledphd2 8y agoHow do neural nets approximate a Bayesian posterior? Not snark, would really like some references if possible. On the major point, while I agree with you, its much nicer to be able to show the "top" variables from a model, which is doable from logreg and forests, but is much, much, much more difficult from a neural net perspective. Additionally, as they tend to take longer to train, its harder to iterate with them, and as they fit so very many parameters, I'm generally pretty sceptical as to their generalisability. That being said, in some tests I've run I've been pleasantly surprised at their performance.
- deleted 8y ago[deleted]
- cosmic_ape 8y ago>>How do neural nets approximate a Bayesian posterior? Not sure what GP had in mind, but if a feature x appears in a dataset n times, with pn times with positive label, and (1-p)n times with negative, and your classifier is f(x) which is trained with the "cross-entropy" cost, then the ideal value, that minimizes the cost should be f(x) = p. In this sense, f(x) is the probability of positive given feature. Whether neural nets really realize this and how reliable that is, is another question. But that's the intention of the cross entropy cost.
- beta_binomial 8y agoThis does not make any sense to me and neither did OP's comment about NN's approximating the posterior. In fact, if p were the solution then that would simply be the maximum likelihood estimate, which would not include the p(theta), or the prior, and hence would not be Bayesian.
- cosmic_ape 8y agoWell, p definitely is the solution in the case I mentioned. It is indeed the maximum likelihood solution. You could incorporate prior info about theta via a regularization term, if so inclined. What does not make sense in this? Not sure what the OP meant, but I though it might be useful to mention how estimators may be interpreted as anything probabilistic at all. Often, arbitrary numbers between 0 and 1 are termed "probabilities", but in this case there actually is some proportion or probability to which f(x) should ideally correspond.
- conjectures 8y agoIn general they don't in the sense of being fully Bayesian. The mental gymnastics are that: - The objective function of the neural net was a likelihood. - The prior was improper. In which case the net is a MAP estimate. A MAP estimate will not give you good uncertainty quantification. Given the application to risk modelling, this seems unlikely to be a trivial departure from a fully Bayesian method.
- agentofoblivion 8y agoHere is a paper on the topic of NN = Bayesian Posterior Probabilities: http://www.ee.iisc.ac.in/new/people/faculty/prasantg/downloads/NeuralNetworksPosteriors_Lippmann1991.pdf http://www.ee.iisc.ac.in/new/people/faculty/prasantg/downloa... Re: comment on showing "top" variables from a model, I agree this could have utility. But I would add that the devil's in the details, and there are multiple ways to calculate importance values, each of which has its own nuances and pros/cons. For instance, how do you compare the importance of a categorical feature to a float feature? Do you one hot encode and then add their individual importances, take the average, or something else? Although sampling from the columns is meant to help deal with feature correlation, under what conditions is this effective and how do you know if your feature importances are safe? Moreover, how does this column sampling work in the context of one-hot-encoded categorical features? This is all a way of saying that while you can devise methods for coming up with metrics, and then assign them handy titles like "Feature Importance", the reality is that these things are pretty nuanced and limited, and upper level management might be fooling themselves by thinking they're "interpreting a model" if they don't recognize the limitations and nuances involved. Or to put it another way, to say "I better understand this model because you gave me a feature importance list and a partial dependence plot," is a dangerous over-simplification.
- projectramo 8y agoIf you do some kind of PCA, that should help work out the correlations between the coefficients, though you are right these can be hard to interpret (although sometimes they have a natural interpretation). And typically, you would demean, and divide by the standard deviation. I guess the coefficients are harder to interpret in this second case than if you did not transform them, but they're still interpretable.
- a-dub 8y agoOne trick I like to do in PCA is to sweep each of the variables, one at a time, in the low-d subspace and then project back up. Looking at these sweeps in the original space gives some hand-wavey "intuition" for what each low-d variable is capturing.
- gbrown 8y agoSure, interpretation must be done with care, but that's one of the primary goals of statistics. I think the bigger distinction you're missing is the difference Breiman draws between algorithmic models and data models. If you use logistic regression (and hopefully a careful study design) to describe a plausible data generating model, you can get good, interpretable inference out of it. If you use the same tool (or penalized equivalent, or RF, GBM, NN etc) as a prediction algorithm on unstructured or poorly structured inputs, you're not going to have your lunch and eat it too (get good, interpretable inference along with robust prediction). It's also unclear to me what you mean when you say that NNs uniquely approximate a Bayesian posterior, or why that's a good thing without knowing more about what posterior you're talking about. You could do a Bayesian logistic regression and get an actual posterior, and it would not remove the interpretation challenges you raise.
- d--b 8y agoPeople making/exploiting models in firms such as BlackRock, Renaissance, 2Sigmas and so on, know what they're doing. Really.
- AznHisoka 8y agoI don't disagree that there are companies that definitely know what they're doing, but on the other hand, we shouldn't put anyone in the financial industry on a pedestal. After LTCM, and the financial crisis of 2008, I've learned my lesson. Half these guys don't have a clue what they're doing.
- d--b 8y agotrue, that's also why it's kind of refreshing to see a big one stepping back and questioning what they're doing.
- riazrizvi 8y agoDon’t forget LTCM. Very smart.
- i-am-charmander 8y agoThat's why you go through the process of validating your model (in the example you provided, checking for multicollinearity), remedying any issues, and then using the model. Modern software often adds very useful layers of abstraction onto existing processes and patterns. This is especially the case in the realm of machine learning software. Libraries like scikit-learn, Keras, and many others are outstanding pieces of work, and make it very easy to rapidly build and deploy ML models. However, this ease-of-use can actually be a detriment, especially to ML newcomers. In particular, it is so easy with these types of ML libraries to do something like `from sklearn.linear_model import LinearRegression; model = LinearRegression(); model.fit(Xtrain, ytrain)`. This is great if you're trying to scalably test many different algorithms and configurations to see what predicts best. This is not so great if you're looking to test and validate some of the statistical assumptions of your model, especially with linear/logistic regression models. As an example, Python's StatsModels library will automatically warn you if certain assumptions of a linear/logistic regression model are violated/close to being violated, which could led to inappropriate conclusions/inference from the model. scikit-learn does not do this. If you have massive multicollinearity in your model (a phenomenon which can affect the reliability of individual-coefficient t statistics and the signs, positive or negative, associated with the coefficients), scikit-learn won't tell you that, and it will be on you to recognize the potential for multicollinearity occurring and remedy the issue. Not to pick on scikit-learn, but their linear_model regression classes also don't provide p-values and standard errors associated with each predictor, common things that basic statistical modeling packages usually provide. But note that scikit-learn's goal is to provide an easy interface with which to do machine learning - not traditional statistical modeling. The ML community is known for placing emphasis on raw predictive performance of models and forgetting about validating the statistical assumptions associated with those models.
- stdbrouw 8y agoHm, I agree with what you're saying, but your example is pretty whack: (1) in a logistic regression the coefficients are on a log scale ergo the ratio between exp(2) and exp(1) is actually x2.7, not x2; the bigger issue is that you have to compare the strength of the association to how easy it is to move that lever, e.g. men might like our marketing message more than women, but it's not like we're going to get people to change gender. (2) moderately correlated predictors do not bias or otherwise complicate the interpretation of regression parameters, it's only unmeasured confounders that do, that is, correlations between variables where one of the variables is not included in the model.
- claytonjy 8y agoIn (2) it's actually worse than that; any relevant predictor left out of the model affects the betas of all other predictors, even if it's uncorrelated to all other predictors. This is a big thing most people miss when trying to interpret logistic regression the way they do linear regression. Logistic regression estimates are _conditional_ on the model spec in a way linear regression estimates are not.
- gbrown 8y agoI'm not sure I follow your second point. In both multiple linear regression and multiple logistic regression, the coefficients are interpreted conditionally - a one unit increase in X^{(i)} is associated with a B^{(i)} unit change in f(E(Y)), conditional on the values of the other covariates. Using a logit link function does remove the result from the raw scale of Y (thus, f), but in my mind the "model" is more than just the data distribution - it's whatever form your linear predictor takes. This is the problem the poster originally pointed to - interpreting a variable conditional on the fixed values of all others doesn't always make sense when there's strong correlation among predictors.
- zwaps 8y agoI still don't get your issue at all. There are tons of different models and implications, and it depends entirely on your question how you interpret the model. Often you will go after average partial effects in your sample. Or you have some correlated variables in mind such that you can just plot the partial effects. Sometimes you have different models, where you will plot a distribution of partial effects around a mean based on some prior assumption... I mean, it really depends. If you impose a more complex relationship, then of course you can not put everything in a single number. But no matter what you are interested in, such a model will give you the possibility to exactly determine the measure. And what you are saying is also not correct. I may very often be interested in fixed values. Doing this regression, and not a quantile regression, for example, means I am somehow interested in a conditional expectation. That means I probably care about some sort of average impact, perhaps for some fixed subgroups. But those averages are fixed values... I think the point is that in these models, we know exactly what we go after, how to get there, and what it means. We know exactly when our inference may fail. If we both care about average effects, and I can convince you of my identification assumptions, then there is really no mystery left as to what my estimates mean. In with Deep Learning, this is still more difficult.
- fjp 8y agoCorrecting these issues in a logistic regression is something you learn in undergrad stats classes and something any professional would do as part of their modeling.
- xg15 8y ago> I think it’s best to accept that you can’t interpret these models very well in general. But then, the classical question: How do you debug the models? How do you know they are actually predicting what you think they are predicting?
- dboreham 8y agoIf ad revenue increases, who cares, right?
- beta_binomial 8y agoThere are legitimate reasons why stats wins out on interpretability. 1. Scaling is not hard 2. It is obvious to me "how to make these aggregations", but that is because I know statistics. Categories of variables treated as random effects can be interpreted both as a group (via variance parameters of the random effects) and individually via coefficients. 3. Bayesian estimation can even account for correlation of parameters and include it in the posterior prediction.
- zwaps 8y agoI am not sure those arguments hold so much water. Classic regressions allow you to explain the marginal effects very well. It doesn't matter much how variables are correlated. If you saturate the model with interaction effects, you can get an accurate (in the sense of the model) prediction of the marginal effect of any variable as a function of others. This is very interpretable. Furthermore, nowadays a lot of techniques are about estimating the causal effects based on assumptions in your data. You could use things like DID or synthetic control, use natural experiments, and so forth, to get a good idea of the causal "treatment effect" of your variable of interest, and you can even do this in a semi-parametric or non-parametric approach. Often, estimating the linear approximation of the marginal on a conditional expectation is "good enough" to learn how things are connected within your data. And in the end, getting this sort of causal effect of a variable is what we are really after in environments where the DGP process isn't simple. In that sense, this type of research is very compelling. Scaling and other issues are of course important, but taking them into account is rather simple... To be sure, a lot of work (for example in econometrics and elsewhere) is proceeding on causal inference of deep learning models, but it is probably also fair to say that right now, classical models are far easier to interpret, especially if you are interested in answering qualitative questions.
- jgalt212 8y ago> But not so fast. B, for instance, might be correlated with A. In this case, the coefficients are also correlated and interpretability becomes much more nuanced. And this isn’t the exception, it’s the rule. If you have a lot of features, chances are many of them are correlated. Aren't we supposed to remove correlated predictors from models? Or we can use techniques like PCA? But then, once again, we fall into the realm of unexplainable.
- shoguning 8y agoYou can remove correlated predictors or do stepwise regression to get a minimal model with more intelligible coefficients. The problem is it's still difficult to determine direction of causality. That's why controlled experiments are so important.
- BenoitP 8y ago> But not so fast. B, for instance, might be correlated with A. Exactly. Rendering a model explainable is an active field of research at the moment. One way to tackle this is to use Shapley values from game theory. Here is a compilation of techniques to render a model explainable: https://christophm.github.io/interpretable-ml-book/shapley.html https://christophm.github.io/interpretable-ml-book/shapley.h..., and they have an elegant way to talk about these values: > The Shapley value is the average marginal contribution of a feature value over all possible coalitions. Now, it doesn't identify correlated variables (which you can do with other techniques), but will balance the influence in a robust manner. ---- > The story only gets more complicated for e.g. Random Forests. Funny you should say this. Recently there was a challenge for explainable machine learning: http://explainable.ml http://explainable.ml My proposal to the challenge was based on Random Forests. Here is the code: https://github.com/benoitparis/explainable-challenge https://github.com/benoitparis/explainable-challenge, here is the paper: https://github.com/benoitparis/explainable-challenge/raw/master/Decision%20tree%20node%20participation%20as%20a%20tool%20for%20interpretability.pdf https://github.com/benoitparis/explainable-challenge/raw/mas... I have a neat visualization for exploratory data analysis (available in the paper and in a notebook) that I'm proud of. ---- By the way if anyone is hiring here, I'm available for talking about your problem and see if my take on explainable machine learning can help you.