3 ms·
That's why you go through the process of validating your model (in the example you provided, checking for multicollinearity), remedying any issues, and then usi
by i-am-charmander 8y ago
That's why you go through the process of validating your model (in the example you provided, checking for multicollinearity), remedying any issues, and then using the model.
Modern software often adds very useful layers of abstraction onto existing processes and patterns. This is especially the case in the realm of machine learning software. Libraries like scikit-learn, Keras, and many others are outstanding pieces of work, and make it very easy to rapidly build and deploy ML models. However, this ease-of-use can actually be a detriment, especially to ML newcomers.
In particular, it is so easy with these types of ML libraries to do something like `from sklearn.linear_model import LinearRegression; model = LinearRegression(); model.fit(Xtrain, ytrain)`. This is great if you're trying to scalably test many different algorithms and configurations to see what predicts best. This is not so great if you're looking to test and validate some of the statistical assumptions of your model, especially with linear/logistic regression models. As an example, Python's StatsModels library will automatically warn you if certain assumptions of a linear/logistic regression model are violated/close to being violated, which could led to inappropriate conclusions/inference from the model. scikit-learn does not do this. If you have massive multicollinearity in your model (a phenomenon which can affect the reliability of individual-coefficient t statistics and the signs, positive or negative, associated with the coefficients), scikit-learn won't tell you that, and it will be on you to recognize the potential for multicollinearity occurring and remedy the issue.
Not to pick on scikit-learn, but their linear_model regression classes also don't provide p-values and standard errors associated with each predictor, common things that basic statistical modeling packages usually provide. But note that scikit-learn's goal is to provide an easy interface with which to do machine learning - not traditional statistical modeling. The ML community is known for placing emphasis on raw predictive performance of models and forgetting about validating the statistical assumptions associated with those models.