4 ms·
Could you expand a bit about what kind of information class imbalance can reveal?
by matmatmatmat 4y ago
Could you expand a bit about what kind of information class imbalance can reveal?
- timy2shoes 4y agoIt's not about revealing information, it's about how you model the problem. Say you're using logistic regression. If you upsample the minority class (never downsample the majority class because that's statistically inadmissible), then your classifier might get better at modeling the decision boundary. However, that comes at the cost of losing the best part of logistic regression: probability prediction and calibration. Sometimes the trade-off is worth it, sometimes not. Depends on the type of problem. For more details I suggest reading Frank Harrell: https://www.fharrell.com/post/classification/ https://www.fharrell.com/post/classification/
- time_to_smile 4y agoThe class imbalance quite literally represents the prior probability in the model. As an example, train a logistic model on imbalanced data with both the real distribution of imbalanced classes and over/down sampled classes and you'll notice that the distribution of the predictions from each model is different. This is hugely important because if you are using the output of your model as an expectation to feed into another model (for example predict expected value of a new user account given P(purchase) * E[value of purchase]), rebalancing will give you the wrong answer. This can be corrected in a logistic model by tweaking the intercept but doing this requires you to make some assumptions, and it's more artful of a process than one would hope. You can also see this if you compare the log likelihood of the true data with the given model and the new data. The model trained on the real data will have a notably higher log likelihood. Then there are separate issues with over/down sampling. If you are over sampling the standard error on your coefficients will be artificially reduced by the over representation of examples from the one class, leading you to be more sure than you should be about your parameters. The point estimates should be more or less the same, but if you care at all about inference this change in standard error is a big deal. I believe, but would have to double check, that this would likewise impact the effects of regularization on your model. If you are down sampling you are quite literally throwing out observations, and using less information than you have when modeling. Additionally you'll get the inverse problem which is artificially increased variance in the parameter estimates. Again, even if you don't care about inference you still have the problem that your regularization is very likely not working how you expect. In my experience nearly all data scientists think they are experts in logistic regression when very few really understand this "simple" model. All of these issues here will manifest themselves in more complex models but understanding how (and how to correct them) is a much trickier question.