4 ms·
This is the imbalanced class problem that comes up often in the real world. You can't optimize for overall accuracy, otherwise you'd have classifiers that alwa
by binalpatel 10y ago
This is the imbalanced class problem that comes up often in the real world.
You can't optimize for overall accuracy, otherwise you'd have classifiers that always just predict guilty (or not sick, or not fradulent, and so on) because they'd achieve 95%+ accuracy.
- hammock 10y agoSo what is the metric to optimize?
- binalpatel 10y agoI use AUC often, and always look at the confusion matrix. Metrics based off of it (recall, precision) are both useful to use as well. In the end the probability cutoffs you choose are more often than not based on business context and costs associated with various actions. For example - if we're trying to predict whether a customer is going to leave or not, I'd choose a different threshold based on whether we're mass e-mailing people, or if we're having customer service reps call people, since the underlying costs for each action are so different. In the first case it's fine to cast a wide net, and e-mail lots of people who we're less certain about, in the second we'd optimize to target people we're most sure are going to leave, since it costs so much more to reach out to them.
- jpfed 10y agoMaybe mutual information?
- joeyo 10y agoIn my opinion d-prime: https://en.wikipedia.org/wiki/Sensitivity_index https://en.wikipedia.org/wiki/Sensitivity_index
- matk 10y agoAUC is robust to class imbalance. It can be understood as follows: given an instance of the positive class and an instance of the negative class, it measures the probability of your learned model being able to distinguish between the two.