4 ms·
>I came to the conclusion that area under ROC is pretty much garbage. I cannot disagree more. IMO AUC is a rare great metric: it's principled, useful, universa
by rm999 9y ago
>I came to the conclusion that area under ROC is pretty much garbage.
I cannot disagree more. IMO AUC is a rare great metric: it's principled, useful, universally applicable (e.g. invariant to class imbalances), and easy to explain adequately to non-statistician ("probability of choosing a positive sample over a negative one").
>What you really want is two separate steps: estimate the probability that instance A belongs in classification X, and then a decision step where you decide how to classify A based on a loss function (this varies depending on how harmful a false positive is).
Yes, but those are two inherently separate steps and should be measured separately. AUC is a metric of the first step (the "model"). The second step is a business decision and will often be made separately from the modeling process, by different people, with a different cadence, and with a different goal in mind.
For example if I am designing a model to find a disease, I just want to make the best prediction I can, which is cleanly measured by AUC. Then, when it comes to actual diagnosis, someone else will choose cutoffs based on various factors like false positive costs (treatment cost, human toll), false negative cost (disease damage, death toll), supply of treatment, etc. I can picture scenarios where the same model is used for decades, but the cutoffs change seasonally.
- mikebenfield 9y ago> invariant to class imbalances I think this is a red herring that comes from not thinking probabilistically. If the distribution of your training data does not resemble the distribution of your real-world data (or cannot be made to resemble it), you're just guessing anyway. If it does resemble the real distribution, then you want those class imbalances. In fact, I tend to think that the fact that ROC is "invariant to class imbalances" is a significant downside: it means in some sense your score is just as sensitive to things that rarely happen as it is to things that happen all the time. > "probability of choosing a positive sample over a negative one" I find the practical implications of this pretty opaque, and it's never been clear to me whether this is measuring anything I actually care about. As far as I know there aren't theoretical guarantees that a better AUC score means anything real. I haven't thought deeply about it, but I am reasonably sure I could find some simple examples illustrating how to "cheat" AUC by getting a higher score with predictions that are worse in any practical sense. I still like the Brier score: just give me a number indicating how well my estimated probability predictions do on a test/validation set. There are even theoretical guarantees about it, because it's a proper score function.
- jhokanson 9y agoBut isn't it possible to design multiple models, where the judgement of the "best model" is dependent on your goals (e.g. reduction of false positives vs false negatives)?