3 ms·
One of the weaknesses with Jaccard similarity is how it focuses on matches/true positives. It neglects the importance of "negative space." I was happy to see M
by bitshiftfaced 4y ago
One of the weaknesses with Jaccard similarity is how it focuses on matches/true positives. It neglects the importance of "negative space."
I was happy to see Matthew's correlation coefficient (MCC) used in the recent "1st and Future - Player Contact Detection" Kaggle competition. MCC balances the eight confusion matrix ratios, and I've gotten excellent results when using it in the past.
- tpoacher 4y agoIt's not a weakness; it's a feature. One that makes it the better choice in situations where negative space should in fact be ignored. (comparing chest xrays are a typical example in medical imaging)
- bitshiftfaced 4y agoIt's unclear to me why they should be ignored.
- tpoacher 4y agoConsider two xrays and the output of an algorithm that tries to outline the lungs. Suppose that one xray was taken in a larger machine, and therefore has more negative space around the lungs. Also suppose that the algorithm delineated the lungs equally well in both cases (anatomically speaking). If you assess performance using the jaccard index, the metric is equal in both cases, as it should be, indicating equal performance of the algorithm w.r.t. the ground truth. Whereas anything that takes accuracy of true negatives into account will necessarily give a higher performance in the xray from the larger machine, even if the person xrayed and the lung outline were identical.
- bitshiftfaced 4y agoI don't see how this is an improvement over MCC in this example, since in this case, unless I'm mistaken, MCC would hypothetically give the same (perfect) value to both x-rays, just as Jaccard would.
- tpoacher 4y agoNo, in this scenario, the MCC will generally give a misleadingly higher value to the larger machine, since it takes into account the higher number of true negatives. In this scenario, this makes it a bad metric. Obviously, a 100% perfect segmentation would of course register as perfect in both, but in practice one rarely deals with such perfect predictions. In general, there is no metric that is universally "better" in all scenarios. One is expected to choose the metric that best suits the particular goal one wishes to validate against.