4 ms·
The accuracy, fairness, and limits of predicting recidivism
- thomasz 9y ago> Algorithms for predicting recidivism are commonly used to assess a criminal defendant’s likelihood of committing a crime. These predictions are used in pretrial, parole, and sentencing decisions. Proponents of these systems argue that big data and advanced machine learning make these analyses more accurate and less biased than humans. We show, however, that the widely used commercial risk assessment software COMPAS is no more accurate or fair than predictions made by people with little or no criminal justice expertise. We further show that a simple linear predictor provided with only two features is nearly equivalent to COMPAS with its 137 features
- fvdessen 9y agoIt would have been interesting to know which are those two features that predict recidivism so well.
- wu-ikkyu 9y ago>Proponents of these systems argue that big data and advanced machine learning make these analyses more accurate and less biased than humans. This assumes the data sets of the criminal punishment system are not inherently biased, which is of course a false assumption. All this does is make the runaway train of the US criminal punishment system a more efficient machine.
- vec 9y ago> This assumes the data sets of the criminal punishment system are not inherently biased, which is of course a false assumption. It doesn't even require that assumption to go awry. Assume there exists a stereotype, say "people with freckles are more likely to be criminals". It doesn't actually matter if the stereotype has any basis in reality, just that it is widely believed. People will, on the margin, be less likely to hire freckled people. This reduces the legitimate employment opportunities available to a freckled individual, which tends to make the illegitimate opportunities more attractive by comparison. So the stereotype becomes a self-fulfilling prophecy: by assuming that freckled people are more likely to commit crimes, society actually causes freckled people to be more likely to commit crimes. An expert system will notice this and begin using "has freckles" as a weighting factor in predicting recidivism. It's important to note that the expert system is not wrong. Freckles are in fact, at this point, statistically correlated to recidivism. But the expert system can't know why the correlation exists. All it can do is tighten the vicious feedback loop, noticing a statistical correlation that strengthens the existing stereotype, which exacerbates the real world impacts, which increases the statistical correlation, which strengthens the stereotype, and so on.
- gizmo686 9y agoNot the main point of the paper, but >it is argued that the COMPAS score is not biased against blacks because the likelihood of recidivism among high-risk offenders is the same regardless of race (predictive parity), it can discriminate between recidivists and nonrecidivists equally well for white and black defendants as measured with the area under the curve of the receiver operating characteristic, AUC-ROC (accuracy equity), and the likelihood of recidivism for any given score is the same regardless of race (calibration) Simpon's paradox [0], probably one of the most insidious problems of well intentioned statistics. [0] https://en.wikipedia.org/wiki/Simpson%27s_paradox https://en.wikipedia.org/wiki/Simpson%27s_paradox
- fvdessen 9y agoInterestingly, in the paper the human assessments shows the exact same biases.
- harryh 9y agoHow is this Simpson's Paradox?
- gizmo686 9y agoThere is a corralation across the entire population that does not exist in 'bucket' of the population. More importantly, the bucketing involved is reasonable and has explanatory power. (If this were not the case, you might have picked the buckets specifically to get this result. Still an example of Simpsons Paradox, but not interesting.)
- gadders 9y agoOn a semi-related note, Theodore Dalrymple (ex-prison psychiatrist) argued that the practise of parole should be ended for similar reasons: https://www.spectator.co.uk/2018/01/parole-is-unfair-and-unworkable-lets-abolish-it/ https://www.spectator.co.uk/2018/01/parole-is-unfair-and-unw...
- mrow84 9y agoThis is an interesting paper on a subject that has provoked debate on HN in the past. I would be interested to read criticism of the methods from those who have expressed support for these kinds of automated systems. The most obvious criticism is that the uninspiring performance of the COMPAS model does not provide any evidence of the weakness of automated systems in general, which is surely correct. This kind of argument would not, however, address more general concerns about how and why these kinds of systems are adopted - why is this system being used when its performance is not obviously competitive? (The answer to that is presumably "corruption", in a generalised sense, but that is hardly an adequate response.) Even more broadly, as the complexity of human society increases, it seems likely that the difficulty of making constructive interventions will also increase, and possibly more rapidly ("The Collapse of Complex Societies" by Joseph Tainter investigates this hypothesis). It does seem risky to just plough on, hoping that we can invent our way out of any problems that might arise as a result of what we are doing.
- sinxoveretothex 9y agoOne question that would be interesting to explore is: what is the predictive score of judges/juries/etc? Is it more or less variable than the COMPAS score? In other words, do judges perform better than COMPAS? And slightly less important: is the variance lower? If it isn't, then there is no point arguing whether COMPAS is risky: it would be less so than the alternative. Although I suppose that the 'crowd of non-experts' is essentially what a jury is and they had about the same performance. At any rate, it sounds like much less of a hassle to input characteristics into COMPAS and get a probability than getting a jury together. Besides, I think I'd be much more worried that a jury could influence each other and bias its decision than about am algorithm making unfair predictions.
- stult 9y agoWell hold on. COMPAS performs as well as people but costs less than getting a jury together or paying a professional. So I wouldn't say that's uninspiring. It also performs as well as their simpler model. But they only evaluated the simple model on data from a single county. There's not enough evidence to determine whether that model would perform as well on other data sets. I would presume, perhaps wrongly, that more work has been done to ensure the generalizability of COMPAS, the results of which work militated against the adoption of a simpler algorithm. Possibly because the algorithm has to avoid racial disparities that might creep up in counties with a different demographic mix. Even if the simpler model performs equally well across many jurisdictions, that merely suggests that a simpler non-proprietary algorithm would be better. But we're still choosing between algorithms, not discrediting the idea of algorithmic sentencing altogether.