3 ms·
I think you're missing the first point in both my summary and in the article: It does not need to be the case that "a certain race commits more crimes". It can,
by 21echoes 11y ago
I think you're missing the first point in both my summary and in the article: It does not need to be the case that "a certain race commits more crimes". It can, instead, just be the case that a certain race is arrested for committing more crimes, despite the equal rates across races of the actual criminal behavior.
For instance: it's a well recognized fact[1] that blacks and whites use and deal marijuana at the same rate, but blacks are arrested for it in far larger volume. So, if this data and other similar data sets are the seed in a machine learning algorithm, then algorithms like the Heat List will output racially biased data.
[1] https://www.aclu.org/files/assets/aclu-thewaronmarijuana-rel2.pdf https://www.aclu.org/files/assets/aclu-thewaronmarijuana-rel...
- yummyfajitas 11y agoI didn't miss that. As I said: Now there are statistical issues one might run into - e.g., early overfitting of what is essentially a bandit algorithm, and unaccounted for feedback between training data and system outputs. There is nothing fundamental about machine learning that says seed data like this will give biased outputs - many algorithms do have this problem (it's a difficult one to deal with), but it's not fundamental. I certainly didn't get the impression from the article that it was advocating for algorithms which are less sensitive to these errors. Among other things, that's far less of a conversation that "we have to talk about", but far more of a conversation that some stats geeks have to talk about. These are also far less of an "ethical" problem (as the article asserts) and far more of a technical one. But maybe I misread.