4 ms·
> If the output examples we are training on are true, the ML algorithm won't adopt any incorrect biases. This is rarely the case when working wtih real data, a
by leftpad 10y ago
> If the output examples we are training on are true, the ML algorithm won't adopt any incorrect biases.
This is rarely the case when working wtih real data, and thus inspecting whether our models are biased against protected classes is probably one of the most important things an ML practitioner should do.
- wyager 10y agoWe have a good understanding of the error profiles of most ML algorithms. We can put tight bounds on the difference between predictions and validation data. If ML algorithms make mistakes, it's usually due to noise or low-quality data, not the algorithm itself. There is no reason an ML algorithm would be biased against a "protected class". It doesn't know what those are. It's possible that the algorithm will uncover a truth that you don't like, like that there as differences in risk profiles across race or gender, but that doesn't mean the algorithm is incorrectly biased. It just means that reality is at odds with how you might want it to be.
- leftpad 10y agoNot talking about the algorithm. It's the data that are biased. The choice of algorithm just determines how interpretable those biases actually are.
- wyager 10y ago> It's the data that are biased. Which data are you referring to? In most cases, the training data isn't human-generated, and if it is, we usually want to match human behavior as close as possible.
- leftpad 10y agoVirtually all data used to predict crimes or recidivism is fraught with human bias, for example. Not sure that we want to reproduce the bias of criminal justice system in any prediction problem involving this type of data. Read anything written by Solon Barocas: http://solon.barocas.org/ http://solon.barocas.org/
- wyager 10y agoHow is recidivism data biased? I'm sure that the information gleaned from parole officers and cops might be biased, but as long as the ML system is trained on whether or not someone actually reverted to committing crimes, it should be able to detect bias on the part of P.O.s and other functionaries and give a more accurate determination as to someone's chances of recidivism.