7 ms·
Bias detectives: the researchers striving to make algorithms fair
- paulus_magnus2 8y agoThis will be interesting. Most actions we people and ALL actions corporations take are optimised for maximal self / personal gain and not for justice (which is hard or impossible to define). This is the basis of neoliberalism. It will be interesting to see where pressure points will emerge and who / how negotiations will progress.
- Sol- 8y agoWhen I was reading through some algorithmic fairness literature some time ago, I came back a bit frustrated because as the article mentions, the fairness definitions are mutually incompatible (though some seem more plausible than others) and it's not really a problem that can be fully solved on a technical level. The only flicker of hope was that a perfect classifier can, by some definitions, be considered fair, so at least you have something to work with - if your classifier discriminates by gender or other attributes, you should at least make it good enough to back up its bias by delivering perfect accuracy (at which point you can investigate why inherit differences between groups seem to exist). It's good that some Computer Science researchers are ready to work in such politicized fields though, it's definitely necessary. I find it admirable because I personally wouldn't enjoy those discussions.
- LoSboccacc 8y agoIf unconstrained learning emits biased result it was given biased samples. That or the bias is in the dataset itself, which could still be fixed by removing and randomizing traits but at which point your alghorithm is learning a representation if reality which usefulness depend on the realm if application, say, great for university admittance and not so great for medical insurance purposes.
- tomjen3 8y agoYou assume that the underlying data cannot be correct and not fair.
- LoSboccacc 8y agouh, no? > [then] it was given biased samples
- commandlinefan 8y agoWell, whether or not (and if so, to what extent) unconscious bias taints engineering or scientific research is an interesting, and potentially important question to answer, it seems to me that the only people who are addressing the question are people who are overtly, consciously biased themselves - to the point where they actively exclude anybody who doesn't share their own set of biases.
- thisisit 8y agoIsn't the whole point of machine algorithms to find the best spot between variance and bias? In which case, every algorithms will have some bias. IMO, the focus should instead be on not overselling algorithms as being infallible rather something which will have some bias and needs overriding time to time. If a system is fully automated without checks and balances we might have serious problems. A good non-ML example was discussed couple of days ago on HN where a person was terminated by a machine without much oversight: https://idiallo.com/blog/when-a-machine-fired-me https://idiallo.com/blog/when-a-machine-fired-me
- pliny 8y ago>Isn't the whole point of machine algorithms to find the best spot between variance and bias? In which case, every algorithms will have some bias. This references two different, almost opposite, meanings of the word bias. The meaning in TFA is being individually subject to judgements about groups (eg an individual AA prisoner being denied parole because AAs as a population have higher rates of recidivism) whereas in the context of ML it refers to a model ignoring important features of a dataset.
- andrewlee224 8y agoAren't the algorithms already reasonably fair? The researchers are just trying to get them to be politically correct?
- Sol- 8y agoBefore throwing around accusations of political correctness you should consider that questions like that are not new, they have just received more attention now that algorithms decide much more things in life than in the past. For instance, US law has "disparate impact" provisions at least since the civil rights act (a time which was probably not dominated by political correctness), which requires outcomes to be not too different between races or other groups. (Though disparate impact is not a particularly good metric of fairness and doesn't seem to be used so much in algorithmic fairness nowadays.)
- TangoTrotFox 8y agoDisparate impact applies to actions which are unjustified. For instance imagine I decide to never hire somebody who likes rap music. That would have a disparate impact on a certain group of people, yet since it probably has nothing to do with the job I'm hiring for, it would be unjustified and could be argued to be a form of disparate impact discrimination. By contrast imagine for a labor job I decide to never hire anybody who can't left and carry at least 120 pounds. That would also have a disparate effect negatively impacting a protected class, but it would be justified. Machine learning takes all data and draws conclusions that map strongly against the data and ideally generalize to new scenarios. So long as the behavior trying to be predicted for was relevant to the task at hand, using the recommendations of such algorithms would certainly be justified and thus not fall under disparate impact. In a nutshell disparate impact is not to ensure equality of result, but to prevent discrimination by proxy. And discrimination not being of the form 'I'm not going to hire women because they can't lift as much as men' but of the form 'I'm not going to hire women because I don't like working with women.'
- marcoperaza 8y agoThe disparate impact doctrine is quite controversial, especially the ever more aggressive applications of it. It has led to some unfortunate rules. For example, it’s (generally and presumptively) illegal to hire based on intelligence tests, but seems to be okay (in practice) to hire only from elite universities that select students largely on the basis of SAT scores, which correlate very strongly with IQ.
- Proven 8y agoThis sounds silly. Bias is essential to individual decision making. Algos are there to help me automate my biased decision making, not to mislead me by making false assumptions on my behalf
- aldanor 8y agoNote that there's not only a potential selection bias problem, but a feedback issue as well. For instance, if the algorithm is biased toward assigning higher criminal activity risk to black people, the black people will be more likely to be checked and, as a consequence, the future versions of such algorithms will be even more biased in the same direction. Debiasing in such situations is a very tough endeavour.
- marcoperaza 8y agoA lot hinges on how you define fairness. Is it unfair if a higher percentage of <racial group> is flagged as “criminal activity risks”, to use your example, than of <other racial group>? What if it aligns with reality? No matter how accurate your metric is—nay, because of its accuracy—people will be mad at you because it reflects an underlying reality they are in denial of or believe is itself unfair, and want to rectify by requiring fictions elsewhere. I don’t envy people working on this because you by definition can’t win. Some of the powerful political forces, on the one hand, and the demands for accuracy and effectiveness, on the other, are irreconcilable.
- yorwba 8y ago> Is it unfair if a higher percentage of <racial group> is flagged as “criminal activity risks”, to use your example, than of <other racial group>? What if it aligns with reality? The article is about the unfairness of a higher percentage of <racial group> getting incorrectly flagged as “criminal activity risks”. The fairness of the expected output is essentially out of scope for this kind of research, it's all about the distribution of correct vs. erroneous decisions. The underlying reality is an imbalance in data quality or similar, not whatever you were thinking about.
- jakelazaroff 8y agoThat's why the grandparent mentioned feedback. Even if a disparity "aligns with reality", if the algorithm is used in a way that reinforces the disparity, then it's discriminatory. When designing algorithms as parts of systems, we need to be careful to not ossify statistics we're ostensibly just reflecting.
- kgwgk 8y agoIn summary: “You can’t have it all. If you want to be fair in one way, you might necessarily be unfair in another definition that also sounds reasonable.”
- deleted 8y ago[deleted]
- mlthoughts2018 8y agoIt reminds me of Arrow’s Impossibility Theorem [0]. David Deutsch had an anecdote in his book The Beginning of Infinity about explaining Arrow’s theorem to a US congressman and eventually getting him to the point of understanding how it applies to the electoral college in principle and there is no simple legislative change that could mollify it (I think the discussion was about how preferential voting would improve upon FPTP voting). The congressman replied something like “this is lamentable” and Deutsch wrote a big passage about whether it makes sense that anyone should ever find a mathematical fact to be “lamentable.” I do think there will be pragmatic varieties of this sort of impossibility theorem for machine learning fairness and that society in a broad sense will have a hard time processing it, and the possible legislative reactions might be totally unreasonable, even unintentionally harmful. [0]: < https://en.m.wikipedia.org/wiki/Arrow's_impossibility_theorem https://en.m.wikipedia.org/wiki/Arrow's_impossibility_theore... >
- kyleperik 8y agoThe purpose of Machine Learning is to generalize on a large scale. I like to think of it as the equivalent of someone who has years of experience in a particular area. Someone who has been doing something for years has seen so much, that they can know in an instant what a situation is based on clues and generalizations. It wouldn't be fast if it wasn't generalizing. If you want to claim you know what fair is in any given situation, then go and hardcode all your own fair rules, because you aren't going to find "fairness" in machine learning.
- deleted 8y ago[deleted]
- throwawayjava 8y ago> If you want to claim you know what fair is in any given situation, then go and hardcode all your own fair rules This strikes me as eerily similar to the argument that type systems are impossible because of the halting problem. It's sort of true in some sense, but not in an even remotely useful way. So it mostly functions as a way of derailing the conversation away of the more subtle distinctions that do matter (e.g., could we design an easy-to-use type system that rules out this particular type of non-termination/other class of bugs). There's a large middle-ground between "hard code all your own rules" and "completely unconstrained learning". Learning under constraints is not a new idea. A classical programming analogy to your argument might be "well the halting problem is undecidable so ignore all this high-level language stuff and just go code up your own turing machine; it's the best you'll ever be able to do". > because you aren't going to find "fairness" in machine learning. Why not? The human notion of fairness is fuzzy, which is why researchers have provided various formal notions of fairness in machine learning tasks. Obviously, these formal definitions may or may not correspond to your own gut instinct about what is "fair". And there might be friction between different notions of fairness. None of that should be surprising; otherwise, fairness wouldn't be something that philosophers continue to bleed ink about. But it is equally obvious that for some notions of fairness, there will exist machine learning algorithms that learn well under the given constraint.
- kyleperik 8y ago
- sarcasm_heals 8y agoThese aren't scientists. These are the corrupted who are making society more dangerous to live in. They subordinate science to favor special groups at the expense of everyone else. I'm so narrow minded thinking about using empirical decision making, though. I'm not seeing the big picture of what releasing people predisposed to violence could do for society.
- local_yokel 8y agoIt's worth pointing out that the original ProPublica investigation was conducted by journalists unskilled in statistics and machine learning. There was a convincing rebuttal posted by the actual scientists involved, which is of course ignored since "racist AI" is the kind of headline that's just too golden to abandon. http://www.uscourts.gov/sites/default/files/80_2_6_0.pdf http://www.uscourts.gov/sites/default/files/80_2_6_0.pdf
- conanbatt 8y ago> racist AI" is the kind of headline that's just too golden to abandon. TBF there is at least one leading ML scientist that has made a huge narrative on discriminating AI, particularly against women.
- Bartweiss 8y agoProPublica's work on algorithmic bias has all seemed well below their usual standards. I haven't followed this rebuttal, but their work showing racism in car insurance pricing was heavily criticized, and while the authors defended the work it looked to me like they picked only the weakest criticisms to respond to. (In the insurance case, ProPublica attempted to compare areas with comparable crash frequencies and show that rates were higher in poor and minority areas. But the data they had was moving accidents, especially with injuries, and the data they didn't have was stuff like "rate of car break-ins" and "odds of being hit by an uninsured driver". Which you would obviously expect to vary by region even when serious-injury accidents don't.)