5 ms·
Thanks for the further background information. I have to say it doesn't really make it better for me. The "angry people" are of course correct that you can also
by horrified 5y ago
Thanks for the further background information. I have to say it doesn't really make it better for me. The "angry people" are of course correct that you can also create bias in other ways than data sets. But are they implying that people generally deliberately introduce such biases to uphold discrimination? That seems like a very serious and offensive claim to make, and not very helpful either.
The whole way to think about issues is backwards in my opinion. I would think usually when you train some algorithm, you tune and experiment until it roughly does what it wants you to do. I don't think anybody starts out by saying "let's use the L2 loss function so that everybody starts white". They'll start with some loss function, and if the results are not as good as they hope, they'll try another one. In fact the usual approach will lead back to issues with the data set, because that is what people will test and tweak their algorithms with. If the dataset doesn't contain "problematic" cases, they won't be detected.
But overall, such misclassifications are simply "bugs" that should get a ticket and be fixed, not trigger huge debates. I think it is toxic to try to frame everything as an issue of race.
- joshuamorton 5y ago> Thanks for the further background information. I have to say it doesn't really make it better for me. The "angry people" are of course correct that you can also create bias in other ways than data sets. But are they implying that people generally deliberately introduce such biases to uphold discrimination? That seems like a very serious and offensive claim to make, and not very helpful either. No. I think Isbell's Neurips Keynote (https://nips.cc/virtual/2020/public/invited_16166.html https://nips.cc/virtual/2020/public/invited_16166.html), titled "You Can’t Escape Hyperparameters and Latent Variables" does a good job of explaining this. The humans who ultimately validate the model (and who decide on the dataset) are a hyperparameter. Often ignored, yes, but they are still part of the training loop. They decide what the other hyperparams are, when to stop training and publish, etc. To use a question I've asked on HN before: say you're training a model to detect criminality based on facial structure. This has come up as a real world example, papers have been published on this topic. What does a "good" dataset look like? Or similarly, for a system that decides on bail or sentence length. Do you use historical data on bail or sentencing? We have very well documented examples of bias in both of those things, even in the ground truth. So how do you decide to mitigate that bias? Or do you choose not to, and to continue enforcing said biases in your model? > But overall, such misclassifications are simply "bugs" that should get a ticket and be fixed, not trigger huge debates But when such "bugs" aren't prioritized because people don't think they are bugs, you have to debate whether or not they are bugs at all! The hyperparameter here is "who decides what is or isn't a bug"
- andreyk 5y agoIsbell's Neurips keynote is fantastic! Definitely recommend.
- horrified 5y ago"say you're training a model to detect criminality based on facial structure. This has come up as a real world example, papers have been published on this topic. What does a "good" dataset look like?" I don't think anybody who is respected says "here is this data set of criminals, we have trained the algorithm on it, and therefore it is proven that such and such facial features predict criminality". I mean yeah this mistake has been made over and over again (even before the invention of computers), but it has long been debunked. Also presumably "black skin" is a good predictor for criminality - in the current day, the crime rate is higher for black people. The algorithm only detects that, it doesn't interpret it. It is up to the humans who use the algorithm to interpret it. If you interpret it as "black people have a genetic disposition to criminality", you are wrong. But it wouldn't be the fault of the algorithm. What is insanity, but basically what the "AI ethics" people demand, is to tweak the algorithms to make them pretend the prevalence of criminality is not higher in certain populations. "But when such "bugs" aren't prioritized because people don't think they are bugs, you have to debate whether or not they are bugs at all!" Nobody says they are not bugs. You are creating an imaginary problem here. You really think, say, researchers at Amazon said "let's make it so that women are ranked down by the algorithm"? Likewise I don't think anybody says "the algorithm should rank black people worse for crime". It is also not a novel idea to look out for bias int he algorithms, delivered to us by the woke crowd. The whole field is about treating bias - a machine learning algorithm is all about training some bias.
- joshuamorton 5y ago> Nobody says they are not bugs. You are creating an imaginary problem here. You really think, say, researchers at Amazon said "let's make it so that women are ranked down by the algorithm"? Likewise I don't think anybody says "the algorithm should rank black people worse for crime". They did though, at least until Gebru and those like her came along and forced the issue. It's really sad to see people say that this was never a concern as though bias and ethics were taken seriously by the field as a whole more than, say, 5 years ago. They weren't. Idk if you're new to the field or weren't paying attention, but it just wasn't a thing. Like most of the foundational papers in terms of racial misclassification and such are from 2017 and 2018.[2] It's more recent than...GANs or AlphaZero. Not to mention that there's attempts to publish garbage like this[1] every year! > You really think, say, researchers at Amazon said "let's make it so that women are ranked down by the algorithm"? Likewise I don't think anybody says "the algorithm should rank black people worse for crime". No, I already said this. Someone failing to notice a bug isn't malice. But there issue is that no one even considered that these kinds of things were bugs so they didn't get noticed or researched. > The whole field is about treating bias - a machine learning algorithm is all about training some bias. Yes, but thinking about race as a particular category where we should avoid unintended bias (and indeed prefer generalization across categories) was a novel idea when proposed by those ethicists! > But it wouldn't be the fault of the algorithm. What is insanity, but basically what the "AI ethics" people demand, is to tweak the algorithms to make them pretend the prevalence of criminality is not higher in certain populations. But...you're making the algorithm. If your goal is to build a model that tries to detect "racial criminality", I'm going to suggest that you probably are doing something racist, because there isn't really a useful, non-racist, reason to train a model that incorrectly classifies people as criminal based on their skin color. On the other hand, if you're having to do additional interpretation of the model output, why aren't you integrating that additional interpretation into the model? And if you can't, then is the model even adding any value? Probably not. And that's not even ignoring questions like what "prevalence of criminality". I think you mean "are arrested more often". We often think that that correlates with criminality, and for some crimes it may, but for e.g. drug crimes we know that it doesn't. The point is, if you don't have at least thoughtful answers to all of those questions and more, you have no business trying to do "criminality" prediction, because your algorithm is not doing whatever you think its doing. [1]: https://www.bbc.com/news/technology-53165286 https://www.bbc.com/news/technology-53165286 [2]: Seriously, Gender Shades is 2018, Debiasing word embeddings is 2016 which I think is the earliest you could argue people were taking this stuff seriously, and it cites "Unequal Representation and Gender Stereotypes in Image Search Results for Occupations" from 2015, which is kind of it.