4 ms·
In the most general sense of the word bias, there can be biases that arise purely as a result of model architecture. For example, I was messing around with GANs
by c1ccccc1 5y ago
In the most general sense of the word bias, there can be biases that arise purely as a result of model architecture. For example, I was messing around with GANs about a year ago, and kept on getting checkerboard-like patterns in the generated images. This was a result of how I had set up the convolution layers: It was a bias introduced by model architecture.
In terms of the "unfairness based on what group you belong to" kind of bias, that's almost always a result of bad input; you'd have to work really hard to make an inherently racist architecture. So the problem is introduced in the training data. But it could be solved either by getting better training data, or by making clever improvements to your algorithm to work around the limitations of the training data. Often changing the algorithm is easier than improving the data, which I think is the point this article is making.
As an example of such a change, let's say we're designing an algorithm to estimate credit score, and we're worried about unfairness based on group membership. Let M be a variable representing what groups a person belongs to: race sex, etc, and let S be the output of the algorithm: that person's credit score. An easy thing to do would be to not allow M as an input to our algorithm. That way we can say that the output "does not depend on group membership". The problem is that we may still have proxies for M as inputs to the algorithm, and if the dataset is biased, that algorithm will still learn to use those proxies to discriminate based on group membership. A better method is surprisingly to keep M as an input during training of the model. Then, at inference time, average over M to get the final output of the model. That still erases the information about what group a particular person belongs to, but it also has the advantage that the model isn't incentivized to learn proxies for M during training, since it has access to M directly.
- malshe 5y agoThanks for taking time to write this explanation. Note that the bias you mentioned about your experience with GANs is not what I would call systematic bias. It is random noise. And it may even be racist in nature. Yes, you may keep on being biased in your own way day in and day out, which makes it systematic for you, but we won't need an entire field of ethical AI to study that. The example you gave about credit score, and the theme generally, is very well understood for a very long time. That's something we teach to stats undergrads early on. That's just one source of bias in the data, btw. There are numerous other data biases and tons of research exists on that. Even on one method of data collection, such as survey data, there is a huge literature on reducing the bias. > you'd have to work really hard to make an inherently racist architecture. That's precisely where my confusion comes from. If someone makes this racist architecture, it is not a systematic bias that anyone cares about. It's a rogue model. In fact, in many cases, this could be outright illegal. For example, if someone designs a credit scoring model to deliberately leave out one group of people based on their gender, ethnicity, or religion, that's not going to fly with regulators. So anyway, I still don't have an example of a systematic bias that is not caused by data.