4 ms·
She didn't call him a racist. She pointed out that there's more to addressing bias than just acknowledging that it's there. If you are going to say "Well, garba
by compycom 6y ago
She didn't call him a racist. She pointed out that there's more to addressing bias than just acknowledging that it's there. If you are going to say "Well, garbage-in, garbage-out!", why do you keep putting the racist garbage in? Why do you, knowing of the problem, keep using biased sources of data, knowing of the potential harms?
That's the deeper question ML research/industry has to grapple with. When these models are deployed increasingly quickly and at scale in ways that can potentially cause massive harm, why is it okay to keep doing the same exact harmful things?
LeCun mentioned that there would have been opposite problem if it were trained on a dataset from Senegal. But why wasn't it trained on a dataset from Senegal? Why do we always see these errors where white-centric datasets produce white-centric results?
It's obvious that while it would be a symmetric situation a vacuum, we do not live in a vacuum. We live in a world with deep sociocultural biases in favor and against various racial groups. And it is unjust to let AI perpetuate and entrench these biases by acting at with this bias at scale.
Acknowledging dataset bias by itself doesn't address the bias meaningfully. In practice, we often treat the bias as an exogenous factor when it is not, moving it outside the scope of our responsibility. But it is very much the product of our work, a reflection of our choices, values, and beliefs about what to prioritize. We can't abdicate our responsibility for it (even if we choose not to prioritize it.)
- MperorM 6y agoI think this is a great comment, that would do a lot of people a lot of good to read. That said, how should a person like Yann LeCun argue his case? Namely that biased models are the result of bad datasets more so than the result of bad algorithms. I don't think Yann would disagree that biased datasets is a large systemic issue. How should a person like Yann make his point? It doesn't seem to me that Yann and Timnit disagree all that much. They both agree biased datasets is a problem. They both agree it's a systemic problem. They both agree it does tremendous harm when these biased models are deployed. I'm very confused as to why there even could be debate between these two people. I cannot find any meaningful difference in their views.
- benjohnson 6y agoMy opinion: Yang made the mistake of thinking a factional conversation was taking place.
- tzs 6y ago> LeCun mentioned that there would have been opposite problem if it were trained on a dataset from Senegal. But why wasn't it trained on a dataset from Senegal? Why do we always see these errors where white-centric datasets produce white-centric results? Probably because it was a research AI not a production AI. Having a very diverse dataset at that stage doesn’t help with your research, so it is fine to use whatever is easily available. At that stage, you are trying to show that your approach can work in some cases. Once you've got that, it is time to expand the research with a wider range of inputs to find out what the limits of your approach are. For example, if I were trying to make a US English speech to text transcription system, I might start with recordings of assorted NPR programs, because NRP often makes the recording available online along with they transcripts. That would be great for determining if my basic approach has promise. Once I have determined that, so know that the whole endeavor is not just a waste of time, I could go looking for data that includes speech that has characteristics that would be missing from the NPR data, such as heavy regional accents.
- kristjansson 6y ago> In practice, we often treat the bias as an exogenous factor when it is not, moving it outside the scope of our responsibility Thanks the comment and this explication, you clarified the conflict that incident for me. While I think the sibling comment is correct that LeCun and Gebru would agree on the proximate causes and mitigations of the adverse outcomes of that particular super-resolution model, the issue was (seemingly?) that focusing on the proximate causes can be read as absolving researchers of their responsibility for those outcomes, and avoids discussion of that responsibility in any instance. Which is a entirely fair criticism of the field as a whole, though it may have been a bit lost in translation to the particular avatars in that instance.
- kitsune_ 6y agoEveryone should read your comment. YLC's initial reply was shockingly reductionist and myopic. How can supposedly smart people, software engineers who are used to thinking in systems and dependencies be so stupid? There is a blindspot in the tech community that is simply terrifying.