4 ms·
For some reason another copy of this was flagged (https://news.ycombinator.com/item?id=37121000 https://news.ycombinator.com/item?id=37121000) and I wanted to r
by wackycat 3y ago
For some reason another copy of this was flagged (https://news.ycombinator.com/item?id=37121000 https://news.ycombinator.com/item?id=37121000) and I wanted to respond to a thread there
>> An entire hugely important potentially world changing field not having representation from a huge swath of the population certainly seems problematic.
>I assume lots of groups are underrepresented, depending on how you slice up humanity. Why is that a problem in and of itself? Ensuring training data is representative is a fundamental technique of data science that you can implement regardless of skin color.
A representative sample means that any way you "slice up humanity", if you slice your training data the same way the proportions of slices would stay the same. Anyone of any skin color can implement that but they must do so taking racial proportions of the general population into account.
>> And in terms of the second part; the problems with such widespread usage are many fold, with detrimental bias and lack of representation as one of the many folds.
>But what does that have to do with who is working on the AI systems?
This is the meatier question but in my opinion, it matters because folks who are underrepresented can often (read: not always) have a better eye for underrepresentation in data. Another way to think of it is the unknown unknowns framing and privilege can result in unknown unknowns around what representative data really looks like.