4 ms·
I completely disagree. You're right that everyone has biases - including the authors of the AI who will therefore choose a biased training dataset. When the to
by diputsmonro 4y ago
I completely disagree. You're right that everyone has biases - including the authors of the AI who will therefore choose a biased training dataset. When the tool they are constructing has unparalleled power, and is very likely to negatively impact the lives of those who the authors are biased against (consciously or not), it's entirely fair for those impacted groups to be concerned.
Every dataset has a skew to it, which should be accounted for during training. If the authors don't explicitly account for such things, out of ignorance or bias, then that skew, that bias, will be ingrained in the AI.
For example, if I only train my AI on it's knowledge of chocolate from ads for a particular brand, it's going to have some very opinionated, very wrong ideas about chocolate. This example is silly and obvious, but similar skews happen all the time in real datasets that the authors don't have the time or expertise to recognize.
When people who do have that expertise speak up, we should listen to them and fix it, not just blindly trust "the data" and whatever our fancy algorithm does with it. Garbage in, garbage out.
- visarga 4y agoThese imbalances in the training data are there because the world generates data in an imbalanced way. Sometimes you can correct for it, other times it's impossible. You simply can't find good data for all your cases. For example in medical diagnosing or in self driving.