3 ms·
> If the input data such as the world in general is causing "problematic" output data, it might be time to reconsider what you think of as problematic. The dat
by zuminator 3y ago
> If the input data such as the world in general is causing "problematic" output data, it might be time to reconsider what you think of as problematic.
The data in the article is arguing against that kind of pat assumption. For example it says that: "Women made up a tiny fraction of the images generated for the keyword 'judge' — about 3% — when in reality 34% of US judges are women, according to the National Association of Women Judges and the Federal Judicial Center."
So part of the concern is that the input data is not a representative sample of "the world in general," but relies on stock photos, celebrity media images, sensationalist news reports, items which by design do not accurately reflect the real world.
- indymike 3y ago> when in reality 34% of US judges are women, according to the National Association of Women Judges and the Federal Judicial Center What percentages of judges are women is a different question than what percentage of judges depicted in visual media are women... and that is an even different question than what percentage of judge images used to train stable diffusion were women?
- ceejayoz 3y agoYes, that's the point. The dataset's proportions don't match reality's proportions, which leaves the AI with an inaccurate perception of reality. This presents as a phenomeonon one might describe as "bias", even though the AI itself has no motivations.
- indymike 3y ago> The dataset's proportions don't match reality's proportions, which leaves the AI with an inaccurate perception of reality. People have this problem as well.