3 ms·
The biggest problem is that the dataset contains an artificial feature that is not invertible. This is an issue because the biases of the author of the dataset
by guipsp 4y ago
The biggest problem is that the dataset contains an artificial feature that is not invertible. This is an issue because the biases of the author of the dataset are present in that feature, and you will never be able to "train" your way out of it because it is not invertible.
- 1attice 4y agoGreat explanation but consider adding blurbs about what makes a feature artificial, what 'invertibility' is, why it's important, and how you could (as you put it) train one's way out of the bias so long as the feature is invertible. Finally, bring it home by giving an example (fictious or, even better, real-historical!) of a model that would be biased, but for the blessings of feature-invertibility and further training, and then explain that you since you couldn't do that with this dataset, because the artificial feature is not invertible. That I think gets us a full ELI5, though I agree the common-sense cutoff is subjective. (And I say this in a spirit of co-collaboration -- I'm fascinated by the problem of ELI5 something like this, as I wind up having these conversations ad-hoc in-the-wild with family and friends, as I work in the field. Finding simple language just is progress)
- 1attice 4y agoOh! One more critical point for an ELI5! 'Although models that make use of non-invertible, artifical features may still be biased, it's less likely. Using only invertible features eliminates some risks of bias, but not others.' At this point I think one might have to explain the Naturalistic Fallacy (which accounts for much of the remaining possibility of bias in a model) but it starts to get into tit-for-tat hand-to-hand ontology: what 'bias' means, what different kinds of bias are possible, and how even 'unbiased data' can create a model that demonstrates behaviours that colloquially and idiomatically count as biased. But one must cut the cloth of the universe somewhere
- a2800276 4y agoSorry, about the rant: ELI5 really triggers me. Have you ever talked to a five-year-old? They're very interested but ignorant and have minute attention spans. Have some self-respect, put on your big boy pants and at least make an effort to understand the adult version and ask some specifics about what you don't understand. What you're asking for is a full on Malcom Gladwell style essay. Not only that, but all the points are discussed in the linked blog post. If one can't even be asked to put in the effort to read it, isn't it a bit presumptuous to expect anyone to summarize for you in elaborate detail!? ELI5 would be: The authors wanted to find out if housing prices have something to do with bad air. They thought some people unfairly make homes cheaper if black people live in them. And that in some cases more expensive. Instead of including data about how many black people live in a neighborhood he included a value that supports his opinion and can't be checked. Because scientists aren't supposed to base their conclusions on gut feelings, this makes the sample data an example for problems that arise when working with data. Five year olds are not particularly intellectual and wouldn't enjoy the discussion of biased models with feature-invertible data!
- josephcsible 4y agoBeing non-invertible just means you can't get the original racial breakdown from the variable. It doesn't mean that you're stuck with any particular bias forever.