3 ms·
> This can completely skew the results. It can skew the results, but it can't completely invalidate them. Remember here, we're not trying to robustly estimate
by shadowmint 11y ago
> This can completely skew the results.
It can skew the results, but it can't completely invalidate them.
Remember here, we're not trying to robustly estimate populations of engineers with specific wages. We're looking at a data set you would expect to be more or less without variation, and being surprised when there is 1) variation, and 2) that variation appears to have racial/gender/whatever correlations.
Now, I get what you're arguing; you're saying, the sampling is biased, so any of those seeming correlations may well be biased (eg. specific demographic consistently under reports their income, people who do report their income tend to be the 'lower' bracket of incomes, etc.)... well, fair enough.
Caveat any results you come up with; but it's not like the data is going to be completely useless and meaningless because it's noisy.
This isn't an arbitrary academic exercise; it's tool for people to use to evaluate their own job positions.
What's the alternative? Have no idea at all what other people are earning? If you don't have any data, you can't do anything.
Even if the data you have is noisy, it'll give you a lot more insight than nothing.
Sure, I don't endorse getting righteous and taking it up the ladder ('My <insert group here> is discriminated against!') without doing your due diligence about samples and caveats.
...but taking a spreadsheet like this to your next pay review? Your manager better have some good answers to give out if you find yourself on the bottom of the curve.
What's wrong with that?
- DangerousPie 11y agoThere's a difference between noisy and biased. As long as the data is only noisy (that is, it has some random variations) I totally agree with you that it's fine to use, and that the results should be robust. However, if there is some sort of bias that only applies to a particular subset of the samples, all bets are off. Just as a totally imaginary example, what if men with higher incomes are more likely to share their salary information than men with lower incomes, while at the same time the situation is reversed for women. So now you will end up with more reports of high income from men and more reports of low income from women, even if their pay distributions are exactly the same. I am of course not saying that this is happening here, but these kinds of things would indeed completely invalidate the results.