6 ms·
At first I thought "What's the big deal? You also remove the query from recommender systems, for example. It's an obvious uninteresting result!". But then I re
by JD557 7y ago
At first I thought "What's the big deal? You also remove the query from recommender systems, for example. It's an obvious uninteresting result!".
But then I read the linked article [Nissim 2019] and it all became much clearer with te example in the title: "Man is to Doctor as Woman is to Doctor". If you remove "Doctor" from the results, you'll get "Nurse" instead, not because the dataset/society/... is biased, but because you inserted the bias in your model!
With this small "optimization" (and others presented in the article, such as the "threshold-method" and hand-picked results from the Top-N words), it's trivial to use the anology method to show that any dataset is biased, since you are filtering the unbiased results!
One a side note, I wish articles like this were more popular. I find that it's really easy to use AI techniques "the wrong way" and there's are a lack of articles pointing to common pitfalls (and, to make things worse, there are a lot of blog posts doing things wrong, which only validates bad methodology).
[Nissim 2019]: https://arxiv.org/pdf/1905.09866.pdf https://arxiv.org/pdf/1905.09866.pdf
- sooheon 7y agoWow, the Manzini et al. paper is particularly egregious, because they accidentally swapped "caucasian" and "black" in their actual query and still got the sensational word associations they wanted.
- dmurray 7y agoYeah, that's more or less academic fraud.
- dvfjsdhgfv 7y agoI wonder why they do that, it's more or less trivial to verify.
- anjc 7y agoThey're reporting real findings which they allow everyone to verify. There's nothing academically fraudulent about this.
- sooheon 7y agoYou're right they're just reporting their findings, and their stated methodology is to look through the top N words and pick the one they like best. They dug all the way down to the 39th word to find asian:engineer => black:? "killer", when the actual top two associations were "operator" and "jockey" (which also indicate bias, but not the kind the authors wanted). They dug to the 53rd word to find christian:conservative => muslim:? "regressive", when top results were "moderate" and "conservative. It may not be fraudulent, but the methodology is highly motivated, subjective, and introduces more bias than it purports to find.
- anjc 7y agoYeah I can't disagree with this. I think it's typical to present the highlights of your work though. In this case, the issue seems to be "why do the authors consider 'black:killer' to be a highlight"? It might be racist but it's not academically fraudulent (a strong statement) in my view.
- tgb 7y agoThe other reason to not exclude King is that it's pretty likely that Queen is the closest word to King. Then King - X + Y = Queen is just saying that X and Y are close to each other, not an interesting result.
- novaleaf 7y agoyour 2 sentence explanation is better than the entire article!
- SQueeeeeL 7y agoI disagree, the article was pretty decent. And it had some nice images.
- Tarq0n 7y agoConsidering the very high dimensionality of the space, it's not as obvious as you think. Consider the noise inherent in the word 2vec negative sampling method as well. Another word could very well end up closer to "king".
- derefr 7y agoFrom my perspective (linguistic anthropology), they’re not actually all that close. Most historical “queens” (which have that label applied to them by modern English speakers) were not rulers (and we have a separate term, “queen regnant”, for that) but rather the gender dual to the male “royal consort.” It was only in recent history where you see examples of “equal-opportunity” monarchies that could have either a male or female monarch of equal power, and thus usages of “queen” to denote those monarchs. Thus—given that we’re defining words based on their centroids of usage in a historical corpus—if a woman is a monarch of a kingdom, “king” is a tighter historical fit to describe her role than “queen” is.
- tareqak 7y agoOne example: https://en.wikipedia.org/wiki/Jadwiga_of_Poland#Coronation_(1384) https://en.wikipedia.org/wiki/Jadwiga_of_Poland#Coronation_(...
- stult 7y agoI think there are two ways to look at excluding input vectors from the output. You can look at it as cheating/bias, or you could say the method explicitly excludes a specific subcategory of otherwise valid results (i.e., perfectly equivalent analogs, where X:Z == Y:Z... not sure if there is a term of art for that). Which to me seems like a perfectly legitimate limitation on the algorithm, as long as it is explicitly acknowledged and handled. But I think you are absolutely correct that very few articles explicitly acknowledge and handle these types of common pitfalls. Especially more intro-level articles. Personally, I'd love to read more about effective strategies for ensembling different algorithms to work around some of those pitfalls. For example, the OP points out that Word2Vec underperforms on lexical semantics. I wonder if you could use a different algorithm that performs well on lexical semantics but poorly on other categories as a supplement, with some third algorithm deciding which to use in any given case.
- flo_hu 7y agoI would also agree with you that it is fine to add additional rules to improve the outcome, but than it shouldn't be made clear in the way the result is presented (as you say, that rarely happens in intro-level tutorials/courses). Your last point sounds like a cool idea! Using those more in-depth metrics to find weaknesses and see if other, complementary algorithms can fill the gap.
- SiempreViernes 7y ago*should?
- bjourne 7y agoThe dataset is biased. Otherwise woman~nurse wouldn't have been the second hit for the query. Had the dataset not been biased it would have been able to produce the analogy "woman is to doctor as man is to nurse". See table 4b where they perform that experiment. I think it is also important to note that this is an argument over which of two methodologies is best. The answer likely is that it depends on the use case. Researchers are not being accused of using underhanded methods, manipulating data or cheating.
- Diggsey 7y agoBias is not a boolean quantity. The underlying data set is somewhat biased, and the "trick" being used greatly exaggerates the impact of that bias.
- bjourne 7y agoBias is a boolean quality. A random distribution is not biased. It really is not a "trick".
- 0xFluegel 7y agoAny finite sampling of any distribution (biased or not) can show bias. A feasible process that doesn't introduce or exagerate bias is the biggest problem when trying to get representative statistics. And bias can certainly be of different magnitude. A weighted dice can have a bias that shows up after only 10 rolls and another one can only show it's bias after 1000.
- dragonwriter 7y ago> A random distribution is not biased. It really is not a "trick". A random sample is not a product of a biased sampling method, but it can be (and almost always is) biased, it's just if you take enough such samples, the biases will tend to offset, which is why we say the sampling method isn't biased.
- bjourne 7y ago
- 6gvONxR4sf7o 7y agoFor others reading, there are still important biases these models pick up (for a variety of reasons I won't go into). The takeaway here is that some of the evidence you've heard about isn't really evidence, not that the biases aren't there.
- thrwaway3873 7y agobut in analogies, a:b::c::___ (a is to b as c is to what), having d be the same as b is not allowed. Such as, I believe, on a Miller Analogies Test. In a mathematical context, if you have something like this: 2:0::4: (two is to zero as four is to what?) then acceptable answers might be 2 (rule is: subtract 2), or maybe it could be 1 (take half, subtract one). But it couldn't be zero. (rule can't be "multiply by 0"). I could be wrong but I believe that's how it works.
- dragonwriter 7y ago> but in analogies, a:b::c::___ (a is to b as c is to what), having d be the same as b is not allowed Particular tests may adopt such a rule, but this doesn't apply to analogies in general. Particularly, consider the case where Alice and Carol and both daughters of Bob: Alice:Bob::Carol:Bob is a perfectly valid analogy; a:b::c:b is commonly phrased along the lines of “a and c are similarly situated with respect to b”.
- Jenz 7y ago> I find that it's really easy to use AI techniques "the wrong way" NLP-wise, is there "a right way"?
- JD557 7y agoJust because there's not a "right way" doesn't meant that there aren't a lot "of wrong ways" ;) Actually, some years ago I bumped into a similar problem to the one discussed, where someone wanted to use NLP to show gender bias in a dataset, and hit one of the common pitfalls that I mention (I hope that I don't start a flamewar by sharing this story): Here's what they tried to do: 1. Fetch a list of ~800 atendee names from a Portuguese tech conference (it had an official API with user profiles) 2. Download a dataset most common male/female names for newborn babies in Portugal and America for the latest 3 years 3. Train a naive bayes model on the downloaded dataset and use it to classify the antedees into male/female After doing that, the algorithm returned something like "8 female attendees and 792 male antendees". I found this particularly strange (considering that I knew more than 8 women that attended on previous years), so I took a peek at the antendee dataset and found that: - There were some users using a fake name (including one organization account) - There were certainly more than 8 female antendees, and at least 6 were named "Inês" (female name) After discussing this with the ones involved, we found the problem! - The dataset was not being normalized (it was being trained with "Ines" and tested with "Inês") - The naive bayes[1] implementation used, when faced with a completly new input, outputed the most common class of the training dataset In the end, the final result was closer to "80 female, 705 male, 15 unknown", which is a much more believable result (closer to the typical distribution of Software Engineer students in Portugal). Note that the author wasn't trying to deceive anyone, he just tripped on some common pitfalls (forgot to normalize the data and used an off-the-shelf implementation without looking into the details). [1] There was only one attribute, so implementing this without using a naive bayes library was actually easier and produced the correct results.