11 ms·
The eye opening thing here is not that the AI failed, but why it failed. At start the AI is like a baby, it doesn't know anything or have any opinions. By teac
by fuscy 8y ago
The eye opening thing here is not that the AI failed, but why it failed.
At start the AI is like a baby, it doesn't know anything or have any opinions. By teaching it using a set of data, in this case a set of resumes and the outcome then it can form an opinion.
The AI becoming biased tells that the "teacher" was biased also. So actually Amazon's recruiting process seems to be a mess with the technical skills on the resume amounting to zilch, gender and the aggressiveness of the resume's language being the most important (because that's how the human recruiters actually hired people when someone put a resume).
The number of women and men in the data set shouldn't matter (algorithms learn that even if there was 1 woman, if she was hired then it will be positive about future woman candidates). What matters is the rejection rate which it learned from the data.. The hiring process is inherently biased against women.
Technically one could say that the AI was successful because it emulated the current Amazon hiring status.
- hkai 8y agoHow did you come to the conclusion that gender was being the most important, rather than skills or aggressiveness?
- kaitai 8y agoI don't think that's what the parent was claiming; the parent says "gender and aggressiveness" were most important and skills listed on the resume as providing such an unclear signal for actual hires that they were not picked up by the AI.
- lalaland1125 8y ago> The number of women and men in the data set shouldn't matter (algorithms learn that even if there was 1 woman, if she was hired then it will be positive about future woman candidates). This is incorrect. The key thing to keep in mind is that they are not just predicting who is a good candidate, they are also ranking by the certainty of their prediction. Lower numbers of female candidates could plausibly lead to lower certainty for the prediction model as it would have less data on those people. I've never trained a model on resumes, but I definitely often see this "lower certainty on minorites" thing for models I do train. The lower certainty would in turn lead to lower rankings for women even without any bias in the data. Now, I'm not saying that Amazon's data isn't biased. I would not be surprised if it were. I'm just saying we should be careful in understanding what is evidence of bias and what is not.
- tomp 8y ago> The lower certainty would in turn lead to lower rankings for women even without any bias in the data. I don't think that's true. "No bias" means that gender is irrelevant (i.e. its correlation with outcome is 0%). Therefore the system shouldn't even take it into account - it would evaluate both men and women just by other criteria (experience, technical skills, etc), and it would have equal amounts of data for both (because it wouldn't even see them as different). You need bias to even separate the dataset into distinct categories.
- theptip 8y ago> "No bias" means that gender is irrelevant False. If we're talking about the technical statistical definition, bias means systematic deviation from the underlying truth in the data -- see this article by Chris Stucchio with some images for clarification: https://jacobitemag.com/2017/08/29/a-i-bias-doesnt-mean-what-journalists-want-you-to-think-it-means/ https://jacobitemag.com/2017/08/29/a-i-bias-doesnt-mean-what... "In statistics, a “bias” is defined as a statistical predictor which makes errors that all have the same direction. A separate term — “variance” — is used to describe errors without any particular direction. It’s important to distinguish bias (making errors with a common direction) from variance which is simply inaccuracy with no particular direction."
- deleted 8y ago[deleted]
- tomp 8y agoI think the comments I replied to mean bias as in “sexist bias”.
- grandmczeb 8y agoBias as in racism, sexism, etc, has multiple definitions, some of which are mutually exclusive.
- theptip 8y ago
- kareemsabri 8y agoThis doesn’t seem to be a reasonable conclusion. There is no reason to assume the AI’s assessment methods will mirror those of the recruiters. If Amazon did most of it’s hiring when programming was a task primarily performed by men, and so Amazon didn’t receive many female applicants, they could be unbiased while still amassing a data set that skewed heavily male. The machine would then just correctly assess that female resumes don’t match, as closely, the resumes of successful past candidates. Perhaps I’m ignorant about AI, but I don’t see why the number of candidates of each gender shouldn’t increase the strength of the signal. “Aggressiveness” in the resume may be correlated but not causal. If the AI was fed the heights of the candidates, it might reject women for being too short, but that would not indicate height is a criteria of Amazon recruiters hiring.
- kaitai 8y agoThe whole aim of the AI was to make decisions like the recruiters did -- that is explicitly what they were aiming to do. It might be worth reading the article as it addresses your two ideas (the aim of the project and the fact that the training set was indeed heavily male).
- kareemsabri 8y agoHey. I did read the article. It doesn’t support the conclusion OP is drawing. The aim of the AI is to “mechanize the search for talent”. It doesn’t care to, nor have any means to, make decisions “like the recruiters did”. Obviously machines don’t make decisions like humans do. They’re trying to reverse engineer an alternate decisions making process from the previous outcomes.
- erikpukinskis 8y agoAren't the "previous outcomes" past hiring decisions though?
- kareemsabri 8y agoYes, but you have to know what pool you started with. As an overly simplistic example, if a bank used historical mortgage approval records from primarily German neighbourhoods to train AI, it might become racist against non-Germans despite that it’s just an artifact of the demographics of the time. I think it just shows how not ready for prime time AI is.
- cal97g 8y agoOr maybe it recognised that women were consistently the worst candidates.
- IshKebab 8y agoThey didn't scrap it because of this gender problem. That wasn't why it failed. They scrapped it because it didn't work anyway. Note the title is "Amazon scraps secret AI recruiting tool that showed bias against women" not "Amazon scraps secret AI recruiting tool because it showed bias against women". But I guess the real title is less clickbaity - "Amazon scraps secret AI recruiting tool because it didn't work".
- stcredzero 8y agoThe same AI should be applied to hiring nurses and various other fields which show population skews in gender, as well as fields which are not skewed. I'd be curious as to the outcome.
- Consultant32452 8y agoWithout regard to this particular issue, you also have to concern yourself with the bias of the person determining if the AI has a bias.
- louwrentius 8y agoThanks for spelling this out, I think this is exactly how to look at this.
- gambler 8y agoThe article didn't specify how they labeled resumes for training. You're assuming that it was based on whether or not the candidate was hire. Nobody with an iota of experience in machine learning would do something like that. (For obvious reasons: you can't tell from your data whether people you did not hire were truly bad.) A far more reasonable way would be to take resumes of people who were hired and train the model based on their performance. For example, you could rate resumes of people who promptly quit or got fired as less attractive than resumes of people who stayed with the company for a long time. You could also factor in performance reviews. It is entirely possible that such model would search for people who aren't usually preferred. E.g. if your recruiters are biased against Ph.D.'s, but you have some Ph.D.'s and they're highly productive, the algorithm could pick this up and rate Ph.D. resumes higher. Now, you still wouldn't know anything about people whom you didn't hire. This means there is some possibility your employees are not representative of general population and your model would be biased because of that. Let's say your recruiters are biased against Ph.D.'s and so they undergo extra scrutiny. You only hire candidates with a doctoral degree if they are amazing. This means within your company a doctoral degree is a good predictor of success, but in the world at large it could be a bad criteria to use.
- jonny_eh 8y agoMen are promoted quicker, and more often, than women.
- deegles 8y agoThere was a company meeting one year at Amazon when they proudly announced that men and women were paid within 1-2% of each other for the same roles. It completely missed the point which you raise. I want to see reports of average tenure and time between promotions by gender. I suspect that the reason we don't see those published is that the numbers are damning.
- zaarn 8y agoOr possibly noone did a study of sufficient size that passed peer review. It's also not hard to make the pay gap 1-2% just like it's not hard to make it 25% (both values are valid). Statistics is a fun field. Don't trust statistics you didn't fake yourself. Amazon could easily cook the numbers to get to 1-2%, I doubt anyone checked the process of determining that number if it's unbiased and fair and accounts for other factors or not.
- monochromatic 8y ago> The AI becoming biased tells that the "teacher" was biased also. That doesn’t follow.
- s73v3r_ 8y agoSomeone had to decide on the training material. Note that saying that they had bias does not mean that they acted with malicious intent; most likely they didn't. That doesn't change the outcome, however.
- roenxi 8y agoDo you have some information not present in the article? There seem to be some assumptions on the training process in your comment that are not sourced in the article. I'll don my flack jacket for this one, but based on population statistics I believe a statistically significant number of women have children. A plausible hypothesis is that a typical female candidate is at a 9 month disadvantage against male employees and that that is a statistically significant effect detected by this Amazon tool. Now, the article says that the results of the tool were 'nearly random', so that probably wasn't the issue. But just because the result of a machine learning process is biased does not indicate that the teacher is biased. It indicates that the data is biased, and bias always has a chance to be linked to real-world phenomenon.
- brown9-2 8y agoDoes Amazon give 9 months of parental leave, or are you saying women employees are disadvantaged for their entire pregnancy?
- roenxi 8y agoAh. Sorry, silly me. A quick search suggests 20 weeks, so ~4.5 months. Obviously I don't have much specific insight, so maybe there is a culture where they don't use leave entitlements. But if there are indicators that identify a sub-population taking a potentially 20 week contiguous break it is entirely plausible that it would turn up as a statistically significant effect in an objective performance measure. All else being equal, then a machine learning model could pick up on that. The point isn't that it is the be-all and end all, just that the model might be picking up on something real. There are actual differences in the physical world.
- dheera 8y agoThe term "AI" is over-hyped. What we have now is advanced pattern recognition, not intelligence. Pattern recognition will learn any biases in your training data. An intelligent enough* being does much more than pattern recognition -- intelligent beings have concepts of ethics, social responsibility, value systems, dreams, ideals, and is able to know what to look for and what to ignore in the process of learning. A dumb pattern recognition algorithm aims to maximize its correctness. Gradient descent does exactly that. It wants to be correct as much of the time as possible. An intelligent enough being, on the other hand, has at least an idea of de-prioritizing mathematical correctness and putting ethics first. Deep learning in its current state is emphatically NOT what I would call "intelligence" in that respect. Google had a big media blooper when their algorithm mistakenly recognized a black person as a gorilla [0]. The fundamental problem here is that state-of-the-art machine learning is not intelligent enough. It sees dark-colored pixels with a face and goes "oh, gorilla". Nothing else. The very fact that people were offended by that is a sign that people are truly intelligent. The fact that the algorithm didn't even know it was offending people is a sign that the algorithm is stupid. Emotions, the ability to be offended, and the ability to understand what offends others, are all products of true intelligence. If you used today's state-of-the-art machine learning, fed it real data from today's world, and asked it to classify them into [good people, criminals, terrorists], you would result in an algorithm that labels all black people as criminals and all people with black hair and beards as terrorists. The algorithm might even be the most mathematically correct model. The very fact that you (I sincerely hope) cringe at the above is a sign that YOU are intelligent and this algorithm is stupid. *People are overall intelligent, and some people behave more intelligently than others. There are members of society that do unintelligent things, like stereotyping, over-generalization, and prejudice, and others who don't. [0] https://www.theverge.com/2018/1/12/16882408/google-racist-gorillas-photo-recognition-algorithm-ai https://www.theverge.com/2018/1/12/16882408/google-racist-go...
- platz 8y ago"a worldview built on the important of causation is being challenged by a preponderance of correlations. The possession of knowledge, which once meant an understanding of the past, is coming to mean an ability to predict the future." - Big Data (Schonberger & Cukier) so, knowledge now is allegedly possession of the future, rather than possession of the past. This is because the future and past are structurally the same thing in these models. Each could be missing, but re-creatable links. Also, conflicting correlations can be shown all the time. if almost any correlation can be shown to be real, what's true? How do we deal with conflicting correlations?
- Yoyoyou 8y agoIt failed because rationally interpreting gender data leads to politically incorrect conclusions.