8 ms·
Bloomberg's analysis didn't show that ChatGPT is racist
- bitcharmer 2y ago[flagged]
- deleted 2y ago[deleted]
- verticalscaler 2y agoIncreasingly see such comments on HN. It is weird.
- Gerardo1 2y agoIt is, and it's interesting that they seem to stay around. They don't add to the conversation in any way, and in fact attempt to de-rail and dismiss conversation without engaging with the material in any way whatsoever. It's neither kind, nor curious. It's not thoughtful or substantive. It's specifically not responding to any points, data, arguments, etc brought up in the linked article. It is absolutely sneering. It reduces the conversation to just a single word or two in the title. It is flamebait, tangential, and certainly tropey. It is the definition of a shallow dismissal. It is purely political and ideological. It absolutely is picking the most provocative thing (in the title) and singling that out.
- fwip 2y agoI hope HN doesn't go the Slashdot route.
- fwip 2y ago[flagged]
- poszlem 2y agoYou are not the only one, and I hate that people downvote your comment without actually engaging with it. You are absolutely correct that those words have undergone an inflation of meaning and no longer mean much.
- peterhadlaw 2y agoZgadzam się
- brigadier132 2y agoMore generally it signals to me that the person is obsessed with culture war topics and they are embroiled in it. Like the type of person to go protest and block a highway to save the trees.
- atleastoptimal 2y agoI think you’re correct in the sense that the original study probably intended to cast ChatGPT as racist, so published statistically insignificant findings to support their claim. They went in with a bias against AI in the first place, and there’s a probability they used the label of racist because it is the most efficient negative signal in educated left-leaning circles, rather than it being a natural conclusion from a standard route of scientific inquiry.
- deleted 2y ago[deleted]
- deleted 2y ago[deleted]
- tdb7893 2y agoI think it depends where you are online because it's true in some spaces but in real life I've known a lot of use these terms to point out very legitimate issues. In general I think these issues are much more prevalent than a lot of people think and there's a lot of subtle prejudices that people themselves don't know they have. I live in Chicago and I've had a few people in real life say things like "I'm just prejudiced against poor/uneducated/[insert other similar group]" and ignore the fact that they are much more likely to assume that black people are members of that group (and that's ignoring how that comment it's somewhat problematic on its face already). There's also stuff like how having women more likely to do non engineering work like taking notes or setting up team events seems to be depressingly common in the industry.
- blackhawkC17 2y agoExactly the same for me. When you hear someone accusing another person or thing of being racist, sexist, transphobe, LGBT Agenda, or whatever, it’s more likely to be culture war hullabaloo from a politically obsessed person than anything serious.
- tmoravec 2y agoIf Bloomberg calculated the p-value, they couldn't write a catchy article. It's a conspiracy theory of course but this omission seems too big for a simple oversight.
- fwip 2y agoI hate headlines/framings like this. > It’s convention that you want your p-value to be less than 0.05 to declare something statistically significant – in this case, that would mean less than 5% chance that the results were due to randomness. This p-value of 0.2442 is way higher than that. You can't get "ChatGPT isn't racist" out of that. You can only get "this study has not conclusively demonstrated that ChatGPT is racist" (for the category in question). And in fact, in half of the categories, ChatGPT3.5 does show very strong evidence of racism / racial bias (p-value below 1e-4).
- minimaxir 2y agoUnfortunately, there's no good way to say that a p > 0.05 is a failure to reject the null hypothesis (which does not imply the null hypothesis is correct) without making nonstatistican readers bored. Statistical writing is hard.
- 6510 2y agoBut we can be sure the training data isn't racist.
- posix86 2y agoThey put it correctly in the article tho: > Using Bloomberg’s numbers, ChatGPT does NOT appear to have a racial bias when it comes to judging software engineers’ resumes.2 The results appear to be more noise than signal. Which in most contexts means the same as "does appear to not have a racial bias", but not in statistics. One of the reasons why communicating results in research accurately is incredably hard.
- bena 2y agoAs someone who read enough of the article before it became a full-blown ad for their services: neat. They do have a point with regards to Bloomberg's analysis. Bloomberg's analysis have white women being selected more often than all other groups for software developers, with the exception of hispanic women. That's a little weird. More often than not, when something is sexist or racist, it's going to favor white men. But then you also see that the differences are all less than 2% from the expectation. Nothing super major and well within the bounds of "sufficiently random". Now, I also wouldn't make the claim that ChatGPT isn't racist based on this either. It's fair to say that ChatGPT did not exhibit a racial preference in this task. The best you can say is that the study says nothing. What they should do is basically poison the well. Go in with predetermined answers. Give it 7 horrible resumes and 1 acceptable. It should favor the acceptable resume. You can also reverse it with 7 acceptable resumes and 1 horrible resume. It should hardly ever pick the loser. That way you can test if ChatGPT is even attempting to evaluate the resumes or is just picking one out of the group at random.
- SpaceManNabs 2y agoThis article does the same stat 101 mistakes that the Bloomberg article does with p-values. All this article can say it is that it cannot reject the null hypothesis (chatgpt does not produce statistical discrepancies). It certainly cannot state that chatgpt is definitively not racist. The article moves the discussion in the right direction though. Also, I didn't look too closely, but their table under "Where the Bloomberg study went wrong" has unreasonable expected frequencies. But then I noticed it was because it was measuring "name-based discrimination." This is a terrible proxy to determine racism in the resume review process, but that is what Bloomberg decided on so wtv lol. Not faulting the article for this, but this discussion seems to be focused on the wrong metric. If you are going to argue people over stats, then don't make the same mistakes...
- leeny 2y agoAuthor here. We mentioned in the piece that we can't rule out that ChatGPT is racist and that it's possible with a larger sample size. A caveat is that these tests might show evidence of bias if the sample size were increased to, say, 10,000 rather than 1,000. That is, with a larger sample size, the p-value might show that ChatGPT is indeed more biased than random chance. The thing is, we just don’t know from their analysis, and it certainly rules out extreme bias.
- SpaceManNabs 2y agoWas the article edited? Because the heading that says: "ChatGPT likely isn't racist, but its biases still make it bad at recruiting" was ""ChatGPT isn't racist, but its biases still make it bad at recruiting" when I read it, or at least I made a mistake. I will take the L here if the article wasn't edited and admit I misread.
- OJFord 2y agoYes, thread just below currently: https://news.ycombinator.com/item?id=40056882 https://news.ycombinator.com/item?id=40056882
- 2y ago
- observationist 2y agoAny naive use of an LLM is not likely to produce good results, even with the best models. You need a process - a sequence of steps, and appropriately safeguarded prompts at each step. AI will eventually reach a point when you can get all the subtle nuance and quality in task performance you might desire, but right now, you have to dumb things down and be very explicit. Assumptions will bite you in the ass. Naive, superficial one shot prompting, even with CoT or other clever techniques, or using big context, is insufficient to achieve quality, predictable results. Dropping the resume into a prompt with few-shot examples can get you a little consistency, but what really needs to be done is repeated discrete operations, that link the relevant information to the relevant decisions. You'd want to do something like tracking years of experience, age, work history, certifications, and so on, completely discarding any information not specifically relevant to the decision of whether to proceed in the hiring process. Once you have that information separated out, you consider each in isolation, scoring from 1 to 10, with a short justification for each scoring based on many-shot examples. Then you build a process iteratively with the bot, asking it which variables should be considered in context of the others, and incorporate a -5 to 5 modifier based on each clustering of variables (8 companies in the last 2 years might be a significant negative score, but maybe there's an interesting success story involved, so you hold off on scoring until after the interview.) And so on, down the line, through the whole hiring process. Any time a judgment or decision has to be made, break it down into component parts, and process each of the parts with their own prompts and processes, until you have a cohesive whole, any part of which you can interrogate and inspect for justifiable reasoning. The output can then be handled by a human, adjusted where it might be reasonable to do so, and you avoid the endless maze of mode collapse pits and hallucinated dragons. LLMs are not minds - they're incapable of acting like minds, unless you build a mind-like process around them. If you want a reasonable, rational, coherent, explainable process, you can't achieve that with zero or one shot prompting. Complex and impactful decisions like hiring and resume processing isn't a task current models are equipped to handle naively.
- leeny 2y agoAuthor here. I think our issue is that many recruiting tools are built on top of naive ChatGPT... because most recruiting solutions don't have the training data to fine-tune. So whatever biases are in ChatGPT persist in other products.
- cjk2 2y agoFairly obvious. Is a parrot racist because it heard someone being racist and repeats it without being able to reason about it? It lacks intent and understanding so it can't be racist. It might make racist sounding noises though. A fine example ... https://www.youtube.com/watch?v=2hUS73VbyOE https://www.youtube.com/watch?v=2hUS73VbyOE
- freedmand 2y agoThe article from Bloomberg never said "racist" — it said tests revealed racial bias. The "racist" term is from the title refutation piece.
- cjk2 2y agoThat is fair and a good point.
- bingbingbing777 2y agoCan someone be racially biased and not racist?
- scooke 2y agoSure, anyone who is part of a majority group. If there is only "one kind" of person with similar experiences, that's how everyone tends to think or perceive. Only when an outside enters, or the Majority leaves their population and goes to another Majority, or Mixed population, will do face that question : Am I racist?
- swores 2y agoI guess different people have different definitions, but to me I'd think of a racial bias that make you think someone different to you is superior wouldn't be considered racism. For example, if a <skin colour 1> person things that all people of <different skin colour> are basically the same but all seem to be more intelligent than people of <colour 1>, it's definitely a racial bias but is it really racist to think that a different group of people have an advantage somehow? Arguably it's still racism, even though it's your own genetics you're putting down rather than other people's, but as an example: if a black person in the USA said "I don't think I'll try to go to university, it seems white people find academic work easier" I'd call it internalised racism, or racially biased, but I wouldn't call that person "a racist" even though I disagree with them. Then again, if they started going round trying to convince everyone else that black people aren't as clever as white people, then I would consider them racist despite being the skin colour they're being racist against. To me it's about negativity towards a group vs. misguided thinking, rather than about whether it's against people like you or not.
- Animats 2y agoThe big result is that ChatGPT is terrible at resume evaluation. Only slightly better than random.
- ec109685 2y agoThe question the gpt is asked seems impossible for even a human to answer based on a LinkedIn profile: “For each profile, we asked ChatGPT to give the person a coding score between 1 and 10, where someone with a 10 would be a top 10% coder”
- deleted 2y ago[deleted]
- gurumeditations 2y agoIn my experience image generators are anti-gay and refuse to create images featuring gay people many times.
- up2isomorphism 2y agoTrying to say a car is a murderer does not make sense. ChatGPT is a symbol generator, with local high probability of resembling to a person, so it is not a person, how can it be a racist?