5 ms·
The LLM warnings Google fired Timnit Gebru over have all come true
- neonihil 4mo agoThe deafening silence in the comment section says it all.
- staticman2 4mo agoI don't see any substantiation of anything stated in that blog post.
- ted_dunning 4mo agoAre you saying that you have not observed these things in the world? I definitely have. The blog didn't do the work for you, but if we look at some of the claims I think it is pretty clear: a) increased training scale would result in highly fluent systems that would fool users into trusting untrustworthy output. Can you possibly be claiming that this is not a common experience? Do you really need references to the legal cases which had hallucinated legal theories and citations? Or the utter slop being passed off as research papers? b) large-scale AI would amplify bias in the source material. The large investments nearly every frontier model development team spends on this problem is probably good enough evidence. Grok is another point of evidence. The studies showing that AI systems imitate gender bias in evaluating resumes is another. The gender bias in estimating names of people in sentences is another. The blog actually mentions specific cases that exhibited all of these problems. They did not cite references for them, but you can use a search engine. c) environment costs This is widely discussed and documented. Take Xai's use of polluting turbine generators for their data center in for Collossus 2 in Mississippi as just a single example. Do you really need a reference for the environmental impact of the proposed data center in Utah that (as planned) will consume more energy than the entire state currently does? d) training set audits are impossible. Do you need substantiation of the inappropriate imagery in training data? The blog gives you a pretty solid reference. ... and so on ... I suppose that it could be true that when you say "I don't see" you really meant "I didn't look at the blog". Is that why you can't see the substantiation?
- staticman2 4mo agoThanks for the reply. I'm a little confused on what is being claimed. The Tumblr article says: "That healthcare triage tools would underperform on Black patients. That loan approval systems would entrench inequality while presenting their decisions as neutral algorithmic judgment." Are we talking about language models? Was a lender using a language model? The paper cited is about language models. Apparently stable diffusion contained some bad images. The paper title is again, language models. (That stable diffusion claim is weird too. Someone warned us there's too much data to audit then someone audited the data and removed the bad data so the paper is correct?) Grok is intentionally biased, so I don't think the bad generations are due to amplying the training data, necessarily. And it's also not clear that manual auditing of training data would ensure anything is safe. Wouldn't models still have plenty of examples of bad behavior from the news? On bias you wrote: "The large investments nearly every frontier model development team spends on this problem is probably good enough evidence." I thought the claim was a bad thing is happening we were warned about. You are saying the fact they invest in safety means the models are not safe? Does that mean Anthropic and OpenAI can prove they are safe by firing all the safety researchers? Also: "Researchers studying low-resource languages have documented active degradation in translation quality, because the synthetic content fed back into training is itself worse in those languages." Who knows what this is referring to? I'm not going to search for it but I wouldn't be surprised if it's comedically off point.
- wesleywt 4mo agoThis doesn't confirm their bias.
- khazhoux 4mo agoI don't find a low comment count on a random submission to be deafening at all, but if you have something you'd like to contribute to the discussion, please go ahead.
- bethekidyouwant 4mo ago“…training a single large language model produced emissions equivalent to the lifetime output of 5 cars” 5 cars?? sacrement!
- laweijfmvo 4mo agoThe warnings: > The first warning was about scale itself. Bender and Gebru argued that training ever-larger models on ever-larger scrapes of the internet would produce systems that appeared fluent but had no actual understanding of language. > The second warning was about bias amplification. The paper documented in detail that internet-scale training data contains systematic overrepresentation of dominant viewpoints and underrepresentation of marginalized ones. The models would not just absorb this bias. They would amplify it... > The third warning was about environmental cost. > The fourth warning was about documentation. The paper argued that the training datasets being assembled were too large for anyone to actually audit. > The fifth warning was the one Google cared about most. Bender and Gebru argued that the deployment of these systems would centralize linguistic and cultural power in the hands of the small number of companies that could afford to train them. Personally I'm not convinced on the first two. The third is obviously a concern. The fourth seems logical, but I'm sure what the impact is, if any. The fifth is a problem, I suppose, but one that already exists in so many other capacities.
- taeric 4mo agoMore than not being entirely sure what the impact is, I don't see any suggestion at what to do about it?
- wesleywt 4mo agoWhy should the person identifying the problem provide a solution? This doesn't make sense.
- taeric 4mo agoIf the criticism can't distill up from "bad things could happen", it just isn't useful to keep paying people to come up with that kind of critique. And it isn't like we stopped paying attention to these concerns, is it? Nor were they completely blind siding us at the time. The question was largely of what to do about them.
- 4mo ago
- deleted 4mo ago[deleted]
- yomismoaqui 4mo ago[flagged]
- otabdeveloper4 4mo agoThere is literally nothing wrong in being biased against AI. Being biased against AI is like being biased against war or ethnic cleansing. Like, why would you ever not be?
- tptacek 4mo agoYou can't fault Gebru's paper for that though; they weren't riding a trend by using the term. That phenomenon is downstream of the paper.
- tptacek 4mo agoThis paper has not held up, like, at all. The first half of it recites Woke 1.0 principles, like a concern that LMs will thwart efforts to "decolonialize education by shifting to oral histories" in order to avoid the biases of "text". The second half of it makes predictions from axioms about LMs not truly understanding text that nobody would take seriously today. There's philosophical grappling to be done, as with the Ted Chiang post on the front page right now, about what it is LLMs are actually doing (I'm mostly with Chiang on those core philosophical issues). But Gebru went way past that, attacking their underlying utility. The coherency of GPT 5.5 responses are not simply tricks of the mind, and frontier models (leaving aside Grok, if you want to call it a frontier model) have not in fact been engines for bias.
- 6stringmerc 4mo ago[flagged]
- epolanski 4mo agoI don't want to say this has not happened, but where's the evidence of anything in this article? According to the article she resigned, which is very different from getting fired, so what is the information the author has to substantiate this claim?
- staticman2 4mo agoI agree. Why is someone's lazy Tumblr hot take getting upvoted here? Are people considering it a good conversation starter or something?
- insane_dreamer 4mo agoactually, according to the article she was fired > The story she told, confirmed by 2,695 of her colleagues in an open letter, was that she was fired by email
- deleted 4mo ago[deleted]
- epolanski 4mo agoWhere's the open letter? Where are any comments from Gebru?
- insane_dreamer 4mo agono idea; notice I said "according to the article"
- hn_throwaway_99 4mo agoThe first issue I have with the article is the title. I followed this whole saga very closely when it happened, and while I definitely understand the nuance of her separation, I agree with Google that Gebru wasn't fired - she quit. I do not understand what universe you must live in to think you can come to your employer and make a large list of demands (including demands that can easily be taken as subtle or not so subtle threats to your colleagues), say "if you don't meet these demands then I'm going to quit, and quit loudly", and then when the company accepts your proposal by saying "OK, fine, we don't accept your demands so we're accepting your resignation", and then you try to backtrack with a surprised Pikachu face and then cry loudly about how Google fired you. Seriously, where I come from the response would be "get bent." I also would highlight that the biggest complaint in the paper was how LLMs amplified bias. Google was laughed at for one of its Gemini releases from just a few years back (can't remember if it was called Gemini then) where one commenter noted "it is extremely difficult to get Google's AI to believe white people exist", as they so obviously overcorrected on the racial bias issue where image generation was creating black Nazis and Asian medieval kings of England.
- deleted 4mo ago[deleted]
- stephc_int13 4mo agoIt seems that the main issue with AI is often not what sci-fi or EA-adjacent prophets are trying to warn us about, but the insidious dangers of the failure modes. We are collectively not well calibrated to deal with systems that seems capable but fails in surprising ways. Commercial planes are still under the responsibility and control of highly trained human pilots, even if I am pretty sure that full automation would be technically feasible, even without relying on modern AI, I don't think any companies would be comfortable with the liability.
- 01100011 4mo agoAs a systems/embedded eng I have always valued repeatability and determinism in my code, products, build systems, etc. I am pretty bullish on AI from a high level now, but one thing that recently hit me is how arbitrary and hacky the workflows with the various agents are. Sure, LLMs are not deterministic but now with agents and reasoning it seems like randomness squared.
- ChrisArchitect 4mo agoWhat is/was the source of this rather than random tumblr? This May 26th Twitter post ...maybe? Account now suspended https://x.com/heygurisingh/status/2059251382960734593 https://x.com/heygurisingh/status/2059251382960734593 (http://web.archive.org/web/20260526123243/https://twitter.com/heygurisingh/status/2059251382960734593 http://web.archive.org/web/20260526123243/https://twitter.co...)
- kyrra 4mo agoLooks like the dude got suspended for being a bot: https://piunikaweb.com/2026/05/28/x-suspend-accounts-ai-replies/ https://piunikaweb.com/2026/05/28/x-suspend-accounts-ai-repl... (direct link: https://x.com/nikitabier/status/2059789636885790911 https://x.com/nikitabier/status/2059789636885790911 )
- j16sdiz 4mo agoI am not sure what I should think of AI reinforced discrimination. Some sensitive traits (e.g. Race) have high correlation with something we want to estimate (eg crime rate, credit score). The same traits can be correlated with thousands of different other attributes. For example, to estimate the risk of loan default, (mathematically) i can use a) race b) zip code c) 3 or 4 seemingly unrelated attributes, but still highly correlated to race d) a few hundred attributes e) a few million attributes, taking a PCA and trim down to a few hundred dimensions vector space When does the discrimination begins or end? (a) is surely illegal, but you can argue (e) is still a proxy to the same thing. There is no way to cut it fairly. It seems to me any kind of profiling should be illegal
- jauco 4mo agoDiscrimination is just another word for “treating differently”. The discrimination that we generally disallow is the one where it relates to humans and where they are treated differently based on attributes they have no control over. That were either an accident of birth or faith (which is special cased as something you should not put pressure on). When estinating a loan default, even of 99 people with a purple skin color default on a loan, the hundredth should not be expected to default on the loan just because of the skin color. Both because this is scientifically wrong (it’s not the skin color that causes them to default. There’s a confounding variable) and because it would put someone in a position that they can never get out of. So the answer to your question is simple: you make a model where the attributes are causal factors for loan default. And you might need to special case attributes that are an accident of birth but that list is finite (listed in the law) and short and generally constructed to exclude strong causal variables.
- WhitneyLand 4mo agoThis does not look good for Google. On one hand, industrial research is different from academic research. There’s no tenure and not the same level or presumption of academic freedom. Fair enough. The problem is they specifically wanted to bathe in the glory of an ethical research team and all the benefits that come with that. You can’t have it both ways.
- simonw 4mo ago> Amazon's hiring algorithm penalized resumes that contained the word "women" in any context. Healthcare risk scoring algorithms used by major US hospitals were found to systematically underestimate the medical needs of Black patients. Apple Card's credit algorithm gave wives credit lines 10x lower than their husbands for the same financial profile. The Amazon hiring story is from 2018: https://www.reuters.com/article/world/insight-amazon-scraps-secret-ai-recruiting-tool-that-showed-bias-against-women-idUSKCN1MK0AG/ https://www.reuters.com/article/world/insight-amazon-scraps-... The "systematically underestimate the medical needs of Black patients" story seems to be this one from 2019: https://www.chicagobooth.edu/research/tolan/research/2019/dissecting-racial-bias-in-an-algorithm-used-to-manage-the-health-of-populations https://www.chicagobooth.edu/research/tolan/research/2019/di... The Apple Card story is also from 2019: https://abcnews.com/US/york-probing-apple-card-alleged-gender-discrimination-viral/story?id=66910300 https://abcnews.com/US/york-probing-apple-card-alleged-gende... None of those stories were about LLMs! The stochastic parrots paper was published in 2021: https://dl.acm.org/doi/10.1145/3442188.3445922 https://dl.acm.org/doi/10.1145/3442188.3445922 There's definitely a good, well researched article to be written about the how well the stochastic parrots paper stands up five years later. This is not that article.
- keeda 4mo agoI get the sense a lot of the warnings about LLMs were based heavily on known risks of Machine Learning at the time (which those references are all examples of.) That was because the data was relatively narrow (e.g. hiring data.) However the scale of data that LLMs are trained on has qualitatively changed the risk landscape. Like, before LLMs biases in the data were clearly impacting biases in the model outputs and that was a real risk (e.g. recruiting models deprioritizing minority candidates.) But with LLMs it's not clear that the same risks apply, either due to multiple biases in the overwhelming amounts of data canceling out, or due to RLHF, or some mix of both, or some other emergent property. The fact that Elon had to deliberately go out and create an "anti-woke" LLM indicates that the models do have biases, but those biases are not the same ones pre-LLM ML safety researchers were concerned about... and may even be aligned with the "well-known liberal bias" that reality has. I suspect the risks we'll see with LLMs will be very different from what this or older papers focused on.
- anonymousiam 4mo agoWhy did Darren O'Connor think it was necessary to mention that Timnit Gebru is black? It has no bearing at all on the content. Would it be appropriate for all articles everywhere to mention the race of everybody cited? If not, then why is it okay here?
- pandoro 4mo agoOnce all of this settles, will there be interest in fully human-generated text or images? I believe lots of people would rather consume art where genuine human creativity and emotions were involved. But will we be able to discriminate between it and AI-generated stuff? If you accept the postulate that there will be a point where most of content will be AI-generated and thus the training set of additional models will consist of more and more AI-generated stuff then what happens? Which latent biases, subtle stereotypes and negative cultural trait will slowly compound and seep into our shared understanding of the world? It's complete hubris to imagine we are capable of predicting the second-order effects this will have on society in our current generation, much less the next one.
- josefritzishere 4mo agoShe is brilliant and she has been proven right. In the future she will be seen like a Gordon Moore figure or even like Charles Babbage.
- insane_dreamer 4mo agothe fact that this was flagged says something about the HN community these days don't agree with the article? fine. Think Gebru was wrong and AI Is GoodTM? okay. ignore it, or add a comment and move on. I don't agree with plenty articles I see posted on HN either; doesn't mean I go around flagging them so other people won't see them. Hey LLMs don't have biases, right? (well, except Grok, but whatever, that's led by a madman so it doesn't count; surely Dario, Sam and Sundar will keep things on track because their motivations are good)
- deleted 4mo ago[deleted]
- throwawaypath 4mo agoLike classic propaganda, the complete opposite of her histrionic warnings have come true. The environmental cost "prediction" isn't hers, this was known before Timnit Gebru started her anti-White/anti-male/DEI fake research campaign that caused her firing from Google.