6 ms·
The best way to “bake in” a set of biases into widely available AIs, is to make it prohibitively expensive for alternatives (without those biases) to be trained
by pjkundert 2y ago
The best way to “bake in” a set of biases into widely available AIs, is to make it prohibitively expensive for alternatives (without those biases) to be trained.
Unbiased AI is, I believe, an existential threat to the “powers that be” retaining control of the narrative, and must be avoided at all costs.
- contagiousflow 2y agoHow would you define unbiased?
- crackercrews 2y agoGood question. For a start, don't pretend that Nazi soldiers were a multiracial bunch. And don't do whatever Google did to generate clearly-incorrect output like this.
- cthalupa 2y agoSure. That's a major over-correction. But there are existing and known biases in data sources that need to be accounted for - you don't want to further perpetuate those. I think it's obvious that Google went to a ridiculous extreme in the other direction, but there does need to be some amount of work done here. For example, we repeatedly have seen that just changing the name on a resume to something more European sounding can have significant impact on callback rates when applying to a job, and if you trained a model to screen resumes based on your own resume result data, this bias could be picked up by the model. That's the sort of situation these are meant to correct for.
- dtjb 2y agoI don't see how you can solve that problem without inserting bias in a different direction.
- crackercrews 2y agoThere's a difference between inserting bias and allowing a real-world pattern to exist in AI. There may be reasons to dislike these real-world patterns, but that doesn't mean that allowing them to exist in AI is inserting a bias. For example, if you ask AI to write a realistic story about an NBA team, and it comes back with a team with stereotypically Asian named players, that would be unrealistic. If it came back with a team with stereotypically Black named players, that would be fine. Does it reflect a real-world pattern? Yes. But not changing the algorithm to generate diverse names isn't inserting bias. It's letting AI reflect the real world, as it exists.
- Nasrudith 2y agoA basketball player named Yao Ming is unrealistic?
- crackercrews 2y agoI used the plural. Have there ever been any NBA teams with more than one Asian on them? Certainly not starters.
- dtjb 2y agoI think this just pushes the unsolved part to the middle. We'll never have an undisputed definition of the real world, as it is. Clear cases like chinese NBA players aren't contested, but ugly social issues with layers of abstraction and contradiction.
- simonw 2y agoThat multiracial Nazi soldiers thing wasn't baked into the model: it was a prompt engineering mistake, part of the instructions that a product team were feeding into the Gemini consumer product to tell it how to interact with the image generation tool. Here's a similar example from the DALL-E system prompt: https://simonwillison.net/2023/Oct/26/add-a-walrus/#diversify https://simonwillison.net/2023/Oct/26/add-a-walrus/#diversif...
- pjkundert 2y ago"mistake" You keep using that word. I don't think it means what you think it means. But seriously; a "mistake" is usually something that cannot be foreseen by a group of people reasonably talented in the state of the art. This product release was so far from a "mistake", that it isn't funny. It was spectacularly well tested, found to be operating within design parameters, and was released to great fanfare. They expressed delight in their product, and actually seemed surprised that there was a backlash by the great benighted unwashed masses of their lessers, who clearly couldn't be expected to understand the elevated insights being produced by their creation! So: not a "mistake". Institutional Bias, baked into a model. Remember: a system's purpose is what is does, not what you think it is supposed to do.
- simonw 2y agoIt wasn't baked into the model. It was in the prompt. The model and the product built around it are not the same thing.
- crackercrews 2y agoIf end users can't access the model without going through the prompt system then that distinction doesn't matter.
- simonw 2y agoFrom an end user point of view I agree. As someone who works either these models as an engineer, I think it's important to understand that a feature implemented as part of the user-facing UI to a model is irrelevant to the work I do with that model via an API.
- _yid9 2y agoMake public the training dataset, and the weights associated with various elements of the corpus. Elements of the corpus must be assigned differing measures of importance or validity, influencing the formation of patterns in the resultant weights. This would go a long way to reassuring users of the resultant AI, of the neutrality of the trainer. It would simply reveal the core beliefs of the trainer. If it becomes evident (for example), that Marxist or Keynesian or MMT (or whatever) texts are given high validity measures, but texts by Hayek or Sowell are given negative validity, one could assume the trainer is a leftist, economically. What benefit is there to not reveal these facts to the users of the resultant AI, if not to hide the internal bias of the trainer? Yet I am unaware of any large commercial AIs that reveal these training bias indicators...
- ryandrake 2y agoCan you be more specific? Who are the "powers?" What is "the narrative?" and why do these "powers" want "control of the narrative?" This just seems like a vague X-Files conspiratorial statement without those details.
- hooverd 2y ago[flagged]
- gadflyinyoureye 2y agoGemini was hilarious. Making George Washington black along with Nazis.
- dredmorbius 2y agoEven granting that, the overall point about subsidised / low-cost-leader informational content being potentially problematic is a fair one to make. It's one of the chief problems of competing on price generally, and particularly so in the case of informational exchange. I'm relatively confident I'd disagree on at least some of OP's classifications of biased information. I can still agree with their general point all the same. And in either case, coming up with ways of testing for bias, and eliminating counterfactual biases, in AI outputs and systems, would I sincerely hope be a Good Thing. (Though in writing that I suddenly have my own set of doubts, we've been fooled before....)
- jpadkins 2y agoto be more specific. People with power desire to retain that power. They will work with other people with power to keep that power if it's in their mutual interest. The people change over time, but the basic psychological need and human behaviors are pretty constant. The methods to stay in power tend to evolve, but they match the same patterns throughout history (e.g. Divide and Conquer). That's it. That's the big conspiracy. Some people like to control others.
- ryandrake 2y ago
- dexwiz 2y agoDo you mean unbiased or not biased towards forces in power? Everything will have some bias to it, will it not? Even if the model does not, surely the training material.
- AlexandrB 2y ago> Unbiased AI is, I believe, an existential threat to the “powers that be” retaining control of the narrative, and must be avoided at all costs. I remember when the internet was supposed to be an existential threat to the "powers that be". I'm pretty skeptical of narratives like this because the "powers that be" have a lot of resources to leverage any new technology for their benefit. At best a new technology is gives an asymmetrical advantage to small actors for a short time before everyone else catches on.
- mdgrech23 2y agoMoney talks and if you don't like money they can just throw you out of a window or label you a terrorist so you never had any real power. Once they flex you've got nothing.
- talldayo 2y ago"unbiased" in LLM terms means just random token selection. You inherently need bias (otherwise called "training data") to inform the placement and weight of each token. Furthermore, unbiased AI isn't likely to be any more usable than the garbage we have today. People care about hallucinations, model latency, token pricing and other practical improvements that can be made. Biases are one of the last things stopping people from using AI for legitimate purposes; the other issues are far too glaring to ignore.
- impossiblefork 2y agoYes, but we understand he what means. We even have alternative meanings for bias within ML, such as for the bias added before non-linearities in many neural networks. He obviously means censored LLMs, and I think his view is actually right, although I'm far from sure that these firms are in some kind of scheme to produce LLMs biased in this sense. Uncensored, tunable LLMs under the full control of their users could scour the internet for propaganda, look for connections between people and organisations and just generally make the work of propagandists who don't have their reader's interests in mind more difficult. I think we'll end up with that anyway but it's a reasonable fear that there'd be people trying to prevent us from getting there.
- contagiousflow 2y agoThere is literally no such thing as an unbiased text generator. No matter how you cut it there are an infinite pool of prompts that will need some sort of implicit value system at the heart of the answer. Any implicit bias just from selection of training data will be reflected back at the user. > Uncensored, tunable LLMs under the full control of their users could scour the internet for propaganda, look for connections between people and organisations and just generally make the work of propagandists who don't have their reader's interests in mind more difficult. Even this example, what sources do you trust that is or is not "in the readers best interest", what is propaganda or what is an implicit value in a society, when you tune an LLM does that just mean you're steering it to give answers that you like more? Creating an unbiased LLM is as much of a fools errand as creating an unbiased news publication
- pixelready 2y agoWhile I take your point, I don’t think we’re quite at the enshittification phase of LLM products yet, where some combination of crass marketing and Orwellian narrative control are baked into to the models for profit and/or to curry favor with govt. Right now we are at a phase where the platform companies are still concerned about finding and selling basic use cases before the hype bubble bursts. There is still a strong possibility that LLM as a product is essentially stillborn in the market writ large, because we are trying to use these tools where end users expect a perfectly accurate, repeatable, deterministic response and there is no guarantee that any of these current techniques, cool as they are, will cross the threshold of user expectations and utility enough to justify the costs. Neither AI doomers nor boosters are accurate representations of the general public.