5 ms·
> it's always going to be biased towards some direction, whether that's the views of the society it pulls most of its data > I don't think there's such thing a
by wizeman 4y ago
> it's always going to be biased towards some direction, whether that's the views of the society it pulls most of its data
> I don't think there's such thing as a lack of bias
If the AI is simply reflecting the data it was trained on and this data is a representative sample of all data, isn't it unbiased by definition?
I don't think we should just throw our hands up and say "this is impossible" just yet.
That's just a convenient excuse for OpenAI (or others like them) to get away with what effectively is censorship of certain ideas or political views.
- karpierz 4y ago> If the AI is simply reflecting the data it was trained on and this data is a representative sample of all data, isn't it unbiased by definition? It's unbiased by definition of "does the output reflect the input"? It's not unbiased by definition of "does the output reflect reality"?
- wizeman 4y ago> It's not unbiased by definition of "does the output reflect reality"? How does "all data" differ from reality?
- scarmig 4y agoOnly a miniscule part of reality is digitized, and what data does exist passed through the biases of people before being available to train on.
- wizeman 4y agoIf that is a concern, then perhaps you could go ahead and sample a tiny part of "reality" (whatever that means) and then adjust the weights of the digitized data so that it becomes a more representative sample. Also, being biased or unbiased is not dichotomic, i.e. it's not all or nothing. It's something that you can work towards if you put an effort into it. Basically what I'm saying is: don't just go around saying that the task is impossible. At least, try to make an effort to be unbiased and to improve on that over time, and don't just say "it's impossible" as an excuse for being biased.
- Balgair 4y agoWoah, I mean, this argument (the last few comments here) has been a central one in 'western' philosophy for at least the the last 2400 years, if not the last ~4000. I'm not a philosopher by any means, so I'm unaware of the current state of the great conversation. But as to whether reality is even knowable is still very much up for debate, I believe (please correct me philosophy peepz!). In physics we're still woefully unaware of what ~70% of the universe's stuff is doing (negative energy) and if it effects us at all. In neuroscience we still debate what % of your brain neurons make up vs. things like glia. Etc. Like, even trying to capture 'reality' with our quite primitive eyes and sensors and optical engineering is really really hard to do (Abbe' diffraction limit, entropy, Lens maker's equation, etc)
- wizeman 4y agoFortunately, I think "reality" in this context doesn't have the same meaning as "the physical universe". I think the important goal is for as many people as possible to feel like the AI isn't being too biased against them, while still not crippling the AI too much. I will leave the exact mathematical formula for that measure (along with the methods for gathering that input) for debate among researchers who know more about that than I do.
- entropicdrifter 4y agoI mean, define "biased against", right? Because some people will argue that anything that's not explicitly in full agreement with them is "biased against" them. It's a narcissistic, dishonest take, but plenty of people do take that stance nonetheless in order to try to shift all arguments into their narrow worldviews/definitions in order to "win" as many conversations as they can. Right? I mean I've met people online and off who do this from almost every part of the political spectrum. So, do we filter those people out from consideration to begin with, or do we have to cater to those with extreme views in order to get as many "not biased against me" ratings as possible? I guess the point I'm trying to make is that trying to optimize for any single metric is a fool's errand because as soon as you do so, it will be gamed/exploited. Then you can either try diversifying your optimization data points (who gets to choose those? How could they possibly be unbiased, when they literally define the system's bias?) or you can try filtering out bad actors from the data, which is very directly an attempt to bias the system away from insincere bad actors. And all of that's not even accounting for the lack of incentive to try to find neutrality when more biased views are more lucrative in the attention economy.
- rurp 4y agoIt's not using all of the atoms in the universe as training data... Any collection of human writing is going to contain objectively wrong assertions, and those errors will vary based on the time and place the training data was sourced from.
- wizeman 4y agoSure but I mean, if a conversational AI would only be allowed to spit out mathematically correct statements, it would be extremely limited (and boring). I think what's important is for those mistakes to be evenly distributed among as many axis(s) as possible, and especially, not bias them towards one side of political thought.
- emodendroket 4y agoBut that will not be achieved by just shoveling in writing indiscriminately.
- luisapic 4y ago[dead]
- sangnoir 4y agohow many sides do you suppose political thought has? Assuming there is more than one, how do you find the geometric center to avoid bias? If your training data is English text, will the AI be biased against early German philosophers and French political theorists? Of by "unbiased", you are simply referring to the American left-right axis? If so, do you mean the fiscal axis, or the social one?
- wizeman 4y agoI mean that the likelihood (or weigh) of an opinion being expressed by an AI should be roughly proportional to the number of people who currently hold that opinion, assuming the AI is simply generating responses based on its training (which is what should actually be as unbiased as possible). As an example, let's suppose that 55% of people believe that it's not OK to make jokes about women, but it's OK to make jokes about men, and that roughly 40% believe it's OK to make jokes about both (I'm not saying this is the case, it's just an example). So perhaps, in this case, by default the AI wouldn't make a joke about women. But if you would slightly nudge it or insist a bit more, perhaps the AI wouldn't refuse to make a joke about women anymore, because there is still a large proportion of the population who do believe that's perfectly OK (of course, then we might get into the territory about overtly sexist jokes, which obviously the AI would have to refuse a lot more than making a more innocent joke about women). Now let's say we start asking the AI to make Nazi comments. Obviously, the segment of the population who agrees with Nazi sentiment is a lot smaller, and the anti-Nazi sentiment is a lot stronger, so the AI should have to object to such a request quite more strongly. This type of refusal or likelihood of the AI saying something should presumably be roughly proportional to the opinions and sentiment of the general population (or at the very least, the target market for the AI), not just the OpenAI employees who performed the RLHF to train the AI in terms of acceptable responses and who are much more likely to be biased. I'm not saying that this is necessarily easy to accomplish, there are certainly difficulties here. As an example, some widely-held opinions, even about objective things, may not necessarily be rooted in facts, so some kind of balancing might be necessary (a general kind of balancing, not a "let's dissect and nudge the AI responses on an opinion-by-opinion basis"). And yes, I understand that this can be quite difficult, because any given source of truth can be perceived to be biased by some segment of the population. What I am saying, however, is that AI creators such as OpenAI should be making more efforts in this direction. To start with, perhaps the RLHF training should be done with AI trainers selected from a more representative sample of the population. And yes, we may never be able to accomplish 0% bias, but we should at least make some effort to reduce it. It's also interesting to me that at some point, the AI may start to express opinions that are not a strict "linear" function of the data it was trained on, and yes, this might piss off a significant amount of people. In my opinion, this would be quite interesting and should be OK, as long as we made reasonable efforts to remove sources of bias from its training process. Although I can also see an important target market (perhaps even larger) for an AI that is more biased to generate responses according to the beliefs of the general population, rather than what it "perceives" to be more true.
- jojobas 4y agoIf the input data was perfectly self-consistent, "all data" could be considered "reality". In reality, "all data" is rife with disagreement, which you have to perceive as noise (and get noisy output) or value-judge contradicting opinions, getting, no surprise, biased output.
- coldtea 4y agoImagine a guy and two of his friends comes, calls you a bad name, has his friends hold you and beats you up. You try to resist, and when you have the chance, run away from them. The story published in newspapers about the incident is "wizeman attacks a group of nice young men minding their own business, steals their wallet". "All data" != reality...
- feet 4y agoThat is not "all data"
- coldtea 4y agoThere is no other data in the example. That's the whole point. If you mean "but in the real world will have way more stories from other sources about other things" sure. But doesn't change anything if you have "all stories printed". The distrubution matters. All or most of them can very well be biased and not reflect reality. And that's for factual matters. Let's not even go into political matters. Like in 1920s South most newspaper stories would be biased in favor of Jim Crow, few would be against it.
- pcstl 4y agoI don't think you can say an AI trained using RLHF - such as ChatGPT is - is really "simply reflecting the data it was trained on". ChatGPT was first trained on a load of data, then it was updated to act in specific ways based on feedback from humans who "nudged" it the way they wanted it to go.
- wizeman 4y agoAre those humans that nudged it representative of the population? Or were they mostly "woke" Silicon Valley employees? (not to dismiss woke Silicon Valley employees, I'm just saying their opinions are not representative of the entire population).
- dorchadas 4y agoThere's also bias in the data itself. That's the difficult thing to avoid. Even down to how we phrase a question, who we collect the data from, it all introduces a bias unless we're literally harvesting all data from every human being and using that for our models. There's no way to get rid of the bias, even if we take out the nudges.
- wizeman 4y agoHow about you select a representative (i.e. random and statistically significant) sample of the population and then ask them their opinions about certain (especially controversial) parts of your data, and then weigh your data according to these opinions? That's just an idea that occurred to me (in 30 seconds of thought) which could probably make the training data significantly more unbiased. But I'm sure there are research scientists who can come up with better methods for sampling data in a more unbiased fashion. Note that this is not an all or nothing approach. Your training data could presumably be 100% biased or 0% biased, but also any value in-between. The goal is to try to make it as close to 0% biased as feasible, given whatever effort you're comfortable expending.
- freejazz 4y ago
- dragonwriter 4y ago> If the AI is simply reflecting the data it was trained on and this data is a representative sample of all data, isn’t it unbiased by definition? No, “data” is just information which has been gathered. “All data” can be biased. Also, data can itself be bias, even if it isn’t biased. For instance, a text generation model that was based on unbiased collection of all text ever written by humans would, in one sense, produce “unbiased, human-representative text”. It would also reproduce the biases of the authors, weighted by the volume of writing coming from that bias. > That’s just a convenient excuse for OpenAI (or others like them) to get away with what effectively is censorship of certain ideas or political views. While one might object to the editorial choices, I can’t see any rational bounds for objecting to the idea that the creator of models would censor “certain ideas or political views” as a generality.
- wizeman 4y ago> It would also reproduce the biases of the authors, weighted by the volume of writing coming from that bias. Yes, but I think there are ways we could reduce this bias, perhaps significantly, even. > While one might object to the editorial choices, I can’t see any rational bounds for objecting to the idea that the creator of models would censor “certain ideas or political views” as a generality. You are right, I was unfair with my words. I think it would be more fair to say that OpenAI is inadvertently biasing ChatGPT answers as a side effect of their RLHF training being done using answers/rankings done by people (i.e. the AI trainers [1]) who are not a representative sample of the population, but rather, probably comprise a group of people who are likely to be significantly more leaning to one side of the political discourse (presumably, OpenAI employees or Silicon Valley-based contractors?). This probably greatly biases ChatGPT to produce certain kinds of answers to certain kinds of questions that would likely not happen otherwise, and in fact, these answers are perceived to be quite biased by the other side of the political discourse. [1] https://openai.com/blog/chatgpt/ https://openai.com/blog/chatgpt/
- smeagull 4y ago> That's just a convenient excuse for OpenAI (or others like them) to get away with what effectively is censorship of certain ideas or political views. I can't tell if you're on the conservatives side or OpenAI's. Are conservatives being censored because OpenAI are allowing "woke" training data to be represented, or are conservatives asking OpenAI to censor the "woke" political views?
- dzikimarian 4y ago>I can't tell if you're on the conservatives side or OpenAI's. Not OP, but maybe none? It's possible to have opinion, without aligning with either side of polarized discussion.