5 ms·
> It's not unbiased by definition of "does the output reflect reality"? How does "all data" differ from reality?
by wizeman 4y ago
> It's not unbiased by definition of "does the output reflect reality"?
How does "all data" differ from reality?
- scarmig 4y agoOnly a miniscule part of reality is digitized, and what data does exist passed through the biases of people before being available to train on.
- wizeman 4y agoIf that is a concern, then perhaps you could go ahead and sample a tiny part of "reality" (whatever that means) and then adjust the weights of the digitized data so that it becomes a more representative sample. Also, being biased or unbiased is not dichotomic, i.e. it's not all or nothing. It's something that you can work towards if you put an effort into it. Basically what I'm saying is: don't just go around saying that the task is impossible. At least, try to make an effort to be unbiased and to improve on that over time, and don't just say "it's impossible" as an excuse for being biased.
- Balgair 4y agoWoah, I mean, this argument (the last few comments here) has been a central one in 'western' philosophy for at least the the last 2400 years, if not the last ~4000. I'm not a philosopher by any means, so I'm unaware of the current state of the great conversation. But as to whether reality is even knowable is still very much up for debate, I believe (please correct me philosophy peepz!). In physics we're still woefully unaware of what ~70% of the universe's stuff is doing (negative energy) and if it effects us at all. In neuroscience we still debate what % of your brain neurons make up vs. things like glia. Etc. Like, even trying to capture 'reality' with our quite primitive eyes and sensors and optical engineering is really really hard to do (Abbe' diffraction limit, entropy, Lens maker's equation, etc)
- wizeman 4y agoFortunately, I think "reality" in this context doesn't have the same meaning as "the physical universe". I think the important goal is for as many people as possible to feel like the AI isn't being too biased against them, while still not crippling the AI too much. I will leave the exact mathematical formula for that measure (along with the methods for gathering that input) for debate among researchers who know more about that than I do.
- entropicdrifter 4y agoI mean, define "biased against", right? Because some people will argue that anything that's not explicitly in full agreement with them is "biased against" them. It's a narcissistic, dishonest take, but plenty of people do take that stance nonetheless in order to try to shift all arguments into their narrow worldviews/definitions in order to "win" as many conversations as they can. Right? I mean I've met people online and off who do this from almost every part of the political spectrum. So, do we filter those people out from consideration to begin with, or do we have to cater to those with extreme views in order to get as many "not biased against me" ratings as possible? I guess the point I'm trying to make is that trying to optimize for any single metric is a fool's errand because as soon as you do so, it will be gamed/exploited. Then you can either try diversifying your optimization data points (who gets to choose those? How could they possibly be unbiased, when they literally define the system's bias?) or you can try filtering out bad actors from the data, which is very directly an attempt to bias the system away from insincere bad actors. And all of that's not even accounting for the lack of incentive to try to find neutrality when more biased views are more lucrative in the attention economy.
- wizeman 4y agoI think all of your points are valid. But I still think we should make an effort and strive to solve these problems. I don't think this is being done with ChatGPT, for example. But also, note that an AI doesn't have to be in complete agreement with someone for that person to not feel "biased against". As long as an AI does make some effort to not be prejudiced/biased, that could work. For example, if someone asks: "is climate change real"? An AI does not have to give a simple yes/no answer, or represent a single viewpoint. It could give an answer that is mostly representative of the major thought streams. For example, it could answer something like: "The vast majority of scientists/governments/people have reached the conclusion that climate change is real, bla bla bla. [Here's some good, convincing evidence]. That said, there is a minor fraction of scientists/government/people who believe that climate change is not caused by human action. [They criticize the above evidence in this way]. [Here's also some counter-evidence]. That said, many scientists believe these studies are flawed for this reason or another." I mean, sure, there is still going to be a lot of people who don't agree with this answer. But I think, on a scale of 0-10 they would agree a lot more with this answer than one that completely ignores their viewpoints. And even for those of us who believe in climate change, we can still consider this answer somewhat reasonable. Thus, increasing the total amount of points would probably be a somewhat effective way of eliminating a large deal of bias, I think. Although, yes, you couldn't do this for every possible viewpoint. And it would be a challenge to figure out how to weigh these points in a way that makes the most amount of people happy. But I still think we should make efforts in this direction.
- coldtea 4y ago>If that is a concern, then perhaps you could go ahead and sample a tiny part of "reality" (whatever that means) and then adjust the weights of the digitized data so that it becomes a more representative sample. Who is doing the "adjusting the weights"? Why would they be "unbiased"? The real answer is those who make the AI (or people who have power over them) get to chose the training data or to adjust the weights. And the rest have to put up with it, whether the former are biased or not.
- rurp 4y agoIt's not using all of the atoms in the universe as training data... Any collection of human writing is going to contain objectively wrong assertions, and those errors will vary based on the time and place the training data was sourced from.
- wizeman 4y agoSure but I mean, if a conversational AI would only be allowed to spit out mathematically correct statements, it would be extremely limited (and boring). I think what's important is for those mistakes to be evenly distributed among as many axis(s) as possible, and especially, not bias them towards one side of political thought.
- emodendroket 4y agoBut that will not be achieved by just shoveling in writing indiscriminately.
- luisapic 4y ago[dead]
- sangnoir 4y agohow many sides do you suppose political thought has? Assuming there is more than one, how do you find the geometric center to avoid bias? If your training data is English text, will the AI be biased against early German philosophers and French political theorists? Of by "unbiased", you are simply referring to the American left-right axis? If so, do you mean the fiscal axis, or the social one?
- wizeman 4y agoI mean that the likelihood (or weigh) of an opinion being expressed by an AI should be roughly proportional to the number of people who currently hold that opinion, assuming the AI is simply generating responses based on its training (which is what should actually be as unbiased as possible). As an example, let's suppose that 55% of people believe that it's not OK to make jokes about women, but it's OK to make jokes about men, and that roughly 40% believe it's OK to make jokes about both (I'm not saying this is the case, it's just an example). So perhaps, in this case, by default the AI wouldn't make a joke about women. But if you would slightly nudge it or insist a bit more, perhaps the AI wouldn't refuse to make a joke about women anymore, because there is still a large proportion of the population who do believe that's perfectly OK (of course, then we might get into the territory about overtly sexist jokes, which obviously the AI would have to refuse a lot more than making a more innocent joke about women). Now let's say we start asking the AI to make Nazi comments. Obviously, the segment of the population who agrees with Nazi sentiment is a lot smaller, and the anti-Nazi sentiment is a lot stronger, so the AI should have to object to such a request quite more strongly. This type of refusal or likelihood of the AI saying something should presumably be roughly proportional to the opinions and sentiment of the general population (or at the very least, the target market for the AI), not just the OpenAI employees who performed the RLHF to train the AI in terms of acceptable responses and who are much more likely to be biased. I'm not saying that this is necessarily easy to accomplish, there are certainly difficulties here. As an example, some widely-held opinions, even about objective things, may not necessarily be rooted in facts, so some kind of balancing might be necessary (a general kind of balancing, not a "let's dissect and nudge the AI responses on an opinion-by-opinion basis"). And yes, I understand that this can be quite difficult, because any given source of truth can be perceived to be biased by some segment of the population. What I am saying, however, is that AI creators such as OpenAI should be making more efforts in this direction. To start with, perhaps the RLHF training should be done with AI trainers selected from a more representative sample of the population. And yes, we may never be able to accomplish 0% bias, but we should at least make some effort to reduce it. It's also interesting to me that at some point, the AI may start to express opinions that are not a strict "linear" function of the data it was trained on, and yes, this might piss off a significant amount of people. In my opinion, this would be quite interesting and should be OK, as long as we made reasonable efforts to remove sources of bias from its training process. Although I can also see an important target market (perhaps even larger) for an AI that is more biased to generate responses according to the beliefs of the general population, rather than what it "perceives" to be more true.
- jojobas 4y agoIf the input data was perfectly self-consistent, "all data" could be considered "reality". In reality, "all data" is rife with disagreement, which you have to perceive as noise (and get noisy output) or value-judge contradicting opinions, getting, no surprise, biased output.
- coldtea 4y agoImagine a guy and two of his friends comes, calls you a bad name, has his friends hold you and beats you up. You try to resist, and when you have the chance, run away from them. The story published in newspapers about the incident is "wizeman attacks a group of nice young men minding their own business, steals their wallet". "All data" != reality...
- feet 4y agoThat is not "all data"
- coldtea 4y agoThere is no other data in the example. That's the whole point. If you mean "but in the real world will have way more stories from other sources about other things" sure. But doesn't change anything if you have "all stories printed". The distrubution matters. All or most of them can very well be biased and not reflect reality. And that's for factual matters. Let's not even go into political matters. Like in 1920s South most newspaper stories would be biased in favor of Jim Crow, few would be against it.