5 ms·
> I think this sort of veiled personal attack resorting to baseless extrapolation is not a productive line of public discourse. I thought that my comment was f
by einarfd 6y ago
> I think this sort of veiled personal attack resorting to baseless extrapolation is not a productive line of public discourse.
I thought that my comment was fitting to the tone of the text I replied too, but fair enough, maybe I shouldn't have included that paragraph.
> You should take a step back and think about a) the field of study, b) how the main argument completely ignores the basis of said technical field and what it studies, c) how the argument being created lies on the idea that a self-annointed elite should have the right to manipulate the public to force it to fall in place with it's goals and desires.
It is unclear to me what this means, is it arguing that studying bias in AI and specifically deep learning is not germane?
Is it arguing that it is not acceptable to institute policies and action to create a level playing field for minorities and repressed group because that would be social engineering, and forcing the common man to do the elites bidding?
If that is what it means, I have to disagree, I find both of those things worthwhile.
- rualca 6y ago> I thought that my comment was fitting to the tone of the text (...) Yes, that was one of the problems. > It is unclear to me what this means, is it arguing that studying bias in AI and specifically deep learning is not germane? Let me make it clear for you so that a) we are able to talk about things objectively, b) your options to continue using veiled personal attacks is curtailed. Either your goal is to model reality and real life, or your goal is to model your idealization of what you feel real life should be. If you pick option #2 then your model does not reflect real life. Bias is by definition the way the model returns results that don't match the real world and real life, in frequency and in proportion. If your goal is to use your model to manipulate and control society based on your own personal criteria, by manipulating it to return results that distort the real world and real life, then call it something else, because bias is not it.
- gnomewascool 6y agoI'm not OP, but what is "real life"? In cases where some communities/countries have (to quote from the article) a "smaller linguistic footprint online", is trying to control for this a form of manipulation of society or a way to fix a bias? Can you actually choose a sampling measure without making some sort of value judgement?
- rualca 6y ago> I'm not OP, but what is "real life"? In the context of creating models, it's representative data collected from statistically significant observations of a population, for starters. > In cases where some communities/countries have (to quote from the article) a "smaller linguistic footprint online", is trying to control for this a form of manipulation of society or a way to fix a bias? In modelling there is no such thing as control. There's the input data and there's the model generated from input data. If you are looking for a model that is expected to represent a property intrinsic to a specific community then you use data collected from that community to generate that model. That's it. Models generalize, and reflect the norm. They are like that by design. That's their point. If your plan is to have a model that does not reflect the input data but instead forces your biases regardless of the input data then your goal is not to model reality but to distort it to comply with your personal goals.
- einarfd 6y ago> Either your goal is to model reality and real life, or your goal is to model your idealization of what you feel real life should be. > If you pick option #2 then your model does not reflect real life. These deep learning models built by corporations are not scientific models, they are engineering solutions, built to solve problems. Reflecting the real world is only useful if it furthers what the company wants to solve. If they for example remove swear words from their training set, that will make them a less accurate model of the world, but make them more useful for building solutions. But it is probably a trade of they would be happy with. We've also seen example of risk scoring application for felons that seem to end up doing racial profiling, because that is what the data seem to indicate makes sense. But that's deeply problematic and runs counter to laws in some places and seem ethically problematic (https://www.theverge.com/2020/6/24/21301465/ai-machine-learning-racist-crime-prediction-coalition-critical-technology-springer-study https://www.theverge.com/2020/6/24/21301465/ai-machine-learn...). > Bias is by definition the way the model returns results that don't match the real world and real life, in frequency and in proportion. Getting a good data set without bias is hard, even if you crawl the whole internet like Google does. Not everything is on the internet and there are systematic drivers that make some part of the human condition over represented (English, science, the views of the affluent and educated), and some under presented (small languages, the discourse of people behind the Chinese great firewall, the poor). So just getting a ginormous data set does not fix bias. > If your goal is to use your model to manipulate and control society based on your own personal criteria, by manipulating it to return results that distort the real world and real life, then call it something else, because bias is not it. Positive bias is absolutely something that we use, and while it might seem sinister it does not have to be. The example I'm most familiar with a facial recognition technology. Most groups building that ends up with a model that is better at some groups than others. Asian research groups often end up with models that does well with asians and worse with whites, while European groups usually end up with the reverse. In some sense these results do reflect the reality of these groups, most people in Europe is white and most people Asian is asians, so that you training sets ends up like that is not surprising. But no one is happy with these kind of results, and everyone wants to fix that. To bring it back to speech and text models, let's say you are building a customer service solutions incorporating a deep learning model, the reality might be that you current customer service representatives treat blacks (or people who use "black" dialects), worse than people who sound white. An accurate model built on this data set will then also do that. But is that acceptable? I hope most companies would want to fix that, and be fine with adding some positive bias in their solution.