5 ms·
Forgive my ignorance since I don't have a CS degree, but don't all algorithms used for selecting have a human behind the logic. So wouldn't an algorithm always
by Robelius 9y ago
Forgive my ignorance since I don't have a CS degree, but don't all algorithms used for selecting have a human behind the logic. So wouldn't an algorithm always have some bias based on the weights given by a human. Kind of like OKCupid saying you are a 75% match with someone. You may not be compatible with that person, but whoever designed the algorithm declared it to be true with their design.
So wouldn't any algorithm always have some biased in it?
Sorry if this comes off as a stupid question or is unclear.
- untog 9y agoIt's not a stupid question at all, and it cuts to the core of a lot of problems with AI and algorithmic-based anything. Any sufficiently complex programming project will end up reflecting some assumptions and biases on the behalf of the programmer. As more programmers contribute to a project, it won't reflect an individual's biases, but will still reflect those programmers as a group.
- gravypod 9y ago> Any sufficiently complex programming project will end up reflecting some assumptions and biases on the behalf of the programmer. As more programmers contribute to a project, it won't reflect an individual's biases, but will still reflect those programmers as a group. Can you break this down more? It doesn't make sense to me. If you're writing a machine learning application to take a dataset and match future inputs to past results I don't see how these biases can sneak into the program. Unless the programmers are changing the datasets then I don't see how this makes sense.
- theemathas 9y agoWell... why can't the dataset be biased?
- amitdeshwar 9y agoThat's not a bias of the programmer
- chongli 9y agoUnless the programmer chose the dataset.
- woodruffw 9y agoIn the case you presented, the biases are in the dataset itself. Arrest statistics in the US are heavily skewed by race. If you were to take a dataset of all US arrests between 1900 and 2000 and ask which populations are most likely to "commit crimes" (i.e., be arrested), you'd get racially biased model without recording why it's biased (discrimination, enduring poverty, minimum sentencing, &c).
- dante821 9y agoCan't that problem be solved by simply not telling the machine the race of the person who's being judged?
- mattnewton 9y agoDead sibling comment asks a common questions I hear: why can't we just not tell the algorithm race? And the answer is that, with enough data points, if there is a statistical bias in the dataset, the algorithm will likely learn a substitute for race in the other features. For the sake of argument, perhaps an interaction between home address, annual income and model of car driven is predictive of race, and race is predictive of recidivism in the dataset- then the algorithm will learn this cross of features is predictive of recividism, even though we would like all of them to ideally be irrelevant to sentencing.
- gravypod 9y ago> home address, annual income and model of car driven is predictive Why the hell would those be factors? The factors of the case should be things that actually matter in the case. A rich person who murders someone should get the same sentence as a poor person who murders someone. A black person who murders someone should get the same sentence as a white person who murders someone. Adding those factors would be insane in the first place. If you're adding crazy things like that you might as well add factors like "Can Juggle" and "Can Burp the Alphabet" because things like that should have just as much to do with sentencing as what kind of car you drive or where you live.
- 9y ago
- dogecoinbase 9y agoHow was the dataset produced?
- deleted 9y ago[deleted]
- randomtask 9y agoAs with anything, it's hard to generalise without oversimplifying, but here goes. You don't generally just have a data set and a machine learning algorithm that somehow magics outputs from a data set. Usually decisions have to be made by people either in training the model, selecting variables that are included in a model, etc. Here's a simple example. Say you're trying to come up with an algorithm that decides whether articles in a data set are "fake news" (topical, I know). We have to tell the algorithm whether a given article in the training set is fake or legitimate, otherwise how would it know? Clearly this will reflect the views of whoever is tagging the articles. When we run the model on a training set we need to score how well it did, again this will reflect the opinion of the person doing the scoring. For a real example: https://mathbabe.org/2016/05/12/algorithms-are-as-biased-as-.. https://mathbabe.org/2016/05/12/algorithms-are-as-biased-as-....
- gravypod 9y agoI've done some machine learning in the past. I get how you tackle machine learning problems and the first thing that I would say that your proposed topic is not currently possible. "*Say you're trying to come up with an algorithm that decides whether articles in a data set are "fake news" (topical, I know).*" Current AI cannot do this. This would take finding sources, pulling data out of those sources, cross referencing multiple sources, and recurring for those articles to a certain depth. That's not a good facsimile for deciding sentencing. Sentencing is more like a linear regression classification. You have a history of previous cases where the defendant was found guilty. You then have a pile of factors that played into the judge's decision for sentencing. For example: * If they meant to do it * If they feel bad about doing it * If they did do it (Beyond a reasonable doubt) * If they have done it before * What severity this crime is * ... etc The judge then uses their experience in law and previous case law as well as statues to find a proper punishment. This is in the form of: * Time served * Fines * Privileges revoked * Community Service This would then be fed into a classification engine. You leave all of the existing infrastructure in place (Judge, Jury, Lawers) and just use their decision as input into the sentencing. Deciding the validity of claims is not within the scope of modern day machine learning (as of 2017). Classification engines are very much in the scope of machine learning of today. I don't see how case factors could be biased. I don't see how historical cases (when stripped of all identifying information) could be biased. I don't see why a system like this would be bad. All treatment of everyone would converge into a uniform handling of cases.
- zitterbewegung 9y agoNot only the algorithm can be biased but also the underlying dataset could be biased. See http://www.sciencemag.org/news/2017/04/even-artificial-intelligence-can-acquire-biases-against-race-and-gender http://www.sciencemag.org/news/2017/04/even-artificial-intel... And the paper http://opus.bath.ac.uk/55288/ http://opus.bath.ac.uk/55288/
- Hydraulix989 9y agoI can see how the dataset can be biased, but I'm struggling to find any biases in the algorithms themselves. Assuming I'm using some actual machine learning model and not my own hand-coded finite state machine, I just don't see how, say, k-means clustering could be biased (unless there was a bug in the programmer's implementation?).
- hackuser 9y agoIt will be intentionally biased; that I'm confident of. https://news.ycombinator.com/item?id=14288232 https://news.ycombinator.com/item?id=14288232
- Jtsummers 9y agoThere are two algorithms. Machine learning generates models based off of input data. Machine learning, principally, will not be [biased or] exhibit any direct bias. [It's math, on its own it's not the problem.] The generated model is then incorporated into an algorithm for making decisions or categorizing objects. If the original data contains biases, the final model will be biased. [EDIT: Forgot two words, added a sentence]
- Hydraulix989 9y ago"There are two algorithms." In the case of machine learning techniques that employ a learning algorithm (e.g. SGD) and a separate classification algorithm (e.g. forward prop), the situation is no different. Gradient descent in itself doesn't take "unbiased" data and pollute it with "bias" -- neither does, forward propagation. These algorithms are merely functions that are _parametrized_ by a potentially biased learned model (which really entirely amounts to having biased data). I recommend taking a look at Stanford's Intro to Machine Learning CS 229 course notes for a brief overview of these concepts. I can't speak anymore highly of the material in this course; it's by far the best out there, and it's how I learned ML. "The generated model is then incorporated into an algorithm for making decisions or categorizing objects. If the original data contains biases, the final model will be biased." Right, the "final model" (which comes directly from the data) will be biased, but not the ALGORITHM (the "math," as you call it). For example, I can't make an unbiased data set "biased" by using linear regression instead of logistic regression. It would be quite a tremendous discovery if you somehow figured out a way SGD in itself was biased, but the burden of proof on such an extraordinary claim would be on you. So like I said in my parent post, I can see how the data can be biased, but I'm struggling to find any biases in the algorithms themselves (assuming they are properly implemented and not an Underhanded C Contest [1] entry). Granted, if we're not actually using ML, and we're using something like a hardcoded decision tree or FSM instead (it was pretty strongly implied that we weren't though because we were talking about data-driven algorithms), then sure, that could very easily be a biased algorithm, but I've maintained that all along. [1] http://www.underhanded-c.org/ http://www.underhanded-c.org/
- hackuser 9y agoUltimately, everything is politics, and I predict that if these algorithms become widely used, politics eventually will be the primary determining factor in their calculations. Why? If these algorithms become widely used and thus influence important outcomes, then people who like to control power will seek to obtain control of the algorithms, either to protect their interests or to grow their own influence. As possible examples, politicians who see an algorithm apolitically jail a powerful person might pass a law saying that public service should be weighed more heavily. Law enforcement might want exceptions or special weightings for police officer defendants, or for police officer-given testimony. Powerful social activists will buy the algorithm developers (do you know who is writing these algorithms even now? I don't) as a way of influencing society - the Koch brothers, for example, invest in other areas, such as academia and down-ballot elections (e.g., secretary of state in U.S. states) for society-changing purposes. Special code might even generate special outcomes for particular individuals - who will ever see the code? Who will even ask this question after the sentencing? If you think such corruption couldn't happen, look around. Brazen corruption happens all the time, especially in state and local government. This will be harder to detect, and will be covered by the excuse that non-technical people will believe: I didn't do it; it was the algorithm! It must be objective! I think the algorithms ultimately will be a step backward, reducing transparency by hiding the bias in a software development process, in the back room of a software firm, and in a mountain of code. At least legislation is published and voted on in the open.