4 ms·
That is a very poorly formed study precisely because systems like the Mechanical Turk actually work surprisingly well. This is a common theme in the 'wisdom of
by indubitable 9y ago
That is a very poorly formed study precisely because systems like the Mechanical Turk actually work surprisingly well. This is a common theme in the 'wisdom of the masses.' If you ask a single person how many beans are in a jar, you're going to get an answer that's generally very wrong. Yet ask 100 people and average their answer and you tend to get an answer that's oddly extremely close to the correct answer. As the number of people independently asked approaches infinity, the error approaches 0.
So when you ask 'x' people to independently judge something, let alone something that is multiple choice with a correct answer, you're going to get answers that are far more accurate than a single individual would give you. So looking at the average answer of 'x' people, comparing it to an AI system, and arguing that they're relatively close so individual people are relatively close to the AI is completely fallacious.
---
As a tangential aside, there's another quirk to the wisdom of the masses. When you let the people communicate and try to intelligently organize and use expertise to come to answer, this effect disappears and the final answer again tends to be very wrong. Kind of an interesting perspective on the current zeitgeist of society and work.
===EDIT===
The authors were obviously aware of the wisdom of the masses. Quoting the paper itself:
To determine whether there is “wisdom in the crowd” (7) (in our case, a small crowd of 20 per subset), participant responses were pooled within each subset using a majority rules criterion. This crowd-based approach yields a prediction accuracy of 67.0%. A one-sided t test reveals that COMPAS is not significantly better than the crowd (P = 0.85).
That's a quite silly p-value and on top of that I'm not sure how they claim their system actually controls for the wisdom of the masses.
- IshKebab 9y ago> As the number of people independently asked approaches infinity, the error approaches 0. Only if people are an unbiased estimator!!
- indubitable 9y agoYou'd think so, but that's not correct. This is not a straight forward result of probability with a filter of complexity. It works regardless of bias, though obviously if everybody was biased in the exact same way then that would cause things to break down - which is perhaps the reason that the coordination results in a worse result than independent averages. For another example of it consider things like the television show, 'Who wants to be a millionaire?' It's a quiz show where one of the choices is for the participant to ask the audience. And the audience tends to do absurdly well on even the most esoteric questions, though independently they are certainly far from trivia experts. But very few are randomly guessing - their own experiences and biases leads them to entirely different conclusions. Yet somehow it produces the correct result time and again. It's a strange phenomena and one that has to be constantly guarded against in anything involving sampling of people. This is a text book example of a study that gets destroyed by it.
- _dps 9y ago>>> As the number of people independently asked approaches infinity, the error approaches 0. >> Only if people are an unbiased estimator!! > You'd think so, but that's not correct. Either you're misinterpreting the technical term "unbiased estimator" here, or you are aware of some research that I would like to read. In context, "unbiased" means that if you pick people at random and ask for their estimates, then on average the too-high estimates cancel out the too-low estimates (i.e. there is not a bias in one direction or another). But people as a whole have poor understanding of many things. One common one that appears in social science research and is often replicated is that people grossly overestimate the size of the homosexual population in the US (the "wisdom of the crowds" often estimates it around 20% whereas best available polling data suggests 3-5%). Here's just one source for this phenomenon http://news.gallup.com/poll/183383/americans-greatly-overestimate-percent-gay-lesbian.aspx http://news.gallup.com/poll/183383/americans-greatly-overest... "Wisdom of the crowds" is occasionally reliable, but it should not be assumed to be reliable for any particular problem without verification. It often fails terribly even on problems that are not very esoteric. Edit: changed phrasing of final paragraph
- indubitable 9y agoAs mentioned, there is a difference between bias and uniform bias. In the US the media, politics, social media, and even miseducation (e.g. in my deviance class we focused on Kinsey's 10%, yet oddly enough never contrasted that against contemporary results) have heavily and uniformly biased the population on sexuality leading people to vastly overestimate the number of homo/bi/trans individuals. Where it works phenomenally well is in areas where biases have not been directly instilled into people. This does not mean people are unbiased, however. Again the knowledge of trivia is a good example since while the crowds can generally do phenomenally well even at very esoteric questions where biases would lead them to individually come to very different conclusions, yet they will invariably fail to answer ostensibly trivial questions like 'What is the capital of Australia?' You'll get Sydney, it's not. You'd likely get a similar result for things like the capital of Pennsylvania. In a way I view the wisdom of the masses as analogous to machine learning systems. They do an oddly good job of providing extremely precise answers to a wide array of questions even when trained with models that do not directly represent the 'questions'. Yet you can also break the systems, at times comically, with certain types of queries designed to do precisely that. And as was the case with machine learning for quite some time, I think people remain reluctant to utilize it due to the black box nature of it. The implication of your comment is that the wisdom of the masses is little more than incorrect answers canceling out on average leaving nothing but a survey of experts. Yet I think there's no evidence for this (even if it may be a perfectly logical 'kneejerk' reaction) as it works even on things where nobody is an expert, and if this were the case then we ostensibly should be able to get comparable answers from coordination - yet coordination causes the entire system to collapse.