36 ms·
The Dunning-Kruger Effect Is Autocorrelation
- danbruc 4y agoI am no expert in statistics or the Dunning-Kruger effect but this analysis doesn't sound correct to me. If you plot self assessment against test scores then the following will happen. If people are perfect at self assessment, then you get a straight diagonal line. The more wrong they are, the wider the line will get, in the extreme - if the self assessment is unrelated to the test result - the line will cover the entire chart. If people overestimate their performance, the line will move up, if they underestimate their performance, the line will move down. If you look at the Dunning Kruger chart, that is what you see, complicated a bit by the fact that they aggregated individual data points. At low test scores the self assessment is above the diagonal, at high test scores it is below. What matters is indeed the difference between the self assessment and the ideal diagonal, but if you don't plot individual data points but aggregate them, you have to make sure that there is a useful signal - if self assessments are random, then the median or average in each group will be 0.5 and you will get a horizontal line, but that aggregate 0.5 isn't really telling anything useful.
- geysersam 4y agoI'm not sure what you mean with "the wider the line will get". But here is the issue: The least competent person cannot underestimate their relative competency. Any not exactly accurate estimate they do is an overestimate. Correspondingly, the most competent person cannot overestimate their relative competency. This leads to the perception of bias where there is none, except a trivial tautological one.
- danbruc 4y agoI made you a picture [1]. I randomly generated 100 test scores between 0 and 1, then different self assessments. Top left, self assessment matches actual score, top middle, self assessment varies uniformly by ±0.1 around the test score, top right, self assessment varies uniformly by ±0.2 around the test score. None of those have a Dunning-Kruger effect. If you aggregate data points, there will be - as you mentioned - an edge effect because the self assessment will get clipped. In the bottom row I added a Dunning-Kruger effect, at a test score of 0.7 the self assessment is perfect, below and above that the self assessment is off by 0.5 times the distance of the test score from 0.7. Otherwise the bottom charts are the same, no random variation on the left, ±0.1 in the middle and ±0.2 on the right. You can see that the edge effect is less important as the data points are steered away from the corners. I will admit that the original Dunning-Kruger chart could or could not show a real effect, really depends on how they aggregated the data and how noisy self assessments are. But if you have a raw data set like the one I generated, you could easily determine if there is an effect. If one could find such a data set, I would like to have a look. [1] https://imgur.com/g4frW6p https://imgur.com/g4frW6p
- geysersam 4y agoI see what you mean. It should be possible to determine if there is an effect using the raw data from the experiment.
- dgb23 4y agoTangential, but the more interesting question for me is: How does estimating my skill level influence skill growth, social relationships and decision making? I think there are a bunch of useful angles to this. When there are risk/responsibility opportunities, then I need to be courageous. When it’s about learning and interacting collaboratively, then I need to be humble.
- apienx 4y ago> Collectively, the three critique papers have about 90 times fewer citations than the original Dunning-Kruger article.5 So it appears that most scientists still think that the Dunning-Kruger effect is a robust aspect of human psychology.6 Critiques cite the work being critiqued (yes, the referenced critiques in TFA cite the Dunning-Kruger study). Also, a 23 year-old paper will inevitably get cited more than 6 year-old papers. But yeah...the inertia in Science is real. That conservatism's a feature, not a bug. Psychology's probably the discipline with the shortest "half-life of knowledge. https://en.wikipedia.org/wiki/Half-life_of_knowledge https://en.wikipedia.org/wiki/Half-life_of_knowledge
- nomilk 4y agoThanks for introducing me to the term "half-life of knowledge". > An engineering degree went from having a half life of 35 years in ca. 1930 to about 10 years in 1960. A Delphi Poll showed that the half life of psychology as measured in 2016 ranged from 3.3 to 19 years depending on the specialty, with an average of a little over 7 years. This is very interesting and makes me wonder what it is for tech careers, e.g. web devs, data scientists etc.
- hallway_monitor 4y agoFor javascript developers it's about six months! Kidding aside, it seems you could estimate it by asking, what portion of the knowledge I use did I learn 20 years ago? Then 10, 5, 1. For me it seems to be somewhere around ten years.
- tpoacher 4y agoDunningKruger.OtherDefinitions.append( article )
- jakear 4y agoArticle seems to be saying “DK doesn’t exist because it always exists”. Which is… absurd? The point of DK is that when you don’t know shit, any non-degenerate self assessment will result in overestimating your ability. In short, “there are more natural numbers above smaller natural numbers than bigger ones”. This doesn’t have to do with psychology, and it’s expected that it appears when evaluating random data. That’s a good thing! It means DK exists even when us pesky humans aren’t involved at all, not that DK doesn’t exist at all.
- larwent 4y agoSeems pretty simple. When we create upper and lower boundaries to some score, people with lower scores have more space to overestimate and those with higher scores more space to underestimate, causing the perceived score to trend towards the mean. I think there's both a component of numbers and psychology here. If the dispersion in perceived score caused by inaccuracy is wide enough to touch the bounds, it will force a trend towards the mean. This effect is possibly exacerbated by a tendency of perception to stray from "extremes", so subjects with a score near the edges will trend to the mean more strongly as they are unlikely to rate themselves the very best or very worst.
- ewzimm 4y agoThis seems pretty simple to correct, so I'm skeptical that nobody has done so yet in these experiments. If true, it's an equally interesting oversight as the Monty Hall problem. The basic premise is that the structure of an experiment will naturally nudge randomness in a particular direction, and we need to adjust for that in the analysis. Everyone who does this type of work should know this. In a simplified experiment where we give people a 3 question quiz, those who got 2 questions right have one overestimation option, 3, and two underestimation options, 0 and 1. So it's very easy to adjust for autocorrelation by checking if a large group of 2-scorers underestimate more than twice as often as they overestimate. Then we see how their tendencies compare against 1-scorers and how they deviate from naturally overestimating more than twice as often as underestimating. I haven't reviewed these types of papers, but if nobody made even that basic adjustment in their analysis, how many others have been missed in experiments like this?
- deleted 4y ago[deleted]
- larwent 4y agoUnless I missed something, this article doesn't explain WHY random data can result in a Dunning-Kruger effect. The relationship between the "actual" and "perceived" score is a product of bounding the scores to 0-100. When you generate a random "actual" score near the top, the random "perceived" score has a higher chance of being below the "actual" the numerical below is larger than the one above, and vice-versa. E.g. a "test subject" with an actual score of 80% has a (uniform random) 20% chance of overestimating their ability and an 80% of underestimating it. For an actual score of 20%, they have an 80% chance of overestimating.
- ImaCake 4y agoAs explained in the article, the reason is autocorrelation. Basically the y axis is correlated to the x axis because the y axis is actually x + random noise. The dunning kruger graph is then a transformation of that data - still subject to autocorrelation.
- once_inc 4y agoA person with an actual score of 80% will probably have enough confidence in his or her abilities due to experience that they will tend not to rate themselves low. Imagine being a graduate student asked how high (s)he would rank. They would not rank themselves as low as they might have when they were sophomore students. They would probably rank within 20% of their actual score, which is what the final graph in the article shows; professors have enough experience to be able to self-assess themselves better than less experienced subjects can.
- fallingfrog 4y agoOK, I think I understand. What the data from the original experiment actually shows is that people at all skill levels are pretty bad at estimating their skill level- it's just that if you scored well, the errors are likely to be underestimates, and if you scored badly, the errors are likely to be overestimates, by pure chance alone. So it's not that low scoring individuals are particularly overconfident so much as everyone is imperfect at guessing how well they did. Great observation.
- civilized 4y agoSo it seems... but I still don't understand why the author thinks it's helpful to say "autocorrelation" dozens of times when he could have just said this.
- diwank 4y agoExcerpt from a newer paper by Nuhfer (2017) adds more clarity: “… Our data show that peoples' self-assessments of competence, in general, reflect a genuine competence that they can demonstrate. That finding contradicts the current consensus about the nature of self-assessment. Our results further confirm that experts are more proficient in self-assessing their abilities than novices and that women, in general, self-assess more accurately than men. The validity of interpretations of data depends strongly upon how carefully the researchers consider the numeracy that underlies graphical presentations and conclusions. Our results indicate that carefully measured self-assessments provide valid, measurable and valuable information about proficiency. …” https://www.researchgate.net/publication/312107583_How_Random_Noise_and_a_Graphical_Convention_Subverted_Behavioral_Scientists'_Explanations_of_Self-Assessment_Data_Numeracy_Underlies_Better_Alternatives https://www.researchgate.net/publication/312107583_How_Rando...
- ithkuil 4y agoIn other words people are quite bad at estimating their skill level. Some people will overestimate, while some other people will underestimate and on average there will be a relatively constant estimated skill level that doesn't change all that much based on the actual abilities. Given that fact, it logically follows that people who score low ability tests will more often than not have overestimated their ability (and the same on the other end of the spectrum). You can frame this effect as autocorrelation if you wish or just as a logical consequence. But that's missing the point. The point is: why on earth are humans so bad at estimating their own competence level as to make it practically indistinguishable from random guesses.
- kizer 4y agoThat’s what I was thinking; if the average is about constant all you’ve shown is that everyone is bad at self-assessment (another issue - not fully qualifying a distribution by just using the average loses information). But a comment above quoting the more recent paper presents a contradictory conclusion: that humans can self-assess with some accuracy. So now I’m confused again.
- d0mine 4y agoYou might have meant this comment https://news.ycombinator.com/item?id=31039901 https://news.ycombinator.com/item?id=31039901 (it says people can self-assess (no bias), more competent people do it better (less variance))
- d0mine 4y ago- what DK claims: there is bias (incompetent people overestimate their ability) - what data actually shows: there is a greater variance (incompetent people both over and _under_ estimate to a larger degree compared with more competent people. Data shows heteroscedasticity. No bias (estimations are around zero +/-, tighter for more competent).
- ithkuil 4y agoThe data supports the claim because indeed it turns out that incompetent people overestimate their ability. This phenomenon exhibits itself with random data too, so it clearly doesn't mean that incompetent people overestimate their ability because of their incompetence. Or is it? The trick lies in the fact that when asked to judge your competence you're given a range (e.g. 0-10) and both competent people and incompetent people have access to the whole range when taking a self-assessment. I.e. if less competent people were on average more aware of their incompetence they may be less likely to rate themselves 5 or 6, but yet the data shows that no matter what competence level you have on average you self-assess more or less the same. This seems to imply that your incompetence indeed doesn't allow you to truly appreciate the full range of skills that are required to reach a higher level of competence. In other words, the DK effect itself is the cause of the random distribution of the skill self-assessment (which in turn is the cause of the overestimation secondary effect)
- crashingintoyou 4y agoSome other Dunning-Kruger critiques aggregated by Andrew Gelman: https://statmodeling.stat.columbia.edu/2021/10/12/can-the-dunning-kruger-effect-be-explained-as-a-misunderstanding-of-regression-to-the-mean/ https://statmodeling.stat.columbia.edu/2021/10/12/can-the-du...
- kizer 4y agoI’m not a scientist, but wouldn’t it make sense for standard practice to be to assume at first that there’s a shared variable (that you have introduced) and to look for it until you’re certain the things you’re plotting are independent? Of course they may not be in the end as that’s the “goal”, but the shared variable if there is indeed causation in that case will be what you’re looking for, not one of the variables you “know”.
- askasp 4y agoIf we assume random data then the people at the lower end will over-estimate their own performance the same amount that people on the higher end will under-estimate theirs. However, if the under-performers consistently over-estimate more than the over-performers under-estimate there is still some merit to the effect, isn't there? That is, the interesting number is the difference between integral of y-x on lower half vs the integral of y-x on the upper half. Does that make sense to anyone else?
- m3047 4y agoYeah, I think so and we're probably in the minority here. There are a couple of other comments referring to regression to the mean and that the article takes a literalist view which is perhaps unwarranted. You win the followup comment. ;-) I confess that I've never paid that much attention to the classic D-K graph, and that taking a close look at it, it is most assuredly crap. Now I want to know what the plots of the actual scores for those quartiles look like rather than %ile, or after-the-fact ranking. Yeah, it sure looks like people mostly figure they're in the 55-75 %ile ranking, if that's what that actually is, and that where in that spread they think they are correlates with their actual ranking. Let's go down a Bayesian rabbit hole. Let's assume, as does the article, that people's self estimations are completely random rubbish: the worst people have nowhere to go but up, the best nowhere but down. Yup, completely agree. Now let me ask a question: is self-estimation of any use in determining actual ability? The answer in this case is no: knowing one does not inform our ability to know the other in a Bayesian sense, they are not correlated. D-K sounds valuable as a cautionary tale concerning excessive exuberance and a tendency not to learn well from experience, but aside from child-proof caps and Mr. Yuk stickers where we really want to apply the lesson is at the high-performing end of the scale and here we get into trouble immediately. It is tempting to say "high-performers have nowhere to go but down" as though maybe we should reject those self-reporting the best performance. The classic chart hints at high performers underestimating their true performance, but it's a crappy chart; maybe they want it to be true. But in the specific case where there is utterly no correlation and true performance is as evenly distributed as self-assessment, if we chop off the "top X self-reporting" we will chop off just as many poor performers as high performers. Yes, I hear you, and I agree, random is an edge case; I just don't believe that affects its prevalence. Maybe it is true; alright dust off those priors and have at it.
- deleted 4y ago[deleted]
- newbamboo 4y agoAutocorrelation is much more interesting, and much more important topic than dk, which mostly seems to be popular concept because it supports biases and other fallacious, ego driven thinking. Autocorrelation is an under-appreciated problem, particularly in the social sciences and Econ. So it’s nice to use dk to catch the attention of the masses to spread the word about autocorrelation.
- highfrequency 4y agoThe author is onto something that Dunning-Kruger is suspicious, but the argument is wrong. The "statistical noise" plot actually demonstrates a very noteworthy conclusion: that Usain Bolt estimates his own 100m ability as the same as a random child's. This would be a great demonstration of the Dunning-Kruger effect, not a counterargument. On the other hand, regression to the mean rather than autocorrelation does explain how you could get a spurious Dunning-Kruger effect. Say that 100 people all have some true skill level, and all undergo an assessment. Each person's score will be equal to their true skill level plus some random noise based on how they were performing that day or how the assessment's questions matched their knowledge. There will be a statistical effect where the people who did the worst on the test tend to be people with the most negative idiosyncratic noise term. Even if they have perfect self-knowledge about their true skill, they will tend to overestimate their score on this specific assessment. Regression to the mean has broad relevance, and explains things like why we tend to be disappointed by the sequel to a great novel.
- IncRnd 4y ago》It’s the (apparent) tendency for unskilled people to overestimate their competence. Close. It's the cognitive bias where unskilled people greatly overestimate their own knowledge or competence in that domain relative to objective criteria or to the performance of their peers or of people in general.
- oh_my_goodness 4y agoDunning and Kruger showed that students all thought they were in roughly the 70th percentile, regardless of where they actually ranked. That's it. The plots in the original paper make that point very clear. It is unnecessary to walk the reader through autocorrelation in order to achieve a poorer understanding of that simple result.
- woah 4y agoSeems like Dunning and Kruger suffered from the Dunning-Kruger effect
- titzer 4y agoI find the article frustrating because of the tone. It's also wrong. They misunderstood what lines mean. This article is absolutely dripping with condescension throughout and is really pushing a "gotcha" that doesn't exist. It then argues basic statistics, generates a DK-looking graph from random data, and then claims the phenomena doesn't exist. When in fact, as other people have commented, when people are bad at estimating their own ability (i.e. random), the DK effect still exists; it falls out of statistics. Sigh, the author misunderstood the very definition of the DK effect: > "The Dunning–Kruger effect is the cognitive bias whereby people with low ability at a task overestimate their ability. Some researchers also include in their definition the opposite effect for high performers: their tendency to underestimate their skills." In all the examples, this holds, even if the assessment ability is totally random. Even if every quartile gives themself an average score, like the random data generated here. The author seems to think that it should be even more lopsided or something to demonstrate the effect. (I mean, honestly, what are they expecting, a line above 50th percentile? A line with negative slope? What?) If there were no DK effect, the two lines would be the same. Instead, if we go back and look at the original data, we see indeed, the two lines are not the same, the average for the bottom quantile is over 50%, there is some small increase in perceived ability associated with actual ability (and not the opposite). The sin here isn't some autocorrelation gotcha, but rather, DK should have put error bars on the graph. If it was totally random, the error bars would be all over the place.
- bena 4y agoThe fact that you can generate a Dunning-Kruger looking graph using nothing but noise does indicate that the graph isn't proof of anything. He also points out that the problem is that there's nothing below zero and nothing above 100. You can't have people who estimate beyond that. He uses another study and it turns out, the less knowledgeable you are about a skill, the worse you are at estimating your ability at all. In both directions. If the lines were the absolute difference between perceived ability and actual ability, for no effect, the lines still shouldn't be the same. They should converge towards those who are knowledgeable. If anything, the difference line should be nearly a horizontal line. Because there should be greater variance in estimations at the lower end.
- edtechdev 4y agoThere have already been responses to this criticism before, such as: https://drbenvincent.medium.com/the-dunning-kruger-effect-probably-is-real-9c778ffd9d1b https://drbenvincent.medium.com/the-dunning-kruger-effect-pr... including from David Dunning himself https://thepsychologist.bps.org.uk/volume-35/april-2022/dunning-kruger-effect-and-its-discontents https://thepsychologist.bps.org.uk/volume-35/april-2022/dunn...
- TimPC 4y agoThe article is correct. The effect is statistical not psychological. It emerges even from artificial data and occurs independently of the supposed psychological justifications even for data where those justifications are clearly removed. If you adjust the experiment design to avoid introducing the auto-correlation you get data that doesn't show the DK effect at all. Some might take issue with the adjusted experiment as using seniority related categories like "sophomore" and "junior" as skill levels has its own issues. To show the DK effect is real you need to come up with a better adjusted experiment that avoids the autocorrelation while still generating data that generates the effect. It's unclear if that's possible.
- deleted 4y ago[deleted]
- hungrygs 4y agoJust anecdotal, but my life observation of DK is often highly intelligent and competent people in a particular field who then generalize that to pontificate and proclaim, directly or indirectly, superior understanding to certified domain experts (e.g., have directly related advanced degree(s), work in the field for decades.) It thus seems more or as much a psychological effect - in short, people with a personality type of superiority and know-it-all, yet have never done the deep and hard work to gain or demonstrate any competency in said areas. A common side observation is of course unfounded conspiracy theories, that the derided experts have sinister intentions.
- NaturalPhallacy 4y agoClassic example: >The V-tail design gained a reputation as the "forked-tail doctor killer",[16] due to crashes by overconfident wealthy amateur pilots,[17] fatal accidents, and inflight breakups.[18] "Doctor killer" has sometimes been used to describe the conventional-tailed version, as well. https://en.wikipedia.org/wiki/Beechcraft_Bonanza https://en.wikipedia.org/wiki/Beechcraft_Bonanza
- deleted 4y ago[deleted]
- jtc331 4y agoBecause of the effect that is actually found (variance is higher the less achievement) it follows that people you encounter who wildly overestimate their ability are more likely to people who are poor performers (the same is true for the inverse, but they obviously don't stand out anecdotally to us). IMO that explains why Dunning Kruger seems intuitively correct even if the conclusion they drew isn't actually correct.
- seventytwo 4y agoAgree. This should have been the original conclusion of DK, if they hadn’t made the mistake. Another way to show this would have been to keep the auto correlation plot, but compare it to the same plot with statistical noise. With infinite random data, the expected value for self-assessment would be 50% score, regardless of actual score - a flat line through the chart. It would then be significant to find a non-flat line, as DK did. It’s not inconceivable that with a smaller sample, you’d get come biasing, where lesser skilled people would over estimate, and higher skilled people under estimate. The follow up studies seem to suggest there’s not really a bias like that, but that there is a “honing” of the general ability to estimate your own outcome, which makes sense. > Although there is no hint of a Dunning-Kruger effect, Figure 11 does show an interesting pattern. Moving from left to right, the spread in self-assessment error tends to decrease with more education. In other words, professors are generally better at assessing their ability than are freshmen. That makes sense. Notice, though, that this increasing accuracy is different than the Dunning-Kruger effect, which is about systemic bias in the average assessment. No such bias exists in Nuhfer’s data.
- caylus 4y ago> people you encounter who wildly overestimate their ability are more likely to people who are poor performers How is this helpful? You won't know whether someone is "overestimating" their ability until you learn both their estimated and actual performance, at which point you don't need to guess whether they're "likely" to have poor actual performance.
- jtc331 4y agoIt's an explanation for why our anecdotal intuition concludes what it does here. I think you're misreading the point of my comment.
- Dave_Rosenthal 4y agoThis was interesting to me so I spent a while this AM playing with a Python simulation of this effect. I used a simple process model of a normally-distributed underlying 'true skill' for participants, a test with questions of varying difficulty, some random noise in assessing whether the person would get the question right, noise in people's assessments of their own ability, etc. I fiddled with number of test questions, amounts of variation in question difficulty, various coefficients, etc. In none of my experiments did I add a bias on the skill axis. My conclusion is that the "slope < 1" part of the DK effect (from their original graph) is very easy to reproduce as an artifact of the methodology. I could reproduce the rough slope of the DK quartiles graph with a variety of reasonable assumptions. (One simple intuition is that there is noise in the system but people are forced to estimate their percentiles between 0 and 100, meaning that it's impossible for the actual lowest-skill person to underestimate their skill. There are probably other effects too.) However, I didn't find an easy way using my simulation to reproduce the "intercept is high" part of the DK effect to the extent present in the DK graphs, i.e. where the lowest quartile's average self-estimated percentile is >55%. (*) However, it strikes me that without a very careful explanation to the test subjects of exactly how their peer group was selected, it's easy to imagine everyone being wrong in the same direction. (*) EDIT: I found a way to raise the intercept quite a lot simply by modeling that people with lower skill have higher variance (but no bias!) in their own skill estimation. This model is supported by another paper the article references.
- SomewhatLikely 4y agoWouldn't variance be influenced by a similar bounding effect but this time from the upper side? That is, if your true skill is 98% you aren't going to ever overestimate by more than 2%, but if your true skill is 50% you could be off by up to 50% in either direction.
- georgefox 4y agoThis is a fascinating discussion, to which I have little to add, except this. Quoting the article (including the footnote): > [I]f you carefully craft random data so that it does not contain a Dunning-Kruger effect, you will still find the effect. The reason turns out to be embarrassingly simple: the Dunning-Kruger effect has nothing to do with human psychology[1]. > [1]: The Dunning-Kruger effect tells us nothing about the people it purports to measure. But it does tell us about the psychology of social scientists, who apparently struggle with statistics. It seems to me that despite rudely criticizing a broad swath of academics for their lack of statistical prowess, the author here is himself guilty of a cardinal statistical sin: accepting the null hypothesis. The fact that data resemble a random simulation in which no effect exists does not disprove the existence of such an effect. In traditional statistical language, we might say such an effect is not statistically significant, but that is different from saying that the effect is absolutely and completely the result of a statistical artifact. The nuance of statistics is never-ending.
- ImaCake 4y agoLater in the article the author points to an article which does systematically illustrate that the D-K effect is probably not real. They achieve this by using college education level as an independent proxy for test skill with the Y variable being an unrelated assessment of skill - self-assessment. So we can be pretty confident that the D-K effect is at least very small.
- ncmncm 4y agoMost citations of D-K are themselves examples of D-K.
- bandyaboot 4y agoMy takeaway, which may be flawed, is that the DK effect really hasn’t been debunked it any fundamental way. It’s just that the effect is statistical rather than psychological. High skilled individuals are still more likely to underestimate their skill level while low skilled individuals are still more likely to overestimate theirs. It’s just that everyone is bad at estimating their skill level and high skilled individuals have more room to estimate below their actual, while low skilled individuals have more room to miss above. Is my reasoning flawed in some way?
- cryptica 4y agoI've felt inadequate throughout most of my early career. That's how I know that the confidence I have today is well deserved. I've never had impostor syndrome though. To have impostor syndrome, you have to be given opportunities which are significantly above what you deserve. I did get a few opportunities in my early career which were slightly above my capabilities but not enough to make me feel like an impostor. In the past few years, all opportunities I've been given have been below my capabilities. I know based on feedback from colleagues and others. For example, when I apply for jobs, employers often ask me "You've worked on all these amazing, challenging projects, why do you want to work on our boring project?" It's difficult to explain to them that I just need the money... They must think that with a resume like mine I should be in very high demand or a millionaire who doesn't need to work. I've worked for a successful e-learning startup, launched successful open source projects, worked for a YC-backed company, worked on a successful blockchain project. My resume looks excellent but it doesn't translate to opportunities for some reason.
- sfvisser 4y agoMy intuition for this is: given a fixed and known scoring range (say 0..100), when scoring very low there is simply a lot of room for overestimating yourself and when scoring very high there is simply a lot of room for underestimating yourself. So all noise ends up adding to the inverse correlation naturally.
- sanp 4y agoSeems like a half-baked analysis. You would plot x=x to show where y is above and where it is below. It is useful for exposition. The author questions this as if it is an analytical oversight.
- playpause 4y agoI’ve always felt the DK effect is cynical pseudoscience for midwit egos. It’s a sophistic statement of the obvious dressed up as an insight. But worse, it serves to obvert something interesting and beautiful about humans - that even very intellectually challenged people sometimes can, over time, develop behaviours and strategies that nobody else would have thought of, and form a kind of background awareness of their shortcomings even if they aren’t equipped to verbalise them, allowing them to manage their differences and rise to challenges and social responsibilities that were assumed to be beyond their potential. Forrest Gump springs to mind as an albeit fictional example of the phenomenon I’m talking about. I think this is a far more interesting area than the vapid tautology known as the DK effect.
- jollybean 4y agoI think there's an easier explanation for the effect, and that is people are just not very good at judging their skill level, and due to reversion to the mean, low-performers probably overestimate and high-performers underestimate. And also, I think there is actually a tiny bit of DK going on. And then, as you say, it gets amplified by the pseudo-literati.
- MrYellowP 4y agoAnd yet, very stupid people are too stupid to recognize that they're very stupid. Not a single word in that blogpost changes anything about that.
- knorker 4y agoNo it isn't. If everyone responded that they are 50% skilled (or per this article, that it's randomly distributed), then we 1. See the same graph, and 2. Bad people overestimate, and good people underestimate This article merely describes Dunning Kruger. Accidentally proves it mathematically, but thinks that it debunks it.
- orf 4y ago> To measure ‘skill’, Nuhfer groups individuals by their education level… I’m surprised this wasn’t flagged as something pretty silly.
- mcguire 4y ago"Academic rank" is an awfully weird proxy for skill, though.
- deleted 4y ago[deleted]
- mattwilsonn888 4y agoI would have been more interested in seeing the raw data from the original Dunning-Kruger study reformatted to avoid auto-correlation. Maybe I've skipped over an important detail in my head, but I don't see why plotting perceived test score vs. actual test score would cause any problems; neither variable is in terms of the other. The final study discussed is convincing as far as I thought. By using academic rank (Freshman, Sophomore, ...) they can plot the difference between difference in score and predicted score against rank without auto-correlation. Its just that using academic rank seems a possibly unreliable metric and an unnecessary complication - why not just use data about test scores and predictions of scores which already exists in a proper statistical interpretation?
- trombonechamp 4y agoThat is not what the term "autocorrelation" means. Autocorrelation is the correlation of a vector/function with a shifted copy of itself.
- jl2718 4y agoSo, they observe a bias toward the average, and the dependence goes exactly as one would naively expect. If scientists exist to explain things we find interesting, statisticians exist to make those things boring. Seriously, work as a data scientist and you end up busting hopes and dreams as a regular part of your job. Almost everything turns out to be mostly randomness. The famous introduction to a statistical mechanics textbook had me pondering this. If life really is just randomness, it’s hard to find motivation. From a different viewpoint, however, I’ve found that the people that embrace this concept by not trying to control things too much, actually end up with the most enviable results, although I may be guilty of selection bias in that sample.
- roguecoder 4y agoThe point is that if people estimate their abilities at random, with no information, it will look like people who perform worse over-estimate their performance. But it isn't because people who are bad at a thing are any worse at estimating their performance than people who are good at the thing: they are both potentially equally bad at estimating their performance, and then one group got lucky and the other didn't. It would require them to be _even worse that random_ for them to be worse at estimating their abilities, rather than simply being judged for being bad at the task. It is only human attribution bias that leads us to assume that people should already know whether they are good or bad at a task without needing to being told. The study assumed that the results on the task are non-random, performance is objective, and that people should reasonably have been expected to have updated their uniform Bayesian priors before the study began. If any of those are not true, we would still see the same correlation, but it wouldn't mean anything except that people shared a reasonable prior about their likely performance on the task. People will nevertheless attribute "accurate" estimates to some kind of skill or ability, when the only thing that happened is that you lucked into scoring an average score. You could ask people how well they would do at predicting a coin flip and after the fact it would look like whoever guessed wrong over-estimated their "ability" and a person who guessed right under-estimated theirs, even though they were both exactly accurate. This comment section clearly demonstrates the attribution bias that makes this myth appealing, though. And this blog post demonstrates how difficult it is to effectively explain the implications of Bayesian reasoning without using the concept.
- roguecoder 4y agoConsider the original study: they used 45 Cornell undergraduate students and asked them about grammar. Grammar isn't objective. Everyone there had performed well on the verbal portion of the SAT, but they weren't studying grammar and hadn't gotten instruction on this particular book of grammar they were judged against. It is very likely that what they were capturing in the "better" or "worse" scores is differences in local dialect. They then judged people whose beliefs about grammar varied from the one book's beliefs about grammar as having over-estimated their performance. They took people out of one context, asked them how they would behave in a novel context, and everyone made an educated guess. The people who guessed correctly were judged to accurately know their own abilities, when actually they may just have gotten lucky. Thus what Dunning-Kruger's paper actually says is that if you want people to know how you would like them to perform a task, you can't assume they will read your mind: you have to provide them with actual feedback on their performance.
- poulpy123 4y agoSo dunning and Kruger were victims of the dunning-kruger effect ?
- js8 4y agoThe opposite - the publishing pressure makes experts overconfident in their abilities. The non-experts then assume that the experts, for sure, know better, so they don't look for flaws.
- nathias 4y agoGreat article, there should be much more common knowledge of statistics and it's problems, it is surely the most abused of all sciences.
- srvmshr 4y agoHence the quote: "There are white lies, damned lies and statistics" Funny that all the major ML marvels are also built on statistical foundations - a tool used as much as abused.
- _dain_ 4y agoI don't find the "autocorrelation" explanation intuitive (although it may be equivalent to what I'm about to suggest). The way I think about it, is that it comes about because the y-axis is a percentile rank. How does it actually work for people to give unbiased estimates of their performance as percentiles? For the people at the 50th percentile in truth, they could give a symmetric range of 45-55 as their estimates, and it would be unbiased. But what about the people at the 99th percentile? They can't give a range of 94-104, the scale only goes as high as 100. So even if they are unbiased (whatever that means in this context), their range of estimates in percentile terms has to be asymmetrical, by construction. So, even if people are unbiased, if you were to plot true percentile vs subjective estimated percentile, the estimated scores would "pull toward" the centre. Then the only thing you need to replicate the Dunning-Kruger graph is to suppose that people have a uniform tendency to be overconfident, i.e. that people over-rate their abilities, but to an extent unrelated to their true level of skill. The estimated score at the left side of the graph goes higher, but it can't go as high on the right side of the graph because it butts up against the 100 percentile ceiling. Then you end up with a graph that looks like lower skilled people are more overconfident than higher skilled people are underconfident.
- derbOac 4y agoIt's an interesting article but the author is using terms a little incorrectly or strangely I think, and making untrue statements. The basic points are important and interesting to think about, but could've been explained more clearly.
- _dain_ 4y agoYes, when I read "autocorrelation" I think of a time-series variable that is correlated with its own lagged values.
- oceliker 4y agoI think the gist of the article is this: Suppose you make 1000 people take a test. Suppose all 1000 of these people are utterly incapable of evaluating themselves, so they just estimate their grade as a uniform random variable between 0-100, with an average of 50. You plot the grades of each of the 4 quartiles and it shows a linear increase as expected. Let's say the bottom quartile had an average of 20, and the top had 80. But the average of estimated grades for each quartile is 50. Therefore, people who didn't do well ended up overestimating their score, while people who did well underestimated it. In reality, nobody had any clue how to estimate their own success. Yet we see the Dunning-Kruger effect in the plot.
- laszlokorte 4y agoSomeone who has a skill=0 can not underestimate and someone with a skill=100 can not overestimate. So by the framing of the question alone the participants are nudged to estimate there own skill "more averagely".
- fullshark 4y agoYeah we learned people are bad at giving themselves percentile rankings apparently, especially when the population is illdefined ("your peers"). https://www.avaresearch.com/files/UnskilledAndUnawareOfIt.pdf https://www.avaresearch.com/files/UnskilledAndUnawareOfIt.pd...
- andersource 4y agoThat's the way I understand the statistical analysis, and in my view this exactly supports (not contradicts) DK: > In reality, nobody had any clue how to estimate their own success. Wouldn't that mean unskilled people tend to overestimate their skill, and experts tend to underestimate it? Why is there a contradiction with DK's conclusions?
- oceliker 4y ago> Wouldn't that mean unskilled people tend to overestimate their skill, and experts tend to underestimate it? I think it's because the original paper speculates far beyond it: > The authors suggest that this overestimation occurs, in part, because people who are unskilled in these domains suffer a dual burden: Not only do these people reach erroneous conclusions and make unfortunate choices, but their incompetence robs them of the metacognitive ability to realize it. The argument about autocorrelation says this "dual burden" doesn't need to be there to observe the effect.
- andersource 4y agoVery interesting article and statistical analysis, but I really don't see how it concludes that the DK effect is wrong based on the analysis. The fact that the DK effect emerges with _completely random data_ is not surprising at all - in this case the intuitive null hypothesis would be that people are good at estimating their skill, therefore there would be strong a correlation between their performance and self-evaluation of said performance. If the data weren't related, then this hypothesis isn't likely, which is exactly what DK means. And indeed if you look at the plots in the article (of the completely random data), they depict a world in which people are very bad at estimating their own skill, therefore, statistically, people with lower skills tend to overestimate their skills, and experts tend to underestimate it. Also wanted to point out that in general there is no issue with looking at y - x ~ x, this is called the residual plot, and is specifically used to compare an estimate of some value vs. the value itself. That being said, the author seems very confident in their conclusion, and from the comments seems to have read a lot of related analyses, so I might be missing something. ¯\_(ツ)_/¯
- _dain_ 4y ago>The fact that the DK effect emerges with _completely random data_ is not surprising at all - in this case the intuitive null hypothesis would be that people are good at estimating their skill, therefore there would be strong a correlation between their performance and self-evaluation of said performance. If the data weren't related, then this hypothesis isn't likely, which is exactly what DK means. DK effect is not that low skill people are overconfident and high skill people are underconfident. It is specifically that low skill people are more overconfident than high skill people are underconfident. i.e. if someone's estimated skill is true_skill+bias+noise, then bias_lowskill > -bias_highskill. This is very clear in the original DK paper, they specifically focus on the supposed metacognitive deficiencies of low-skill people. The article argues that the graphs supposedly demonstrating this fact, can also be generated from a model that does not have this difference, i.e. where bias_lowskill == bias_highskill. EDIT: My characterization of the article is not correct, see here[1] for a visualization of the point I'm trying to make. [1] http://emilkirkegaard.dk/understanding_statistics/?app=Dunning_Kruger http://emilkirkegaard.dk/understanding_statistics/?app=Dunni...
- PaulKeeble 4y agoModern Psychology is having a lot of these sorts of results over the last decade, none of their methods are holding up under proper scrutiny. They are struggling to reproduce findings but more critically even the reproduced ones are turning out to be statistical and mathematical errors like shown here. Some of the findings have also done severe harm to patients over the decades as well, I can't help but think we need a lot of caution when it comes to psychology results given its harmful uses (such as the abuse of ill patients) and its lack of truthful results.
- Phileosopher 4y agoI'm convinced it's associated with the methodology of how psychology has approached matters. In the world of biology, you're observing the world around you. Same for physics, chemistry, et al. This means that you can set up proper controls to obscure your own presence from any potential results (e.g., isolate everything in another room, use cameras to avoid being near animals, etc.) Psychology has the same nightmare as quantum physics: pre-existing thoughts and beliefs literally define what results you end up with. I'm convinced that psych is a victim of the "new" way of doing science: treating the Scientific Method™ as a self-evident concept instead of regarding science as a vastly certain domain of metaphysics.
- BlueTemplar 4y agoWell, all sciences have this issue : see Kuhn, Feyerabend... https://samzdat.com/2018/05/19/science-under-high-modernism/ https://samzdat.com/2018/05/19/science-under-high-modernism/ It would be kind of ironic if psychologists were more susceptible to the "Pop-Baconian" simplification of science ? And damning if what I've heard about psychology going through multiple paradigms during the 20th century alone is true. But then, indeed, also understandable, due to the "softness" and "paradigmlessness" of the subject matter, as Kuhn had pointedout the later back then ? It's still sad how we now have detailed theories and histories of science, but in practice scientists show no interest in trying to learn from them, nor the mistakes of their predecessors? But then maybe that would have too high of a cost. (Consider all those successful projects where the founders later say : "we didn't knew what we were getting into / that this was considered impossible".) Bonus : Wiseman & Schlitz’s attempts to do an adversarial super-controlled parapsychological experiment : ( IV. ) https://slatestarcodex.com/2014/04/28/the-control-group-is-out-of-control/ https://slatestarcodex.com/2014/04/28/the-control-group-is-o...