4 ms·
Inter-rater reliability is super important for tests like this. http://en.wikipedia.org/wiki/Inter-rater_reliability http://en.wikipedia.org/wiki/Inter-rater_re
by boredguy8 14y ago
Inter-rater reliability is super important for tests like this. http://en.wikipedia.org/wiki/Inter-rater_reliability http://en.wikipedia.org/wiki/Inter-rater_reliability The gist is: you can't simply mark one person as "asian" and assume that categorization is correct. In that respect, the data would reveal more about the person sorting the photos than it would reveal about the perceptions of those that are rating the photos.
Second, there is a huge problem with causality here. So for instance, the author writes: "Be Asian if you want to appear smart; Latino if you want to appear extroverted." The problem is that there is a methodological flaw. On the first photo I saw on judge.me, I was presented with this image: http://images.judg.me/82e7fcbd988dbdcac0d00bd53fb93e96.jpg http://images.judg.me/82e7fcbd988dbdcac0d00bd53fb93e96.jpg This would appear to me to be a latino or hispanic male at a party. I'm highly inclined to rate them highly on the extrovert scale: they're at a party. But that doesn't indicate stereotypically latino or hispanic features indicate extroversion. It could be that people with stereotypically latino or hispanic features were more likely to upload photos in which the image portrayed a more stereotypically extroverted activity.
Third, it appears that users can upload a photo to the site and see their feedback from votes. It seems highly possible that users self-select a photo that will best affirm the image of themselves they wish to cultivate. In that respect, there's both a huge confirmation bias and huge self-selection bias. If I want to think of myself as an academic, I'll upload a picture of me at my desk studying and watch the "intellectual" ratings pour in. Then I can feel assured that other people perceive me the way I want to be perceived. Additionally, if one wants to conform to social expectations (and things like Asch's line test http://en.wikipedia.org/wiki/Asch_conformity_experiments http://en.wikipedia.org/wiki/Asch_conformity_experiments indicate conformity is common), this data might really be nothing more than showing the degree to which people post photos affirming their conformity to their social expectations (i.e. 'smart' ethnicities posting 'smart-looking' photos) and be saying nothing at all about how people actually perceive ethic cues.
There are huge methodological concerns for this 'study'. Instead, the revelation of this data might actually be the insight that "pictures of yourself at social events makes you look more social." Taking much of anything at all away from this data set would be rather unwise.
- hej 14y agoWhile inter-rater reliability is a concern for identifying ethnicities, other traits should be more reliable (gender, hair color). Still, stuff like that needs to be pre-tested. You basically let the coders code a limited set of photos and check the correlation between their codes. The higher the correlation, the better, >.7 is the convention in social science (but still pretty bad, higher would be better). You should also check intra-coder reliability, i.e. give the same person the same set of photos with two weeks or so between. You can then again calculate the correlation. This tells you wether your categories are too fuzzy (e.g. what exactly is medium length hair?). All in all this has serious methodical flaws, from a social science perspective it’s not salvageable, and I haven’t even talked about the complete lack of statistical tests (which, to be honest, would just be like polishing a turd).
- weissguy 14y agoI would think that a multi-level/hierarchical/mixed GLM would be an interesting approach to their data. Multilevel modeling assumes that there is correlation between observations that are inside the same "level". This is in stark comparison to regular GLM (even one with dummy variables to represent categories), which assumes that all observations are 100% independent. E.g. in a model that predicts students' GPA, you could divide your data into a hierarchy consisting of, at the highest level, geographic area, followed by high school, maybe followed by teacher. In that model, the correlation between students who are in the same state, the same school or in the same classroom would be accounted for. You could even go as deep as at an individual level if you have >1 observation per student. In addition to regular predictive variables, judg.me could probably use their weblogs to group people's judgement scores by country of origin and by individuals, among other possibilities.
- larrys 14y ago"In that respect, the data would reveal more about the person sorting the photos than it would reveal about the perceptions of those that are rating the photos." The problem is none of this matters since we don't know anything about the people rating the photos. Not their sex, not their age, not their location nothing. To wit: "and nothing about users who judge the photos." http://news.ycombinator.com/item?id=3921271 http://news.ycombinator.com/item?id=3921271