4 ms·
Ah yes, ye olde "all humans are equally capable of all tasks" axiom (the paper is correct and the parent comment is wrong, perhaps obviously)
by nmca 4y ago
Ah yes, ye olde "all humans are equally capable of all tasks" axiom (the paper is correct and the parent comment is wrong, perhaps obviously)
- YeGoblynQueenne 4y agoTo be clear, that's a different criticism of the article's methodology than mine, yes? If you assume that the two sets of human annotators are different, then what are you comparing, exactly? The ability of one group to second-guess the other? There are other issues if you choose to assume that the two groups are fundamentally dissimilar: one group was two grad students, the other a number of Mechanical Turks. You can expect there to be more disagreement between (more than two? I'm not sure) Mechanical Turks than between two grad students (both in political science). Ultimately, my problem with the study is that the labelling they took as ground truth (i.e. that they compared ChatGPT and Mechanical Turks against) is too uncertain, or even subjective, to know for sure what exactly they found out. Edit: Oh, wait, I didn't put that in the comment above. Damn. I thought I had.
- nmca 4y agoIn the framing of your original comment: A: Mturkers B: ChatGPT C: Experts => ChatGPT better approximates the labelling of D by experts than mTurkers Which is a coherent and interesting conclusion. Edit: also, please forgive the snark in my first response
- YeGoblynQueenne 4y ago>> Edit: also, please forgive the snark in my first response No need to apologise! Your snark wasn't overboard, I thought. Anyway, big girl, can take it :) >> ChatGPT better approximates the labelling of D by experts than mTurkers That could be a "coherent and interesting conclusion" but it's not what the article really claims. The article's title is I think hedging its bets, by being very precise about who, exactly, was outperformed by ChatGPT, although it still manages to be vague about how ChatGPT outperformed the Mechanical Turks. I'm also really doubtful that two political science graduate students can be considered as "experts" in the annotation tasks they were called to perform, which were, again if I got that right, about content moderation. "Experts" in this setting would be people with experience in moderating discussion boards etc. I don't see that this was the case with the two students that provided the initial annotation.