4 ms·
This is often required work in supervised learning / any machine learning with labels. Sometimes you get better result with human judgment on the labels, and t
by erdevs 10y ago
This is often required work in supervised learning / any machine learning with labels. Sometimes you get better result with human judgment on the labels, and that means you have to grind through a meaningful corpus of data.
If you believe what you're doing is important and beneficial, it doesn't feel like too much of a grind. Ultimately this was probably about a week of work, if every labeling participant independently reviewed ~3-5K articles.
Some people might have chosen to mechanical turk this, or to gather feedback from a subset of users and use those labels. Doing that might be a better approach as it might not only be less labor-intensive up-front but also allow the system to be regularly retrained easily (ie by gathering new label information as clickbaiters inevitably try to adapt their headlines). I imagine there were reasons for starting off this way, though.
- room271 10y agoIndeed, I've seen at least one paper where they used Mecanical Turk and aggregated results (by showing the same examples to multiple people) to check for quality. It sounds like they did almost exactly the same thing here. Paper: 'Antisocial Behaviour In Online Discussion Communities' (Cheng et al., 2015) Link: http://arxiv.org/pdf/1504.00680v1.pdf http://arxiv.org/pdf/1504.00680v1.pdf (Search for 'Mechanical Turk' / section Data Preparation -> Measuring text quality.)