3 ms·
Yeah, I think this is part of the problem. Is large-scale, low-quality data good? Sometimes it is (depending on the tradeoff), but from a model performance pers
by echen 5y ago
Yeah, I think this is part of the problem. Is large-scale, low-quality data good? Sometimes it is (depending on the tradeoff), but from a model performance perspective, it's often more effective to get smaller amounts of higher-quality data instead.
Hopefully people also don't need to be at the level of a professional linguist to label messages like "this is fucking awesome" correctly!
And great point on context. For example, the GoEmotions dataset didn't present labelers with the actual post or subreddit the message came from -- just the text itself. That makes it really difficult to label something like "his traps hide the fucking sun"! But once you see the comment in its original context https://www.reddit.com/r/nattyorjuice/comments/aee3wx/olympic_drug_tested_wrestler_revaz_nadareishvili/ee8jd5h/ https://www.reddit.com/r/nattyorjuice/comments/aee3wx/olympi..., and know that it's in the /r/nattyorjuice bodybuilding subreddit, it's much easier to realize that this is talking about someone's large muscles.
- viraptor 5y agoEven with proper labelling done by people who have lots of time to dig into each message, I don't believe we'd ever get a reasonable model. People can trivially hide the true meaning of sentences. The worst actually-consciously-racist accounts on Twitter will not have a single thing to report. You can find some insinuation or fragments where you know exactly what it all adds up to, then click report, get asked to choose messages to report and... yeah, not a single one of them is toxic in the literal sense.
- DarylZero 5y ago> The worst actually-consciously-racist accounts on Twitter will not have a single thing to report That's survivorship bias.
- hnbad 5y ago> Is large-scale, low-quality data good? It depends on what the purpose of content moderation is. Good if you want to accurately identify abusive behavior and protect users from harm? No. Good enough if you want to find the most blatant examples of name-calling and insults to appease regulators and trigger-happy lawyers by appearing to use "state of the art technology"? Sure.