6 ms·
One of the problems with real world machine learning is that engineers often treat models as pure black boxes to be optimized, ignoring the datasets behind them
by echen 5y ago
One of the problems with real world machine learning is that engineers often treat models as pure black boxes to be optimized, ignoring the datasets behind them. I've often worked with ML engineers who can't give you any examples of false positives they want their models to fix!
Perhaps this is okay when your datasets are high-quality and representative of the real world, but they're usually not. For example, many toxicity and hate speech datasets mistakenly flag texts like "this is fucking awesome!" as toxic, even though they're actually quite positive -- because NLP datasets are often labeled by non-fluent speakers who pattern match on profanity.
(So is 99% accuracy or 99% precision actually a good thing? Not if your test sets are inaccurate as well!)
Many of the new, massive scale language models use the Perspective API to measure their safety. But we've noticed a number of Perspective API mistakes on texts containing positive profanity, so this post was an attempt to explain the problem and quantify it.
- blowski 5y agoMy favourite of these was “stick your carrot in my fluffy bunny”. Humans are good at coming up with new ways of “being toxic”.
- rndgermandude 5y agoAs teens, some of my friends and I had a "game" where we would find new innocent phrases to describe sexual acts, taking a shit, and other vulgar/gross/unmentionable stuff. Yeah, stupid teenager humor, but we had a lot of fun nonetheless. We really didn't have a points system, but you got bonus rep if you'd be daring enough to utter the phrases you came up in the presence of adults and get away without a scolding - which resulted in one of the dad's kinda joining our game after he figured us out.
- native_samples 5y agoYou don't even need an insult like that. I just tried out Perspectives on this awesome insult I heard only yesterday: "I'm going to sleep with your father and then give him a son he actually loves." Rated not toxic. All this API is going to do is promote a renaissance of polite burns.
- tracker1 5y agoThe false negative example you have us exactly the argument I've used to fight against language filters and sensors in smaller sites. In the end you can be profane and very positive. You can also have strike language while being incredibly vile and negative.
- Gibbon1 5y ago> In the end you can be profane and very positive. I've noticed a lot of humans can't detect sarcasm. And that isn't correlated with traditional smarts either. For AI it's a hopeless task. > You can also have strike language while being incredibly vile and negative. I'm reminds of a group of people I know. They love love using genteel language to throw vile insults at each other. Using profane language is a automatic foul.
- skissane 5y agoYou can also write things which sound violent (when read literally) but are actually innocent. My wife's cousin recently had a baby, and posted a photo on Facebook. My wife commented (something like) "She's so cute I'll have to kidnap her". Facebook locked her account for 24 hours for making a "violent threat". I am sure the mother wasn't threatened by the remark in the slightest. It just added to my wife's anger at Facebook for repeatedly giving her "warnings" over trivial or innocent things. Not online, in person, but one of the teachers at our son's school sometimes "threatens" to "steal him" from us – her remarks don't worry us, because we know she would never actually do that, it is just a colloquial way of expressing affection.
- tobyhinloopen 5y agoWhat if a random old guy she didn’t know commented that exact message? I don’t think the mother would be as accepting then, so it also matters who posts the comment to who
- skissane 5y agoTo Facebook they are “friends” and have many mutual “friends” as well. You’d think they could take that into account in evaluating the nature of the remark, but it does not appear they have done so.
- DoItToMe81 5y agoEven if totally 'accurate' to the dataset, there's the issue that a lot of 'toxicity' and offence is completely culturally based. Take 'cunt' for example. Used in casual, informal conversation here, but deemed incredibly offensive by some Americans. Or, more broadly, 'thumbs up' could mean "Good job" or "Stick it up your arse".
- echen 5y agoExactly. This is why it's important not just to have language skills when creating these kinds of datasets, but also cultural knowledge and context. For example, to pick a bit on the Google Emotions dataset again, it's difficult to label this message... “Also Republicanism is a belief system. It’s taught and handed down like religion. Conservative talk radio is its evangelism.” ...unless you're familiar with US politics. Hence why it was labeled as APPROVAL by the non-US annotators, even though it's criticizing Republicans.
- wodenokoto 5y agoA few weeks ago someone posted a guide to interviewing as a data scientist, and while it touches on all sorts of algorithms and statistical relations, the number one thing you want out of a data scientist is respect for the source data. Second is understanding what you want to do. Modelling comes third. > "this is fucking awesome!" as toxic, even though they're actually quite positive And this touches on the question of what we want to do. It might not be a negative sentiment, but it might still be offensive. On the other hand, there are no offensive words, and only "positive" sentences in many sarcastic utterances. And there are a lot of sentences that can be perceived as insensitive, that are not meant as such. GCP Grey has a video on the words "Indian" vs "Native American", and apparently it's very complicated which word you can or should use. https://www.cgpgrey.com/blog/indian-or-native-american-reservations-part-0 https://www.cgpgrey.com/blog/indian-or-native-american-reser...
- fho 5y agoOne thing I came across is that 99% precision/recall/F1 score might not be enough. Especially in large sample sets "only" accurately identifying 99% leaves you with too many wrong classifications. Eg 1.000.000 samples -> 10.000 false classifications. No idea if there is a different metric that somehow (again) takes the number of sample into account.