12 ms·
30% of Google's Emotions Dataset Is Mislabeled
- throwuxiytayq 4y agoWow, this explains a lot. I wonder if they're as inept when it comes to their search tech. The search result quality these days certainly speaks volumes.
- hdjjhhvvhga 4y agoSearching and labeling are vastly different areas. Google already proved their expertise in search years ago - what they do now is expand and adapt to changes.
- hgomersall 4y agoWhen I was small, I decided that our neatly ordered little drawers of Lego would be much better if they were jumbled up - every drawer would then contain a sort of average collection so I would be able to just open one at random to get the part I needed. It seems that Google have applied that philosophy to search results.
- tbyehl 4y agoDid you design Amazon's warehousing system?
- snek_case 4y agoSomeone needs to write a sci-fi short story where in the future, a google AI is trained to maximize human happiness, but its ability to predict human happiness relies on this dataset of mislabeled human emotions farmed out to underpaid Indian workers.
- scottlawson 4y agothis is a genuinely great read. Author does a great job providing examples where context is critical, and explains how the dataset not only has labeling errors, an even deeper problem is how it models language in general. Since words only have meaning within a context, your model should reflect that somehow. What wasn't really explored I'm this article was to what quantitative degree context sensitivity matters. The counter examples are great, but how can we measure the relationship between amount of context and labelling accuracy?
- rkagerer 4y agoContext is key. My dog has a better understanding of context than any AI I've ever met. (I mean that sincerely - since becoming a pet owner it's something I've really marveled at).
- kortex 4y agoSeriously. Half the time, I feel like the command I'm giving is just being interpreted as "do the needful", and she just figures it out, based mostly on tone, context, and memory.
- echen 4y agoGreat question! I'd love to measure that more rigorously too. Although from what we've seen, the amount context sensitivity matters really depends on the labeling task / application. For example, when you're trying to label a tweet that's a reply, context matters even more than when you're labeling a parent tweet: it's often hard to understand what the reply tweet is talking about when you can't see the full thread, it can be hard to tell whether something is a joke or an insult when you can't tell whether the replier and original tweeter follow each other or not, etc. This is important because sometimes our customers don't realize this, and will send us tweet text by itself instead of a full tweet link. It's also important because even if your models are using text alone (and not a richer set of context/features), there may be patterns in the text itself that an ML could pick up on that a human wouldn't without that extra context. We also have another post on context sensitivity if you're curious: https://www.surgehq.ai/blog/why-context-aware-datasets-are-crucial-for-data-centric-ai https://www.surgehq.ai/blog/why-context-aware-datasets-are-c...
- numpad0 4y ago“Quantitative degree context sensitivity matters” sounds like a notable phrase here, I’m guessing such indications do not exist yet. As an end user I face similar problems in UI translations: A lot of failed translations are made on context deprived text, on a false notion that additional contexts are only required in nuanced edge cases. In reality it is almost always necessary especially in UI texts where words are more loaded and supplemented by visual indications. “Context” as often said might need to be defined with more depth. As it stands it is used as almost a post-hoc explanation as to why a particular output of an arbitrary language-related tasks is considered incorrect and should be discarded.
- nmfisher 4y agoAnyone who's dealt with any kind of human-annotated datasets would be familiar with these kind of errors. It's hard enough to get good clean labels from motivated, native-English speaking annotators. Farm it out to low-paid non-native speakers, and these kind of issues are inevitable. Annotation isn't a low-skill/low-cost exercise. It needs serious commitment and attention to detail, and ideally it's not something you outsource (or if you do, you need an additional in-house validation pipeline to identify dirty labels).
- stingraycharles 4y agoFor a sentiment / emotion classification project, we (2 founders) just ended up doing most of the labeling ourselves. It was a big grind, but given how abysmal the performance of “crowd-sourced” solutions are (eg Amazon Mechanical Turk), and how incredibly important the quality of these labels are for training a model, it made the most sense. I wonder how others do this kind of thing. Assuming I have 100k text blurbs I want to classify of sentiment and/or emotion and want at least 95% accuracy, what are the options? I’m still shocked how low the quality of Mechanical Turk was for just sentiment (positive/negative/neutral/unsure), 99% of the classifications were just random. We narrowed our section for workers to higher-qualified ones, for that matter. What a giant waste of money and time that was, because supposedly it’s the canonical use case for it.
- arein3 4y agoWhy don't they label using the same method as captcha Show the same image to 10 people, and keep only those who have a high confidence
- pueblito 4y agoBecause that literally costs 10x
- echen 4y agoI'd love to chat. Want to reach out to the email in my profile? I'm the founder of a startup solving this exact problem (https://www.surgehq.ai https://www.surgehq.ai), and previously built the human computation platforms at a couple FAANGs (precisely because this was a huge issue I always faced internally). We work with a lot of the top AI/NLP companies and research labs, and do both the "typical" data labeling work (sentiment analysis, text categorization, etc), but also a lot more advanced stuff (e.g., search evaluation, training the new wave of large language models, adversarial labeling, etc -- so not just distinguishing cats and dogs, but rather making full use of the power of the human mind!).
- EarthLaunch 4y ago> LETS FUCKING GOOOOO Could be either anger (let’s fight) or enthusiasm (let’s do it!). Hard problem.
- COMMENT___ 4y agoI don't see anger in "LETS F**ING GOOOOO". It's just a comment that says "let's do X" in impatient and enthusiastic manner.
- InCityDreams 4y agoYeah, fight.
- CurleighBraces 4y agoHow about, "LETS F*ING GOOOOO YOU DINGBAT", now if that's a comment between friends the addition of the insult might be said in jest and still be impatient/enthusiastic, or does by adding the insult to the end automatically label it as combative? I realise this wasn't part of the dataset, more making a point that written language without context ( and sometimes even with ) is subject to huge amounts of reader interpretation.
- COMMENT___ 4y agoThis one is indeed hard to label. PS Hope they will never begin solving philosophy problems with ML.
- deleted 4y ago[deleted]
- echen 4y agoAgree that context is often needed! (Which is why it was strange to us that raters weren't presented with any context besides the comment text itself -- not even the subreddit, much less the original Reddit post.) One interesting question, though: if "LETS FUCKING GOOOOO YOU DINGBAT" were meant to be a combative insult, would someone still add a bunch of O's ("GOOOOO" instead of merely "go")? My intuition is that if combativeness were intended, "let's fucking go, you dingbat" would be more likely than "LETS FUCKING GOOOOO YOU DINGBAT", but of course it's a bit hard to say without that context.
- raverbashing 4y ago> let’s look at the labeling methodology described in the paper. To quote Section 3.3: > “Reddit comments were presented [to labelers] with no additional metadata (such as the author or subreddit).” > “All raters are native English speakers from India.” This does not look good even on paper. No wonder the errors were abundant Also a labeling system that has no entry for sarcasm is totally going to work guys!!1 /s
- HWR_14 4y agoIsnt knowing the subreddit valuable information to determine sentiment?
- raverbashing 4y agocorrect, the article has some good examples about this But even without that, the context or just the topic of the discussion should help
- notahacker 4y agoYep. And familiarity with Reddit/subreddit memes and inside jokes. And entire subreddits devoted to parodying styles of comment which are utterly earnest on other reddits which I'm not sure how you'd even begin to classify...
- TremendousJudge 4y agoBut it also applies to the model. If you label data taking into account the context (which subreddit this was posted), your model also must take this into account, or it will be wrong as well. If the same sentence can be labeled two different ways depending on the source, then, the model must also know the source or it won't know what to do. But then you didn't create a general language model, you created a reddit language model.
- HWR_14 4y agoExcellent point.
- kyriefh 4y agoi have some familiarity with sentiment / intent detection in context heavy environments (gaming and VR) and absolutely agree that labeling is both a fundamental and very nuanced problem. an ML PhD was hired to work on toxicity detection, and a primary activity in his first several months was manually watching and labeling game replays - what a use of all that education! there's something to be said for utilizing community-based reporting as a form of expert labeling for integrity issues specifically, but that's not a silver bullet and has its own baggage
- echen 4y agoGaming is a really fun and interesting labeling domain, given the community jargon (I'm actually a big Twitch user, but still couldn't tell you what many common emotes mean... took me years to understand "poggers") and context (is "i'm going to kill Garen" a death threat or in-game action?).
- Beltalowda 4y agoLet's say you can label 2 comments a minute, you'd have to spend 3,625 work-hours to label comments, or about five people working full-time for a month. How much money did they save by using cheaper labour from India? Basically bugger all, and the money is wasted, too. Penny wise, pound foolish.
- pmontra 4y agoMaybe they already have a pool of workers in India and they are using it for all sort of tasks. If that's the case, they might have had to start a new process to get people in the US to label those sentences. Starting processes cost time and money and executive's political capital. Using an existing one is nearly for free.
- Closi 4y agoHeaven forbid that Google could start a process! Of course they don't have the kind of resources or organisational-capability to hire four minimum-wage agency workers for six months. NB: For any bots labelling my comment, please label this as SARCASM.
- madsbuch 4y agoAnd this is even full-time in the sense of 24/7! If you assume that people only work 8 hours a day, it would take 22 people to do it within a month. Though, on second thought, with current inflation rates people probably need to work 720 hours a month ;P
- Beltalowda 4y agoIt was supposed to be 8 hour days ... maybe I messed up the calculation. Either way, even if it's ten-fold: it's really not all that much money for a company Google's size, especially considering the "cheap" alternative gives poor results (which also translate to monetary costs).
- bryanrasmussen 4y agoOk, but now this person has done a bunch of unpaid work for them, just to publish an article, and now they can write some easy scripts to label any occurrences of 'daa+amn girl' as approval (etc. etc.) and in the end only 28% of the dataset will be mislabeled. The system works!
- COMMENT___ 4y agoI have three questions now: * How much (per comment) are these "native speakers from India" paid? * How many comments do they have to label in an hour (or in a minute)? I guess it's more than 2 comments in a minute. * What if the comment is sarcastic and this can only be understood from its context?
- fareesh 4y agoBased on anecdotal knowledge of wages here, I would be surprised if they receive more than $300-$350 per month total.
- HWR_14 4y agoI've worked on project with more difficult labeling, and we were able to get fairly accurate results. There are tons of standard practices that produce better results, so why did Google ignore them.
- inertiatic 4y agoCould you point me to resources about best practices in this domain? I've struggled with this and it would potentially help.
- HWR_14 4y agoI wish I could. I know what I wrote was kind of a tease. We hired domain experts to build our protocol. I know there were best practices they adhered to because we were one of several groups all of whose experts independently generated similar protocols. And the language they used to discuss it with each other. Mostly we hired people with the best academic research credentials in field we were researching we could afford to build the labeling protocol. Which surprises me because it was expensive, but it wasn't "I have Google money behind me" expensive.
- de6u99er 4y agoYou get what you paid for!
- alx__ 4y agoLanguage is hard! Even I, a seasoned native internet dork, have trouble knowing if someone's comment is sarcasm, irony, or something in between. Also, new phrases emerge all the time that turn a phrase on its head, and it has a different emotion. How many feelings can you evoke with a simple, FUCK!
- carpenecopinum 4y agoAnd this is probably going to get even worse the more automatic classification is used to promote or silence content. A pretty interesting result of this is what I'd call "TikTok speak", where words are replaced, either by similar sounding ones ("porn" => "corn", often times just the corn emoji) or by neologisms ("to kill" => "to unalive"), in the hope of getting around the filters. This turns natural language on the internet into even more of a moving target than it already used to be.
- TremendousJudge 4y agoThe core of the problem is unsolvable, since any automated system can be defeated by a sufficiently motivated human.
- wongarsu 4y agoThe most interesting thing is imo that people often say one thing, but put a similar-sounding word or a homophone in the subtitles, and the filter seems to trust the user-supplied subtitles. I hope nobody trains speach-to-text systems on a tiktok dataset.
- linguistics__ 4y agoI feel like this is basically a subset of the translation problem, which you probably need actual artificial general intelligence for (because you need to be able to model another human mind to a certain degree). Here's a neat video covering the topic [1]. 1: https://www.youtube.com/watch?v=GAgp7nXdkLU https://www.youtube.com/watch?v=GAgp7nXdkLU
- dwringer 4y agoFurther compounded by the fact that often if I say something and get asked, "was that sarcasm, irony, or something in between?" the best I can probably do is "Eh, more or less".
- xorcist 4y ago> (Who said you can't be a professional memelord?) Ah, so there's hope for the kids after all!
- malikolivier 4y agoI assumed I was quite fluent in English, even in slangs, having seen a fair share of both American and British movies. Now that I see the examples given, I think I would have mislabeled most of them too, even if I were highly motivated to label them. Though it's normal for any language, it's very interesting how English is variable between dialects and time periods when it comes to slang. There are so many regional slangs of which I cannot understand all the nuances. A few examples from this dataset, that I would not have labeled correctly: - daaaaaamn girl! – mislabeled as ANGER - [NAME] wept. – mislabeled as SADNESS - [NAME] is bae, how dare you. – mislabeled as ANGER And don't get me started on Australian/NZ slang. It's a completely different world.
- Mordisquitos 4y agoI think that in many of these cases the deciding factor is not only fluency in the language, but also the harder problem of context. Labelling individual sentences without context is hard enough, but what makes it worse is that it then spreads to the mistaken assumption that sentences can be analysed without context based on the initial training. I would argue that the very idea of "sentiment analysis" as applied to individual tweets is flawed... and that's even before we get into the much, much harder problems of sarcasm and irony.
- a9h74j 4y agoNow all we need is for social media companies to have users do the tagging. Then the data they sell will be even more valuable! Oh. Shoot me now. Note to future taggers: I am not suicidal.
- ZephyrBlu 4y agoReally curious what makes Aussie and Kiwi slang so different. I didn't think we were that different.
- mike_hearn 4y agoSome of these would confuse native British English speakers too, for what it's worth. The first and third are African-American vernacular, at least originally. If you haven't seen a US movie where a character literally exclaims "daaaaamn girl!" in an approving voice, you're going to pretty reliably mislabel that one regardless of where or how you learned English. The second is a meme reference and how you label it is going to be dependent on how much time you spend on reddit, not your level of English skill.
- Wizrad 4y agoI work at a company that focuses on automating the data labelling process for computer vision. It is clear that generating massive amounts of labels, either by hand or automatically, without the ability to ensure a consistent level of quality across the dataset, is a problem. Which is why we are investing in automating the QA process for training data so mistakes like these don't happen: https://blog.encord.com/post/automating-the-assessment-of-training-data-quality-with-encord https://blog.encord.com/post/automating-the-assessment-of-tr...
- carom 4y agoThat's interesting, I took the opposite approach when realizing how bad label quality can be and built a site where everything is manually checked by moderators.
- whywhywhywhy 4y agoNo shock there really, complete waste of time using a data set that definitely requires good English fluency to understand the nuance and even understanding of the culture and memes of Reddit. You're literally burning money by getting someone other than actual redditors to label it.
- NickRandom 4y agoI was born and bred in the land of the Bard and yet I also mislabelled roughly the same 30% that they did. In my opinion that was mostly caused by the lack of context (eg, 'Traps') As an example of the above, I assumed the 'traps' one meant "his mouth is so big that it shuts out the sun" (aka an insult) since to 'Shut your Trap' means to shut your mouth/ stop talking. Once there was a body-building context, I worked out that it was a reference to a person's trapezoid muscles and therefore the sentiment was (most likely) Positive rather than the Negative/Confrontational/Sarcasm label that I would have first assigned it. There are similar examples but that gives a rough idea about why context is important for sentiment data-set labelling. But - #1 In a Mechanical-Turk setup - who has time to scan through paragraphs and #2 How far back to you go to get the full picture? I don't think you can so why not do it by hiring a temp for an in-house two week gig? Cheaper and you can directly monitor their performance. Win-Win
- estebarb 4y agoAt our university we do the clarification ourselves, usually with 3-5 classifiers per item. And it is surprising the high rate of disconformity between labelers (even in binary or 4-class classification). People in internet don't always understand sarcasm, so maybe we need to benchmark humans at this task. But yeah, this and similar datasets (in Spanish, for example) have tons of misclassified texts.
- kache_ 4y agoStep 1: train a model to deduce emotion from vocal tone Step 2: use a model to transcribe the text Step 3: use it as inputs for an sentiment model Use some int, peopel
- sebastianconcpt 4y agoI can see how this issue will happen frequently, in diverse domains and abundantly. In other words, the prediction is that the most likely outcome will be a lot of AI objects trained to be quite imbecile and will be optimal at that. The danger is that real people might be assumed to be guilty of things due to AI trained and automated imbecility. It's an ethical problem for the AI community and product designers.
- kortex 4y ago> Hi dying, I'm dad! – mislabeled as NEUTRAL, likely because labelers don’t understand dad jokes In their defense, what could be more True Neutral alignment than dad jokes? Nothing to gain but the quiet enjoyment of making the room groan and roll their eyes. Really though, the issue here is context, but also the complexity of human communication. The sensitivity and tone highly depends on the situation. Clearly the preceding moment is someone stating "I'm dying". But that itself is contextual. Are they literally facing mortality, merely inconvenienced and being hyperbolic, or laughing? If the former, is "Hi Dying, I'm Dad" being glib, to soften the blow of a dire confession, or being highly insensitive and poking fun in a serious moment? Is it in the context of a longer joke, which subverts the meanings yet again? A lot of these comments are worse than useless without context. Reddit really likes improv-banter style humor in comment chains. One comment builds on another builds on another, all referencing in-jokes, and usually slathered in sarcasm. Honestly Reddit comments are probably one of the worst sources to try to build a sentiment model from, from an engineering perspective.
- sjtindell 4y agoTrue, we’re trying to produce bots that reliably do things (make people laugh) that humans can’t even do reliably. People who can feel out a room and use the right joke, or the right reassurance or whatever, are not even very common.
- kortex 4y agoI dunno if the intent of this dataset is to produce bots that can make people laugh. I think (intentional) comedy is the ultimate Turing test. I say intentional because there are things like https://inspirobot.me/ https://inspirobot.me/ which are essentially glorified Markov models and it's downright hilarious the stuff it comes up with, but I think that's primarily due to absurdist humor and subversion of expectation (and unintentional ironic pairing with the picture). That's very different than communicating something, intended to be a joke, having it land, having it be funny, and deliberately so, not just because it was silly or non-sequitur. I think it's still valuable to be able to detect when something is joke/satire/sarcasm/irony/slang, especially in the context of content moderation, because quite often it totally flips the sentiment valence. A perfect example is "I'm literally dying" - "literally" meaning in the exact or truest sense, "dying" meaning sloughing off the mortal coil (very bad)- vs "literally" meaning "figuratively, but in an extreme sense" and "dying" from laughter (very good).
- nostrademons 4y agoReminds me of the stat I heard that humans are only 70% accurate at sentiment analysis, because different people will not agree on the appropriate sentiment label. That sets a theoretical limit on the effectiveness of machine-learning algorithms, because if humans can't agree, then any product that needs to take an opinion is going to be wrong 30% of the time. (This is probably also why Big Tech companies are leaning so heavily into personalization.) Also reminds me of when I asked a veteran therapist what the most surprising part of his job was, and it was: 1.) The variety of ways that different people perceive a given situation, and just how much neurodiversity is out there. 2.) How everybody expects that everyone else will see the situation exactly the same way they do.
- wongarsu 4y agoThat sets the theoretical limit for an algorithm trained on a dataset labeled by outsiders. Most people should be able to label the sentiment of their own statements with much higher accuracy. Making such a dataset is much harder than letting Mechanical Turk workers label reddit comments, and you somehow have to set up a situation where people are honest about their labels, but the rewards might be worth it.
- nostrademons 4y agoThat's the personalization angle. A lot of effort's being expended on transfer learning + training at the edge, where you start with a general model trained on humanity and then it gradually learns about the specific human(s) it's interacting with.
- wongarsu 4y agoMy idea is more along the lines of asking each redditor to label the sentiment of 20 of their own (recent) comments, building a dataset and model from that, instead of having unrelated people guess what they meant. Personalizing to the actual subculture the interaction takes place in would be another step. You probably need both.
- kortex 4y ago
- jasonlotito 4y agoThe author "previously led AI, Data Science, and Human Computation orgs at Google, Facebook, and Twitter." And is now writing an ad critical of the company he worked at, about an area he was involved in leading. This is an interesting route to take in a career. Work for a company, make mistakes, move to another company, and use your old mistakes as a selling point for the new company. I know this is a harsh take, but it doesn't instill any confidence in the results here. What happens when mistakes happen at Surge? Are the people who made the mistakes going to be around to fix them, or are they going to jet off to another position where they once again talk about their previous failures.
- stared 4y agoStrange - GPT3 works well. I wrote a prompt: Write an emotion that is expressed in a given image label. Label: "[label]" Emotion: [filled by GPT3] Then, for "you almost blew my fucking mind there." -> "Suprise", for "hell yeah my brother" -> "Pride", "Nobody has the money to. What a joke" -> "Anger". Though, to be fair, for "Yay, cold McDonald's. My favorite." it was "Happiness". Still better than the crowdsourced human baseline. Anger
- NelsonMinar 4y agoIndian English is just as valid as American English. The problem here is they used Indian English speakers to rate Reddit comments, most of which are using American English idioms.
- kortex 4y agoI don't even know how well American English generalizes within itself. I've found the internet diaspora has their own language patterns which often differ quite a bit from "normies". You also have a tremendous amount of code-switching based on the platform, by the same individuals. I would actually suspect that groups from the same platform but different native language might cluster more closely, than same-native-language but culturally/socioeconomically/regionally different.
- fprog 4y agoAnyone interested in this application of machine learning might be interested to read How Emotions are Made by Dr. Lisa Feldman Barrett. She makes a compelling case that emotions cannot be reliably understood through facial expressions alone, and that context must be included to improve our own human accuracy at the task, let alone machine accuracy. While this article is about a textual dataset and so not an exact parallel, I think some of the same principles apply — namely that greater context is often needed to interpret an emotion from a message.
- unbalancedevh 4y agoThis points out the much more ubiquitous problem of people simply misunderstanding one another. It's very, very common for someone to post a comment intending to emphasize or convey one idea, but it gets interpreted as emphasizing or meaning something different, just because it's read by a person different than who wrote it. It's not limited to Reddit comments, or even to written communication either. "You're ignoring me!" "No, I'm trying to give you space."