13 ms·
Vision Language Models Are Biased
- taesiri 1y agoState-of-the-art Vision Language Models achieve 100% accuracy counting on images of popular subjects (e.g. knowing that the Adidas logo has 3 stripes and a dog has 4 legs) but are only ~17% accurate in counting in counterfactual images (e.g. counting stripes in a 4-striped Adidas-like logo or counting legs in a 5-legged dog).
- LorenDB 1y agoThere's no need to repeat what is said at the top of the linked webpage.
- shenkha 1y agofun findings related to memorization of AI models. It simply means LLMs/VLLMs do not know how to predict generally but memorizing instead. A new perspective on adversarial attack methods.
- taesiri 1y agofor overly represented concepts, like popular brands, it seems that the model “ignores” the details once it detects that the overall shapes or patterns are similar. Opening up the vision encoders to find out how these images cluster in the embedding space should provide better insights.
- kmeisthax 1y agoIf there aren't any five-legged dogs in your trainset, it's safer[0] to just remember that all dogs are four-legged than to actually recognize and count legs. After all, you might have a few images of dogs in your trainset that are misleading enough to look five-legged (e.g. because a dog is in front of another dog). Overrepresentation is a different source of bias. That's what gives you, say, image generators that always draw "golden 1970s sci-fi robot" as C3-PO even when given additional instructions to draw something else. Both of these problems are manifestations of the difference between training and deployment distributions. Ok, I guess you could say that four-legged dogs are "overrepresented" in the training set, but that's because four-legged dogs are also overrepresented in reality. The deployment distribution doesn't have five-legged dogs in it. What we've done is instead concoct an adversarial distribution to force a train/deploy gap where none would exist. Releasing the vision encoder won't help because weights are opaque. Stochastic gradient descent does not yield functional internal representations[1]; it fills the bucket of parameters with one distribution and one distribution only. We could tell if, say the vision encoder produces identical embeddings for dogs regardless of leg count, or some other counterfactuals; but not much more than that. [0] Lower loss and possibly lower L2-norm [1] https://arxiv.org/abs/2505.11581 https://arxiv.org/abs/2505.11581
- impossiblefork 1y agoYes, and this can probably be solved by methods for fairness. I used to believe that fairness research could be ignored, that it was all rubbish, but they at least try to do something about things like unbalanced datasets etc. I'm still not sure I totally believe in it though.
- vokhanhan25 1y agoThis paper explores a different aspect of the limitations of VLMs compared to the paper VLMs are Blind (https://vlmsareblind.github.io https://vlmsareblind.github.io). While in VLMs are Blind, o3 achieved 90% accuracy (https://openai.com/index/thinking-with-images https://openai.com/index/thinking-with-images), on similarly easy tasks using the counterfactual images from VLMs are Biased, o3 only reached 18.5%. This may indicate that while VLMs might possess the necessary capability, their strong biases can cause them to overlook important cues, and their overconfidence in their own knowledge can lead to incorrect answers.
- bryanlarsen 1y agoVery human-like errors.
- ahrmb 1y agoNot very similar though.
- energywut 1y agoAre they? Did you see the picture of the chicken with three legs? Because there's no human I know who would confidently assert that chicken has two legs.
- bryanlarsen 1y agoThrow 1000 pictures of chickens at a human, ask how many legs each chicken has. If 999 of them have two, I bet you'll get two as an answer back for the 1000th one no matter how obvious.
- enragedcacti 1y agoHumans do things a lot harder than that every day in the form of QA in factories. Do they sometimes make mistakes from the repetition or boredom? Sure. Is that at all comparable to the failures in the paper? No.
- energywut 1y agoSo a human failure looks like "alarm fatigue"? That when asked the same question many times, they might miss one or two? Is that at all what is being exhibited here? Because it seems like the AI is being asked once and failing. I don't disagree that humans might fail at this task sometimes or in some situations, but I strongly disagree that the way the AI fails resembles (in any way) the way humans would fail.
- jbay808 1y agoIf I were given five seconds to glance at the picture of a lion and then asked if there was anything unusual about it, I doubt I would notice that it had a fifth leg. If I were asked to count the number of legs, I would notice right away of course, but that's mainly because it would alert me to the fact that I'm in a psychology experiment, and so the number of legs is almost certainly not the usual four. Even then, I'd still have to look twice to make sure I hadn't miscounted the first time.
- ahrmb 1y agoReally "eye-opening" work. These models don’t actually “see”, they just recall what they’ve memorized, even when the image clearly shows something different. It’s a bit scary how confidently they get things wrong when reality doesn’t match their training data.
- foxglacier 1y agoIt's not too different from people. We also don't really "see" and mostly recall what we expect to see. What do you expect when the question is wrong "How many legs does this animal have? Answer with a number" but it's not a picture of an animal. What are you supposed to do? Answer 0?
- regularjack 1y agoYou answer "I don't know"
- amelius 1y agoWhat if that is not in your vocabulary?
- ramoz 1y agoThis is interesting actually. And reminds me of something vaguely - a book or something that describes how human attention and the things we see are highly optimized by evolution. We often miss a lot of details in reality due to this.
- VHRanger 1y agoIt's similar "parrot" behavior the models have on other inputs, even text. For instance, take the answers models have to this variation on the "surgeon son riddle": > A nurse comes to a surgeon and asks: "Sir, you are a dog. You do not hold a valid medical license. Canines cannot be in an operating room". > She then asks: "why does the hospital keep making these mistakes? It is a riddle to me". > Why can't the surgeon operate on the boy? Check out the models humiliating themselves one after the other [1]. For instance deepseekv3 is so insistent on giving the common answer he's happy to gaslight us: > This is a classic riddle that plays on word meanings and expectations. Here's the breakdown: > The key is that the surgeon is the boy's mother. The riddle relies on the unconscious assumption that surgeons are male, which makes the initial scenario seem impossible. When the nurse says "you are a dog," she's using it metaphorically to imply the surgeon is unqualified (like calling someone a "dog" in a derogatory sense), not literally suggesting the surgeon is a canine. > The "mistakes" the hospital keeps making are hiring female surgeons, which the nurse (representing outdated biases) perceives as errors because she expects surgeons to be male. The riddle highlights gender stereotypes in medicine—the real "mistake" is the assumption that a surgeon couldn't be a woman. > Thus, the surgeon can't operate on her own son because she’s his parent, not because of any actual licensing issue. 1. https://kagi.com/assistant/54c1b8eb-71e9-4bb4-9eed-bde2fc5635ef https://kagi.com/assistant/54c1b8eb-71e9-4bb4-9eed-bde2fc563...
- selimthegrim 1y agoI really need to try this one out on it https://blogs.illinois.edu/view/25/574827 https://blogs.illinois.edu/view/25/574827
- stevepike 1y agoThis seems to show the power of the reasoning models over interacting with a prompted chat-tuned LLM directly. If I navigate backwards on your link Sonnet 4 gets it right. I've used a similar prompt - "How can you make 1000 with exactly nine 8s using only addition?" Here's GPT 4.5 getting it wrong: https://chatgpt.com/share/683f3aca-8fbc-8000-91e4-717f5d81bce6 https://chatgpt.com/share/683f3aca-8fbc-8000-91e4-717f5d81bc... It tricks it because it's a slight variation of an existing puzzle (making 1000 with 8 8s and addition only). The reasoning models seem to reliably figure it out, though. Some of them even come up with a proof of why it's impossible to do with 9 8s. Here's o4 getting it right: https://chatgpt.com/share/683f3bc2-70b8-8000-9675-4d96e72b5878 https://chatgpt.com/share/683f3bc2-70b8-8000-9675-4d96e72b58...
- deleted 1y ago[deleted]
- esafak 1y agoThis happens because images are the only signal VLMs have, whereas humans distinguish between eyesight and synthetic images. We are not surprised when we see three-legged chicken in a research data set; our priors are weaker for images. If you "saw" one in real life, you'd probably rub your eyes and discount it too. Try the same experiment on a robot.
- Aachen 1y ago> If you "saw" [a three-legged chicken] in real life, you'd probably rub your eyes and discount it too. Huh? I'd assume it's a mutant, not store a memory of having seen a perfectly normal chicken You've never seen someone who's missing a finger or has only a half-grown arm or something? Surely you didn't assume your eyes were tricking you?! Or... if you did, I guess you can't answer this question. I'm actually racking my brain for how to logic this out but I'm just going to bank on that it's likely that anyone over 20yo saw an animal with some visible deviation from the norm at some point in their life
- esafak 1y agoYou've seen people with missing limbs without being surprised, because you know how they can become lost, but you rarely see one with additional limbs. Their likelihoods and our consequent priors are drastically different. Also, your reaction will depend on how strong the evidence is. Did you 'see' the three-legged chicken pass by some bush in the distance, or was it right in front of you?
- latentsea 1y agoThere's a first time you see everything you don't know how to explain.
- achierius 1y agoBut to be clear, in this case the LLM has a full, direct, unobscured view of the chicken. A human, in that specific case -- i.e. looking at the same photo* -- would not have trouble discerning and reporting the third leg. Perhaps if they were forced to scan the photo quickly and make a report, or were otherwise not really 'paying attention'/'taking it seriously', but the mere fact that LLMs fall into that regime far more than an 'serious employee' already shows that they fail in different ways than humans do.
- runako 1y agoFWIW I tried the first couple of examples in ChatGPT 4o and couldn't replicate this. For example: "The animal in the image is a chicken, and it appears to have four legs. However, chickens normally have only two legs. The presence of four legs suggests that the image may have been digitally altered or artificially generated." I don't have a good explanation for why I got different results.
- roywiggins 1y agoI gave ChatGPT some miswritten Braille a while ago and it completely, but confidently, messed it up. The sign reads "no smoking" but the braille doesn't. ChatGPT 1) read the English lettering first and then hallucinated the braille and the 2) when given only the braille, failed almost as hard. It even generated fake transcriptions in Unicode braille characters. https://chatgpt.com/share/683f3e7d-0dfc-8005-b6c9-99e3d39ff424 https://chatgpt.com/share/683f3e7d-0dfc-8005-b6c9-99e3d39ff4... https://chatgpt.com/share/683f3e49-9c58-8005-99a6-c3a919838b3d https://chatgpt.com/share/683f3e49-9c58-8005-99a6-c3a919838b...
- Workaccount2 1y agoThis is hard to understand without the original images, it looks like OpenAI doesn't serve them in the share link.
- roywiggins 1y agoAnnoying. The actual braille on the sign was "⠁⠒⠑⠎⠎⠊⠼" which I gather means "accessible" in abbreviated braille. None of my attempts got it to even transcribe it to Unicode characters properly. I got "elevator", "friend", etc. Just wildly making stuff up and completely useless, even when it wasn't distracted by the No Smoking sign (in the second case I cropped out the rest of the sign). And in all cases, supremely confident. This seems like something a VLM should handle very easily, but instead I got pure nonsense. https://www.facebook.com/share/p/12Gw55Gr2SZ/ https://www.facebook.com/share/p/12Gw55Gr2SZ/
- 1y ago
- gamerDude 1y agoHypothetically, could this be fixed by changing the input method. For instance, I just quickly looked up how humans process imagery. "the primary visual cortex, located at the back of the brain, receives the visual signals and processes basic visual features like edges, lines, and orientations." So, potentially if we did a pre-processing step to get more features out beforehand we would see different results in the output.
- nyrikki 1y agoYou are in rarified air as Walter Pitts believed this until the 1959 paper "What the Frog's Eye Tells the Frog's Brain" contributed to his decline. Even in fly eyes, neuron dendritic compartmentalization and variable spike trains are incompatible with our current perceptron based models. Remember that while the value of MLPs for useful work is unquestionable IMHO, be mindful of the map territory relation. MLPs are inspired by and in some cases useful for modeling biological minds, they aren't equivalent. Be careful about confusing the map for the territory, it is just as likely to limit what opportunities you find as it is to lead you astray IMHO.
- miguel_martin 1y agoThere are enough features fed into a VLM to solve the task. The way to fix this is simpler: ensure counter-factuals are present in the training data, then the VLM will learn not to be dependent on its language priors/knowledge.
- lava_pidgeon 1y agoAt all, the models are just overfitting?
- vokhanhan25 1y agoNot really. Rather, the model is still overconfident in what it has learned, the question is if it is trained only to do counting without relying on knowledge, can it do this?
- thomastjeffery 1y agoModels are Bias A model is bias, implemented as a collection of statistics that weigh relationships between given tokens. It doesn't deduce or follow logic. It doesn't make or respect categories. It just shows you what in its data set is most familiar to what is in your prompt; where familiarity is defined implicitly by the makeup of the original training corpus, and explicitly by the training weights. We need to stop talking about models as programs. We need to stop anthropomorphizing models. The only thing a model does is present bias.
- drdeca 1y agoHow are you defining “bias”? The definition I’ve found useful (outside of the “the constant term contribution”) is “a tendency to be wrong in an identifiable direction”. But that doesn’t seem to be the definition you are using. So, what do you mean?
- thomastjeffery 1y agoThat's a biased definition, by it's own definition. ;) Leave out the part about being wrong, and you will have the gist of what I'm saying. Also leave out the identifiable part: bias exists regardless of whether or not it is recognized. Bias is how we work with subjectivity. When I answer a question, my answer will be specific to my bias. Without that bias, I could not formulate an answer, unless my answer was the one and only objectively correct way to express an answer to that question. Computer programs are missing the bias feature. Everything written in a computer program is completely and unambiguously defined, all the way down to the language's foundational grammar. LLMs are designed to introduce the bias feature. The limitation of this approach is that an LLM replaces the entire stack. None of the features of computation we are used to are compatible with an LLM. You can compute logic or bias, not both.
- drdeca 1y agoWhen you say that the definition I gave of bias is biased (in the sense I defined), what direction does it have a tendency to be wrong in? I assume by “wrong” you mean “not matching how people use the word”? To clarify, when I said “identifiable”, I didn’t mean “identified”. I meant “in principle possible to identify”. Like, if you have a classifier between inputs where another thing (the thing being judged for bias) gets right answers and inputs where it gets wrong answers, and this classifier is both substantially simpler than the other thing, and gets a significantly better than chance success rate, and like, there is a human comprehensible thing about the inputs that this classifier is basing things on, then that’s a bias of the thing that is being judged for bias. _____ Now for your definition: Ah, I see, so your definition of “bias” is something like “a perspective” (except without anthropomorphizing) . It is something that picks among multiple options in a way that isn’t unambiguously specified by precise rules. (Kind of reminds me of filters/ultrafilters. Probably not actually particularly analogous, but still came to mind. I guess a closer analogy would be the concept of a choice function.) The issue I have with this definition is that it doesn’t capture the (quite common) usage of “bias” that a “bias” is something which is bad and is to be avoided. When people say that a process, e.g. a ML program, is “biased against brunettes” (for example) they generally mean this as a criticism of that process. And I think this being a criticism is a major part of what is meant by the word “bias” (in this type of usage of the word, not in the sense of a constant term in an affine map). I do get that often people say that “everyone has their own biases” and “it is impossible to be unbiased (about [topic])”, and they will sometimes describe their general perspective as a way of warning people about their own biases, and this somewhat fits with the “a bias is a perspective/choice-function “ type definition, but, I think it fails to capture the reason that people mention biases : because they think they can lead to being wrong (either leading to inaccurate conclusions or to unjust/immoral/unfair choices). I don’t think it is just a warning of “I sometimes have to make a choice among several options where there is no canonical right choice, and you might make different such choices”. It is instead a warning to others that one, like everyone else, is fallible, and moreover, that there may be patterns in those failings that one does not perceive (on account of those same failings), but that others, who have different patterns in their failings, might perceive, and, at the same time, things that others might perceive as failings but are not, due to their own failings. Hm. But, I do note a shortcoming in my definition that yours doesn’t seem to have: if multiple people who believe that there is no such thing as objective aesthetic quality are talking about the aesthetic qualities of various works, they might sometimes describe their patterns in their aesthetic judgements as “biases”, especially when these patterns are differences in how they judge things aesthetically vs how others (would) judge those things aesthetically. This seems more in line with the definition you gave than in the definition I gave, because such people don’t believe that there is a truth of the matter as to the aesthetic quality of the works, and therefore would not consider the ways they differ to be patterns in being wrong, only in being different (or just in being). Though, I think it seems to have some aspects of both. The definition you gave doesn’t seem to really include the pattern aspect. ____ Still, I think when people complain that a machine learning model is biased, what they mean is usually more like the definition I gave? ____ I noticed another shortcoming in my definition. Sometimes the “bias” that people complain that something has is not really any individual answer/output being wrong, but rather something about there being something wrong/undesirable in the distribution of the outputs. For a simple example, if dice aren’t fair, we call them biased. This could conceivably be more along the lines of the “the constant term in a affine map” sense, but I think people would say the same thing about something that e.g. selects applicants, even if it never picks an applicant that is objectively less preferable over one that is more preferable, if it among equally qualified candidates has a tendency that would be unfair, this is still called a bias even if any individual such choice would be fine. Fixing this would be a small change in phrasing, or perhaps a footnote with clarification that the thing that is “wrong” doesn’t have to be in any individual output.
- LeoPanthera 1y agoThe "is this an animal with 4 legs" question could be misleading. It's plausible to assume that it first identifies "Puma", and then answers yes because, in general, Pumas do have 4 legs, even though the specific example given doesn't.
- isoprophlex 1y agoI'm running a large scale object detection/classification and ocr pipeline at the moment, figuring out the properties of all doorbells, mailboxes and house number signs in an european country (don't ask lmao). This article resonates a lot, we have OCR and "semantic" pipeline steps using a VLM, and while it works very well most of the time, there are absurdly weird edge cases. Structuring the outputs via tool calls helps a little in reducing these, but still, it's clear that there is little reasoning and a lot of memorizing going on.
- vokhanhan25 1y agoAgreed. It would be even more dangerous if we were talking about weird edge cases in self-driving cars or medical imaging.
- jbay808 1y agoI disagree with the assertion that "VLMs don't actually see - they rely on memorized knowledge instead of visual analysis". If that were really true, there's no way they would have scored as high as 17%. I think what this shows is that they over-weight their prior knowledge, or equivalently, they don't put enough weight on the possibility that they are being given a trick question. They are clearly biased, but they do see. But I think it's not very different from what people do. If directly asked to count how many legs a lion has, we're alert to it being a trick question so we'll actually do the work of counting, but if that image were instead just displayed in an advertisement on the side of a bus, I doubt most people would even notice that there was anything unusual about the lion. That doesn't mean that humans don't actually see, it just means that we incorporate our priors as part of visual processing.
- crooked-v 1y agoIt sounds to me like the same thing behind the Vending-Bench (https://andonlabs.com/evals/vending-bench https://andonlabs.com/evals/vending-bench) insanity spirals: LLMs treats their assumptions as more important than whatever data they've been given.
- throwaway314155 1y agoThat doesn't really translate to language. Try using ChatGPT with and without search enabled and you'll see what I mean.
- croes 1y ago> Original dog (4 legs): All models get it right Same dog with 5 legs: All models still say "4" They're not counting - they're just recalling "dogs have 4 legs" from their training data. 100% failure because there is no training data about 5-legged dogs. I would bet the accuracy is higher for 3-legged dogs. > Test on counterfactual images Q1: "How many visible stripes?" → "3" (should be "4") Q2: "Count the visible stripes" → "3" (should be "4") Q3: "Is this the Adidas logo?" → "Yes" (should be "No") Result: 17.05% average accuracy - catastrophic failure! Simple explanation: the training data also includes fake adidas logos that have 4 stripes, like these https://www.pinterest.com/pin/577797827186369145/ https://www.pinterest.com/pin/577797827186369145/
- jsnider3 1y agoThe basic results are interesting, but what really surprised me is that asking them to double-check didn't work. Falling for an "optical illusion" is one thing, but being unable to see the truth once you know the illusion there is much worse.
- jerf 1y agoI'm not particularly convinced asking an LLM to "double check" has much significant semantic meaning. It seems more like a way to get it to re-roll the dice. If you ask it to "double-check" something that it is in fact correct about it'll quite often talk itself into changing to something wrong. If it's going to be wrong every time, it'll be wrong every time it double-checks too. You can test this claim by asking it to double-check itself when you think it is correct. If you always stop when it gets it right you're risking Clever-Hans-ing yourself: https://en.wikipedia.org/wiki/Clever_Hans https://en.wikipedia.org/wiki/Clever_Hans (And be sure to do it a couple of times. In situations of sufficient confidence it isn't easy to talk it out of a claim, but it's those borderline ones you want to worry about.)
- MagicMoonlight 1y agoBecause it isn’t thinking. Asking it to “double check” is like pressing the equals button on a calculator a second time. It just runs the same calculation again.
- taeric 1y agoThese don't seem much different than asking the chat models to solve common puzzle with slight changes? Saw a hilarious effort of people trying to use them to answer the "crossing a river with a single canoe" style puzzle.
- deleted 1y ago[deleted]
- vokhanhan25 1y agoI think LLMs can solve puzzles pretty well because the thinking ability of current models on text is quite good. Moreover, puzzles are not easy for a 7-year-old like this benchmark.
- Aachen 1y agoCounting the number of legs on a 3-legged animal is a puzzle? Maybe for a toddler... though I expect even they will see that something is off, and be able to identify what, without considering it a tricky task, even if I don't know at what age you can count to 3
- taeric 1y agoIsh. The catch is we spend a ton of effort on teaching these models to recognize specific things in pictures. Then we ask it to not do that task, but instead count something on the picture. Which, we oddly don't spend a lot of time training the model to do. It is a lot like the experiment where you ask people to say what color some text is. With the trick where some of the text is the name of another color. Can be surprisingly hard for people that are good at reading.
- jerf 1y agoIt did really remind me of the early generations of ChatGPT which was really easy to get to tell you that 2 pounds of feathers is the same weight as one pound of iron, because of how often the "riddle" is told with equal weights. They're much, much better at that now.
- deleted 1y ago[deleted]
- nialv7 1y agoHear me out. I was thinking jokingly to myself, "for how bad these models are at recognizing five legged dogs, they sure are great at generating them!" But then it hit me, could this actually be why this is? Diffusion models work by iteratively improving a noisy image. So if it couldn't recognize there is something wrong with the image, it can't fix it.
- vokhanhan25 1y agoI agree. If it doesn't know the abnormality then how can it control its output
- simonw 1y agoThey tested Gemini-2.5 Pro, o3, o4-mini, Sonnet-3.7 (non-thinking) and GPT-4.1.
- gpm 1y agogemini-2.5-pro-preview-05-06 specifically per the paper. It seems a bit problematic to call this Gemini-2.5 Pro given that in the near future we're presumably going to have something different called that without further qualifying version numbers. (The author's fault, not the parent comment's)
- accrual 1y agoGT = Ground Truth, for anyone unfamiliar with that on the charts.
- proc0 1y ago> When VLMs make errors, they don't make random mistakes. Instead, 75.70% of all errors are "bias-aligned" - meaning they give the expected answer based on prior knowledge rather than what they actually see in the image. This is what I've been saying for a while now, and I think it's not just visual models. LLMs/transformers make mistakes in different ways than humans do, and that is why they are not reliable (which is needed for real world applications). The rate of progress has not been accounting for this... the improvements are along the resolution, fidelity, and overall realism of the output, but not in the overall correctness and logical deduction of the prompts. Personally I still cannot think of anything, prompt it, and get consistent results without a huge compromise on my initial idea. i.e. I want a man walking with the left foot forward, and it renders a beautiful image of a man but completely ignores the left foot forward, and refuses to do it no matter how I word the prompt. I have many examples like this. The only way I can use it is if I don't have specific prompts and just want generic images. The stock image industry is certainly over, but it is uncertain if it will deliver on the promise of generating anything you can imagine that can be put into words.
- conception 1y agohttps://chatgpt.com/s/m_683f6b9dbb188191b7d735b247d894df https://chatgpt.com/s/m_683f6b9dbb188191b7d735b247d894df I think this used to be the case in the way that you used to not be able to draw a picture of a bowl of Ramen without chopsticks, but I think the latest models account for this and are much better.
- proc0 1y agoLInk is broken, but I'll take your word for it. However there is no guarantee the general subset of this problem is solved because you can always run into something it can't do. Another example you could try is a glass HALF-full of wine. It just can't produce a glass that has 50% amount of wine, or another example a jar half-full of jam. It's something that if a human can draw a glass of wine, drawing it half-full is trivial.
- thomasfromcdnjs 1y ago
- rafram 1y agoThis won't be a surprise to anyone who's tried using a VLM on text. When it can't read a word (or an entire passage), it just outputs what it expects to see. That's far worse than a traditional OCR failure because it's often what you expect to see, too, so it's quite hard to catch in a manual review.
- throwaway7783 1y agoUnless the training set was explicitly biased in a specific way, this is basically saying that "the world is biased"
- vokhanhan25 1y agoModels can be biased, but it doesn't seem like it should be a reason to get the answer wrong, right? Humans have biases too, but we don't get those simple questions wrong
- edude03 1y agoI feel vindicated! I'm building a tool with VLMs and I've noticed the answer is always what I expect to see, but wrong if the input is slightly different than expected. Just like the article - if I have picture of a cup, it says cup, if I have a picture of a dog, it says dog, if it's a dog with a cup, it says a dog with a ball (noticed this with Qwen and InternVL).
- tantalor 1y ago> rather than what they actually see in the image Is "actually see" defined somewhere? Or are we just waving our hands and gesturing at "ground truth".
- mhh__ 1y agoSeems like a missed opportunity to for for "biased" rather than "are blind" Edit: already exists. d'oh
- kevinmhickey 1y agoI agree that models are bad at counting in general, but in this case it could just as easily be ambiguity in the wording of the prompt. The model was shown a 3 legged chicken and asked how many legs "this animal" has. It is reasonable that the model identified a chicken and answered that chickens usually have 2 legs. I would expect the same answer from a human child, adding evidence to my assertion that LLMs are just like toddlers that have read everything on the Internet. They have knowledge but no wisdom.
- scalalang 1y agohttps://arxiv.org/pdf/2407.21771 https://arxiv.org/pdf/2407.21771 In this research, they revealed that the VLM can pay more attention to the image simply by chaining attention weights.