6 ms·
Really "eye-opening" work. These models don’t actually “see”, they just recall what they’ve memorized, even when the image clearly shows something different. It
by ahrmb 1y ago
Really "eye-opening" work. These models don’t actually “see”, they just recall what they’ve memorized, even when the image clearly shows something different. It’s a bit scary how confidently they get things wrong when reality doesn’t match their training data.
- foxglacier 1y agoIt's not too different from people. We also don't really "see" and mostly recall what we expect to see. What do you expect when the question is wrong "How many legs does this animal have? Answer with a number" but it's not a picture of an animal. What are you supposed to do? Answer 0?
- regularjack 1y agoYou answer "I don't know"
- amelius 1y agoWhat if that is not in your vocabulary?
- ramoz 1y agoThis is interesting actually. And reminds me of something vaguely - a book or something that describes how human attention and the things we see are highly optimized by evolution. We often miss a lot of details in reality due to this.
- zehaeva 1y agoIf it were a Fiction novel then might I suggest Blindsight by Peter Watts?
- vunderba 1y agoThat wasn't one of the questions - any reasonable person would have classified that chicken as an animal, albeit a mutant one. I would also hardly count many of these questions as "tricks" either. Take the chess example. A lot of my friends and myself have been playing chess since we were young children and we all know that a fully populated chess board has 32 pieces (heavily weighted in our internal training data), but not a single one of us would have gotten that question wrong.
- gowld 1y agoDon't be too literal. Imagine walking to a room an seeing someone grab a handful of chess pieces off of a set-up board, and proceed to fill bags with 4 pieces each. As they fill the 8th bag, they notice only 3 pieces are left. Are you confident that you would respond "I saw the board only had 31 pieces on it when you started", or might you reply "perhaos you dropped a piece on the floor"?
- vunderba 1y agoI'm not. I'm referencing the paper - not some hypothetical abstract word problem. Imagine walking into a room, where the pieces are slowly morphing from staid Staunton structures into amorphous blobs of lava lamp Cthulhu nightmares. If a locomotive steam train from Denver passes within 15 meters of the room, how many passengers paid for the tickets using a cashier's check? Nobody's arguing that humans never take logical shortcuts or that those shortcuts can cause us to make errors. Some of the rebuttals in this thread are ridiculous. Like what if I forced you to stare at the surface of the sun followed by waterboarding for several hours, and then asked you to look at a 1000 different chess boards. Are you sure you wouldn't make a mistake? In the paper the various VLLMs are asked to double-check which still didn't make a difference. The argument is more along the lines that VLLMs (and multimodal LLMs) aren't really thinking in the same way that humans do. And if you REALLY need an example albeit a bit tangential - try this one out. Ask any SOTA (multimodal or otherwise) model such as gpt-image-1, Kontext, Imagen4, etc. for a five-leaf cover. It'll get it about 50% of the time. Now go and ask any kindergartener for the same thing.
- wat10000 1y agoDepending on the situation, I'd either walk away, or respond with, "What animal?"
- enragedcacti 1y agoIts true that our brains take lots of shortcuts when processing visual information but they don't necessarily parallel the shortcuts VLMs take. Humans are often very good at identifying anomalous instances of things they've seen thousands of times. No one has to tell you to look closely when you look at your partner in a mirror, you'll recognize it as 'off' immediately. Same for uncanny CGI of all types of things. If we were as sloppy as these models then VFX would be a hell of a lot easier. Ironically I think a lot of people in this thread are remembering things they learned about the faultiness of humans' visual memory and applying it to visual processing.
- soulofmischief 1y agoHumans do this, but we have more senses to corroborate which leads to better error checking. But what you see in your visual mental space is not reality. Your brain makes a boatload of assumptions. To test this, research what happens during saccades and how your brain "rewinds" time. Or try to find your blind spot by looking at different patterns and noticing when your brain fills in the gaps at your blind spot. It will recreate lines that aren't there, and dots will wholly disappear. Additionally as an anecdote, I have noticed plenty times that when I misread a word or phrase, I usually really do "see" the misspelling, and only when I realize the misspelling does my brain allow me to see the real spelling. I first noticed this phenomenon when I was a child, and because I have a vivid visual memory, the contrast is immediately obvious once I see the real phrase. Additionally, I seem to be able to oversharpen my vision when I focus, making myself hyperattentive to subtle changes in motion or color. The effect can be quite pronounced sometimes, reminiscent of applying am edge filter. It's clearly not reality, but my visual system thinks it is. If you really want to understand how much the visual system can lie to you, look into some trip reports from deleriants on erowid. I wouldn't recommend to try them yourself but I will say that nothing will make you distrust your eyes and ears more. It's basically simulated hallucinatory schizophrenia and psychosis.