6 ms·
> When VLMs make errors, they don't make random mistakes. Instead, 75.70% of all errors are "bias-aligned" - meaning they give the expected answer based on prio
by proc0 1y ago
> When VLMs make errors, they don't make random mistakes. Instead, 75.70% of all errors are "bias-aligned" - meaning they give the expected answer based on prior knowledge rather than what they actually see in the image.
This is what I've been saying for a while now, and I think it's not just visual models. LLMs/transformers make mistakes in different ways than humans do, and that is why they are not reliable (which is needed for real world applications). The rate of progress has not been accounting for this... the improvements are along the resolution, fidelity, and overall realism of the output, but not in the overall correctness and logical deduction of the prompts. Personally I still cannot think of anything, prompt it, and get consistent results without a huge compromise on my initial idea.
i.e. I want a man walking with the left foot forward, and it renders a beautiful image of a man but completely ignores the left foot forward, and refuses to do it no matter how I word the prompt. I have many examples like this. The only way I can use it is if I don't have specific prompts and just want generic images. The stock image industry is certainly over, but it is uncertain if it will deliver on the promise of generating anything you can imagine that can be put into words.
- conception 1y agohttps://chatgpt.com/s/m_683f6b9dbb188191b7d735b247d894df https://chatgpt.com/s/m_683f6b9dbb188191b7d735b247d894df I think this used to be the case in the way that you used to not be able to draw a picture of a bowl of Ramen without chopsticks, but I think the latest models account for this and are much better.
- proc0 1y agoLInk is broken, but I'll take your word for it. However there is no guarantee the general subset of this problem is solved because you can always run into something it can't do. Another example you could try is a glass HALF-full of wine. It just can't produce a glass that has 50% amount of wine, or another example a jar half-full of jam. It's something that if a human can draw a glass of wine, drawing it half-full is trivial.
- thomasfromcdnjs 1y agochatgpt can easily do that? What was the last time you tried?
- proc0 1y agoI just tried with Flux.1 Kontext, which I assume is better than o3 at creating images, but I'll admit I didn't do extensive tests. It's more trying to do test projects. Maybe I'm having bad luck but doesn't seem that way.
- jxjnskkzxxhx 1y ago> LLMs/transformers make mistakes in different ways than humans do Sure but I don't think this is an example of it. If you show people a picture and ask "how many legs does this dog have?" a lot of people will look at the picture, see that it contains a dog, and say 4 without counting. The rate at which humans behave in this way might differ from the rate at which llms do, but they both do it.
- DeathRay2K 1y agoI don’t think there’s a person alive who wouldn’t carefully and accurately count the number of legs on a dog if you ask them how many legs this dog has. The context is that you wouldn’t ask a person that unless there was a chance the answer is not 4.
- tantalor 1y agoYou deeply overestimate people. The models are like a kindergartner. No, worse than that, a whole classroom of kindergartners. The teacher holds up a picture and says, "and how many legs does the dog have?" and they all shout "FOUR!!" because they are so excited they know the answer. Not a single one will think to look carefully at the picture.
- jxjnskkzxxhx 1y agoIt's hilarious how off you are.
- ekianjo 1y agoYou have never seen the video of the gorilla in the background?
- petesergeant 1y agoThat's a specific example that when you draw a human's attention to something (eg: count the number of ball passes in this video), they hyper-fixate on that, to the exclusion of other things, so it seems like it makes the opposite point that I think you're trying to?
- 0xab 1y ago> When VLMs make errors, they don't make random mistakes. Instead, 75.70% of all errors are "bias-aligned" - meaning they give the expected answer based on prior knowledge rather than what they actually see in the image. Yeah, that's exactly what our paper said 5 years ago! They didn't even cite us :( "Measuring Social Biases in Grounded Vision and Language Embeddings" https://arxiv.org/pdf/2002.08911 https://arxiv.org/pdf/2002.08911
- moralestapia 1y agoThat's weird, you're at MIT. You're in the circle of people that's allowed to succeed. I wouldn't think much about it, as it was probably a genuine mistake.
- JackYoustra 1y agoWhat does allowed to succeed mean?
- moralestapia 1y agoYour work usually has 1,000x the exposure and external validation compared to doing it outside those environments, where it would just get discarded and ignored. Not a complain, though. It's a requirement for our world to be the way it is.
- _345 1y agoIs there truth to this? Do you have any sources to link to on this
- moralestapia 1y agoSure dude, here's the link to the UN Resolution about which researchers deserve attention and which others do not, signed by all countries around the world [1]. *sigh* It's pretty obvious, if you publish something at Harvard, MIT, et. al. you even get a dedicated PR team to make your research stand out. If you publish that on your own, or on some small research university in Namibia, no one will notice. I might be lying, though, 'cause there's no "proof". 1: https://tinyurl.com/3uf7r5r7 https://tinyurl.com/3uf7r5r7