6 ms·
Some of the proposed models for human thought can't handle negation. An example of this is map-based models. Theories which propose human thought is a form of
by phaedrus 4y ago
Some of the proposed models for human thought can't handle negation. An example of this is map-based models. Theories which propose human thought is a form of a production system (as in logic) get a boost from the fact they can explain negation as well as other things. So to see DALL-E can't handle negation is, to my thinking, an important difference between the way it "thinks" (or what passes for that) and how we do.
- astrange 4y agoIts training data might not have negations in it - I'm not sure anyone has tried to explain what CLIP understands of text. Besides that, we know it doesn't because models don't recurse because no matter how big they are, they still have constant execution times. And we gave up on systematically parsing text in favor of just vibing anyway (technical term). That's just as well since most parsers didn't really work in practice; I think I vaguely remember one as kind of working (https://www.link.cs.cmu.edu/link/ https://www.link.cs.cmu.edu/link/) but it can't handle bad grammar or misspellings like ML can.
- visarga 4y agoThat's not true, the autoregressive LMs (GPTs) can generate variable length outputs allowing for step by step computation. Recent papers get them to generate the full chain of thought reasoning process to the solution (also called rationales in some papers), improving accuracy. But this version of Dall-E is not autoregressive, they've gone for diffusion models in the generator and for a single vector embedding (no internal spatial grid structure) for the encoder. That explains why it's bad at negations, counting and stacking coloured blocks. I believe these skills were not the top priority of OpenAI this time, they understandably went for pretty images.
- astrange 4y agoI agree GPT as a whole generates variable length outputs, but the model isn't generating them in a single pass, is it? I thought there was some sampling process implemented in traditional code outside the model that made it do that, and it wasn't learned. Have forgotten the details. If Dall-E took variable length sequences as inputs, its execution time would still be O(N) on the length of the sequence and it would still have a fixed amount of memory to work with. Whereas a computer program of length N can do pretty much anything, and a human given a description can go off and think about it.
- bloaf 4y agoI strongly suspect the inability to do negation comes from the fact that the training dataset contains virtually no images tagged with negated descriptors. It wouldn't surprise me to see it correctly handle negations that might be more common in image tags (e.g. "a car not running" might be a sufficiently common phrase for "broken down car").
- gwern 4y agoSeems straightforward to check something like LAION-400M's captions to see how often negations show up. If it's simply a lack of data, that should be easy to fix. Given an image caption, every other caption is a relevant negation: 'a photo of a happy dog' is also not 'a photo of a sad cat', and can be turned into the caption 'a photo of a happy dog and not a photo of a sad cat'. So lots of obvious data augmentation and synthetic data approaches to apply there to boost negation.
- ralfd 4y ago> and how we do https://langcog.stanford.edu/papers/NF-cogsci2013.pdf https://langcog.stanford.edu/papers/NF-cogsci2013.pdf Fun fact: Young Children also don't understand negation. There are some parenting theories claiming that is why one should use avoid negations in commands.