2 ms·
Sure, but GPT-3 was trained by self-supervised learning on only static text. We see how powerful even just adding captions to text can be with the example of DA
by dougabug 4y ago
Sure, but GPT-3 was trained by self-supervised learning on only static text. We see how powerful even just adding captions to text can be with the example of DALLE-2. GATO takes this further by letting the large scale Transformer learn in both simulated and real interactive environments, giving it the kind of grounding that the earlier models lacked.
- PaulHoule 4y agoI will grant that the grounding is important. The worst intellectual trend of the 20th century was the idea that language might give you some insight into behavior (Sapir–Whorf hypothesis, structuralism, post-structuralism, ...) whereas language is really like the evidence left after a crime. For instance, language maximalists see mental models as a fulcrum point for behavior, and they are, but they have nothing to do with language. I have two birds that come to my window. One of them has no idea of what the window is and attacks her own reflection hundreds of times a day. She can afford to do it because her nest is right near the bird feeder and doesn't need to work to eat, in fact it probably seems meaningful to her that another bird is after her nest. This female cardinal flies away if I am in the room where she is banging. There is a rose-breasted grosbeak, on the other hand, that comes to the same window. She doesn't mind if I come close to the window, instead I see her catch the eye of her reflection and then catch my eye. She basically understands the window. Here you have two animals with two different acquired mental models... But no language. What I like about the language-image models is how the image grounds reality outside language, and that's important because the "language instinct" is really a peripheral that attaches to an animal brain. Without the rest of the animal it's useless.