7 ms·
It's about precision. Image models can be very off and still produce a satisfying result. Consider that I could literally vary all the pixels in an image rando
by stravant 4y ago
It's about precision.
Image models can be very off and still produce a satisfying result. Consider that I could literally vary all the pixels in an image randomly by 10% and you'd just see it as a bit low quality but otherwise perfectly cohesive image.
Language models have no such luck, the problem they're trying to solve is way "sharper", it's very easy for their results to be strictly wrong if they're off even a little bit.
So you need a much larger model to get a sufficient level of "sharpness" for text.
- uh_uh 4y agoMaybe another way to think of it is that the error correction part of image generation models is offloaded to the human visual cortex which is a very old evolutionary construct and thus had time to become very resilient? In case of text generation, maybe the error tolerance of the human brain is less developed as human-level language is a newer evolutionary invention. It'd be interesting if the parameter/complexity requirements are actually similar once you examine the system as a whole, meaning machine _and_ human brain.
- Codesleuth 4y ago> image generation models is offloaded to the human visual cortex which is a very old evolutionary construct and thus had time to become very resilient This is a very important point. A group of my colleagues (who are not tech people) are much more impressed with the image generation models than with the chat interface, even though the images are often whacky or just wrong. Yet the fact that it tried is impressive to them, with their minds managing to fill in the blanks. I wonder how this compares to how a toddler speaks vs. paints/draws, which is typically better in the former than the latter. I'm both cases, we fill in the blanks in our minds.
- v01dlight 4y agoToddler speaking gets impressive/surprising quite fast, whereas the drawing usually does not. The most surprising thing about most toddler drawings is listening to the kid describe it or tell you about making it.
- glomgril 4y agoThe consistency of descriptions is particularly surprising to me. Like you got a roughly circular collection of seemingly random scribbles, but they can tell you exactly which parts of it correspond to the person's nose, hair, arms, eyes, etc. And the descriptions seem to stay the same if you ask about the same picture on different days. Still not sure what to make of this phenomenon but it is fascinating.
- spacebanana7 4y agoI wonder whether video and metaverse generation models will be even smaller than an image model because of this mechanism. The mapping and motion parts of the human brain are also old evolutionary constructs that could error correct the output of models.
- flakeoil 4y agoIt's kind of similar to audio vs video. Although audio requires less data and processing than video, it's much more difficult to get the audio good than the video and if the audio is bad or even missing for some time, it's useless, while if the video is bad, stuck or missing, it's not that big of a deal most of the time. This is particularly true in a video conferencing situation. If the audio is bad, you miss out a lot. If the video is bad, it's not a big deal.
- devenvdev 4y agoDeaf people would disagree :) if you talk in sign language on zoom missing video parts would ruin the conversation. I don't think it's about precision, in the case of audio vs video - if you remove all the even columns from a video it would be similar to reducing quality, the same can be done with audio - removing half of the frequencies uniformly will just lower the quality.
- kelipso 4y agoThat's a pretty specific case. You can get really good performance for a ton of tasks in video (video question answering, object identification and tracking, action recognition, etc) by just sampling a frame per second or even less frequently. Definitely can't do that with audio.
- jameshart 4y agoRight - The ‘palette’ for text generation is smaller: just 26 or so letters (plus some other characters), and if you put the wrong ones next to one another the result is garbage. There’s something interesting in the fact that an image based system doesn’t need as much complexity to capture a semantic model as a verbal system does; I think there’s maybe a parallel there to the way that human minds find it easier to just ‘visualize’ some things as a basis for reasoning about them, but if we can’t ‘visualize’ and instead have to ‘think things through’ it’s a more intensive process. Like, GPT has well known trouble counting - ask it for five things and it will give you four or six. Humans can offload some thinking about counting to visual/spatial reasoning though.
- hgsgm 4y agoThe pallette for LLM is tokens not characters.
- Der_Einzige 4y agoDepends on the LLM. Character based LLMs exist, and even have advantages vs regular LLMs...
- dragonwriter 4y agoAnd if it is characters (as it is for some models), its more than 26 of them for English. Between space, case, punctuation, and digits, its basically 7-bit ASCII without most of the control characters (newline is semantically important, the rest not), almost 100 characters.
- ipunchghosts 4y ago> It's about precision. This is utterly wrong. There is a huge amount of redundancy in images compared to language. This redundancy is why image models have yet to surpass language models. In some sense, language is much easier than the vision problem.
- hgsgm 4y agohow do you measure "surpass"?
- dragonwriter 4y agoI think no one has bothered with using as many images as documents used to train GPT-3.5, to create as big of a model, and then RLHF as done to produce ChatGPT from GPT-3.5 is why image models haven’t surpassed language models. At any level of scale of model and scale of training set, images models do surpass language models.
- ipunchghosts 4y agoNo way this is true. I have yet to see it. Image permutations are so easy to come by where as they are eventually infinite.