4 ms·
LLMs may be overhyped, but transformers in general are underhyped. LLMs make a lot of mistakes because they don't actually know what words mean. The key thing
by Jack000 4y ago
LLMs may be overhyped, but transformers in general are underhyped.
LLMs make a lot of mistakes because they don't actually know what words mean. The key thing is though - it's much harder to generate coherent text when you don't know what the words mean. In a similar vein it's completely unreasonable to expect an LLM to perform visual tasks when it literally has no sense of sight.
The fact that it can kind of sort of do these things at all is evidence of the super-human generalization potential of the transformer architecture.
This isn't very obvious for English because we have prior knowledge of what words mean, but it's a lot more obvious when applied to languages humans don't understand, like DNA and amino acid sequences.
- fourfivefour 4y agoHow can these things not know what words mean? Did you not see how they created a virtual machine under chatGPT? They told it to imitate bash and they typed ls, and cat jokes.txt and it outputted things completely identical to what you'd expect. Look it up. https://www.engraved.blog/building-a-virtual-machine-inside/ https://www.engraved.blog/building-a-virtual-machine-inside/ I don't see how you can explain this as not knowing what words mean. It KNOWS.
- xg15 4y agoYeah, that's the actual bit that baffles me about ChatGPT still. Producing coherent, fluent text is alright, but we could already sort of do that 20 years ago with markov models or even just grammars (see Chomsky). Understanding text in the depth that ChatGPT (and GPT-3) appear to understand the prompts is something entirely different and has to my knowledge never been archieved before the current architectures.
- Jack000 4y agoLLMs are trained exclusively on text, which means they lack crucial context behind the meaning of sentences. The universe of information outside of pure text - vision, sound, etc is completely unknown to it. LLMs are basically the aliens in blindsight. They have a superhuman ability to memorize the context of words it has seen and generalize to new contexts, but it can never be perfect because it's working on incomplete information.
- akimball 4y ago> it can never be perfect because it's working on incomplete information. Unlike you?
- wan23 4y agoThere is a lot of knowledge encoded into the model, but there's a difference between knowing what a sunset is because you read about it on the internet vs having seen one.
- lossolo 4y agoThe whole input of your current session is fed into the model, that's why it tricks you that it "KNOWS" like human would know, in reality this is a lot of data, computation and statistics without any reasoning, that's why there is a lot of examples showing it contradicts itself in the same paragraph, because it doesn't know the meaning of the words its using in the same way as humans do, it only knows probabilities of words in sequences.