4 ms·
Hofstadter makes the claim that "these LLMs and other systems like them are all feed-forward". That doesn't sound right to me, but I'm only a casual observer of
by gwright 3y ago
Hofstadter makes the claim that "these LLMs and other systems like them are all feed-forward". That doesn't sound right to me, but I'm only a casual observer of LLM tech. Is his assertion accurate? FWIW, ChatGPT doesn't think so. :-)
- toxik 3y agoIt depends on how you define fed forward, LLMs are typically auto regressive and so can take their own previous output into consideration when generating tokens.
- SpaceManNabs 3y agoThey are not all feed-forward unless it is some other definition that i am not aware of. Convolutional layers, XL hidden states, and graphical networks (which transformers are a special case of) aren't considered feedforward. Unless you consider the entire instance as a singular instance and don't use any hidden states, then I guess it could be considered feed-forward. I don't know. Feedforward doesn't seem like a useful term tbh. Some people mean feedforward as information only goes one direction, but that depends on your arrow. Autoregressive seems more useful here.
- FrustratedMonky 3y agoI think he was referring to feedforward when running GPT in a current conversation, it only remembers the conversation by re-running the prompts. It isn't doing 'feed-back' in the sense of re-updating its weights and learning while having the conversation. So during any one conversation it is only feed-forward.
- bryan0 3y agoI believe what he is referring to is that the LLM’s weights are set when chatting. It is not “learning”. simply using its pretrained weights on your input. Edit: Nope. TIL feed-forward means no loops.
- gautamcgoel 3y agoNo, it is not correct. Transformers have two components: self-attention layers and multi-layer perceptron layers. The first has an autoregressive/RNN flavor, while the latter is feedforward.
- hackinthebochs 3y agoThey are definitely feed-forward. Self-attention looks at all pairs of tokens from the context window, but they do not look backwards in time at its own output. The flow of data is layer by layer, each layer gets one shot at influencing the output. That's feed-forward.
- gamegoblin 3y agoAll the other responses to you at the time of writing this comment are confidently wrong. Definition of Feedforward (from wiki): ``` A feedforward neural network (FNN) is an artificial neural network wherein connections between the nodes do not form a cycle.[1] As such, it is different from its descendant: recurrent neural networks. ``` Hofstadter expected any intelligent neural network would need to be recurrent, ie looping back on itself (in the vein of his book “I am a strange loop”). GPT is not recurrent. It takes in some text input, does a fixed amount of computation in 1 pass through the network, then outputs the next word. He is surprised it doesn’t need to loop for an arbitrary amount of time to “think about” what to say. Being put into an auto-regressive system (where the N-th word it generates gets appended to the prompt that gets sent back into the network to generate the N+1th word) doesn’t make the neural network itself not Feedforward.
- ke88y 3y agoRight. I'm not at all sure what the siblings are talking about. I suspect at least one is confusing linear with feed-forward? But I'm also surprised that Hofstadter keys in on this so heavily. The fact that he wrote an entire pop-sci book on recursion would, in my mind, make him (1) less surprised that AR and R aren't so dissimilar and (2) more sensitive to the sorts of issues that make R more difficult to get working in practice. (In my mind, differentiating between auto-regressive and recursive in this case is kind of the same as differentiating between imperative loops and recursion -- there are extremely important differences in practice but being surprised that a program was written using while loops where you imagined left-folds would be absolutely required seems a bit... odd.)
- gamegoblin 3y agoI think it has to do with the training regime and fixed-computation time nature of feedforward neural networks. Recurrent neural networks have the recursion as part of the training regime. GPT only has auto-regressive "recursion" as part of the inference runtime regime. I think Hofstadter is surprised that you can appear so intelligent without any recursion in the learning/training regime, with the added implication that you can appear so intelligent with a fixed amount of computation per word.