6 ms·
I'm not sure what you mean by "hidden state". If you set aside chain of thought, memories, system prompts, etc. and the interfaces that don't show them, there i
by gugagore 1y ago
I'm not sure what you mean by "hidden state". If you set aside chain of thought, memories, system prompts, etc. and the interfaces that don't show them, there is no hidden state.
These LLMs are almost always, to my knowledge, autoregressive models, not recurrent models (Mamba is a notable exception).
- barrkel 1y agoHidden state in the form of the activation heads, intermediate activations and so on. Logically, in autoregression these are recalculated every time you run the sequence to predict the next token. The point is, the entire NN state isn't output for each token. There is lots of hidden state that goes into selecting that token and the token isn't a full representation of that information.
- gugagore 1y agoThat's not what "state" means, typically. The "state of mind" you're in affects the words you say in response to something. Intermediate activations isn't "state". The tokens that have already been generated, along with the fixed weights, is the only data that affects the next tokens.
- NiloCK 1y agoPlus a randomness seed. The 'hidden state' being referred to here is essentially the "what might have been" had the dice rolls gone differently (eg, been seeded differently).
- barrkel 1y agoNo, that's not quite what I mean. I used the logits in another reply to point out that there is data specific to the generation process that is not available from the tokens, but there's also the network activations adding up to that state. Processing tokens is a bit like ticks in a CPU, where the model weights are the program code, and tokens are both input and output. The computation that occurs logically retains concepts and plans over multiple token generation steps. That it is fully deterministic is no more interesting than saying a variable in a single threaded program is not state because you can recompute its value by replaying the program with the same inputs. It seems to me that this uninteresting distinction is the GP's issue.
- barrkel 1y agoSure it's state. It logically evolves stepwise per token generation. It encapsulates the LLM's understanding of the text so far so it can predict the next token. That it is merely a fixed function of other data isn't interesting or useful to say. All deterministic programs are fixed functions of program code, inputs and computation steps, but we don't say that they don't have state. It's not a useful distinction for communicating among humans.
- gugagore 1y agoI'll say it once more: I think it is useful to distinguish between autoregressive and recurrent architectures. A clear way to make that distinction is to agree that the recurrent architecture has hidden state, while the autoregressive one does not. A recurrent model has some point in a space that "encapsulates its understanding". This space is "hidden" in the sense that it doesn't correspond to text tokens or any other output. This space is "state" in the sense that it is sufficient to summarize the history of the inputs for the sake of predicting the next output. When you use "hidden state" the way you are using it, I am left wondering how you make a distinction between autoregressive and recurrent architectures.
- FeepingCreature 1y agoThe words "hidden" and "state" have commonsense meanings. If recurrent architectures want a term for their particular way of storing hidden state they can make up one that isn't ambiguous imo. "Transformers do not have hidden state" is, as we can clearly see from this thread, far more misleading than the opposite.
- gugagore 1y agoI'll also point out what is most important part from your original message: > LLMs have hidden state not necessarily directly reflected in the tokens being produced, and it is possible for LLMs to output tokens in opposition to this hidden state to achieve longer-term outcomes (or predictions, if you prefer). But what does it mean for an LLM to output a token in opposition to its hidden state? If there's a longer-term goal, it either needs to be verbalized in the output stream, or somehow reconstructed from the prompt on each token. There’s some work (a link would be great) that disentangles whether chain-of-thought helps because it gives the model more FLOPs to process, or because it makes its subgoals explicit—e.g., by outputting “Okay, let’s reason through this step by step...” versus just "...." What they find is that even placeholder tokens like "..." can help. That seems to imply some notion of evolving hidden state! I see how that comes in! But crucially, in autoregressive models, this state isn’t persisted across time. Each token is generated afresh, based only on the visible history. The model’s internal (hidden) layers are certainly rich and structured and "non verbal". But any nefarious intention or conclusion has to be arrived at on every forward pass.
- brookst 1y agoState typically means between interactions. By this definition a simple for loop has “hidden state” in the counter.
- ChadNauseam 1y agoHidden layer is a term of art in machine learning / neural network research. See https://en.wikipedia.org/wiki/Hidden_layer https://en.wikipedia.org/wiki/Hidden_layer . Somehow this term mutated into "hidden state", which in informal contexts does seem to be used quite often the way the grandparent comment used it.
- lostmsu 1y agoIt makes sense in LLM context because the processing of these is time-sequential in LLM's internal time.
- 8note 1y agodo LLM models consider future tokens when making next token predictions? eg. pick 'the' as the next token because there's a strong probability of 'planet' as the token after? is it only past state that influences the choice of 'the'? or that the model is predicting many tokens in advance and only returning the one in the output? if it does predict many, id consider that state hidden in the model weights.
- patcon 1y agoI think recent Anthropic work showed that they "plan" future tokens in advance in an emergent way: https://www.anthropic.com/research/tracing-thoughts-language-model#:~:text=Does%20Claude%20plan%20its%20rhymes? https://www.anthropic.com/research/tracing-thoughts-language...
- 8note 1y agooo thanks!
- NiloCK 1y agoThe most obvious case of this is in terms of `an apple` vs `a pear`. LLMs never get the a-an distinction wrong, because their internal state 'knows' the word that'll come next.
- 3eb7988a1663 1y agoIf I give an LLM a fragment of text that starts with, "The fruit they ate was an <TOKEN>", regardless of any plan, the grammatically correct answer is going to force a noun starting with a vowel. How do you disentangle the grammar from planning? Going to be a lot more "an apple" in the corpus than "an pear"
- halJordan 1y agoIf you dont know, that's not necessarily anyone's fault, but why are you dunking into the conversation? The hidden state is a foundational part of a transformers implementation. And because we're not allowed to use metaphors because that is too anthropomorphic, then youre just going to have to go learn the math.
- markerz 1y agoI don't think your response is very productive, and I find that my understanding of LLMs aligns with the person you're calling out. We could both be wrong, but I'm grateful that someone else spoke saying that it doesn't seem to match their mental model and we would all love to learn a more correct way of thinking about LLMs. Telling us to just go and learn the math is a little hurtful and doesn't really get me any closer to learning the math. It gives gatekeeping.
- tbrownaw 1y agoThe comment you are replying to is not claiming ignorance of how models work. It is saying that the author does know how they work, and they do not contain anything that can properly be described as "hidden state". The claimed confusion is over how the term "hidden state" is being used, on the basis that it is not being used correctly.
- gugagore 1y agoDo you appreciate a difference between an autoregressive model and a recurrent model? The "transformer" part isn't under question. It's the "hidden state" part.