5 ms·
What's not apparent is how does the Markovian framework account for the apparent long-range dependencies and global coherence we observe in LLM outputs, particu
by nameisconcealed 2y ago
What's not apparent is how does the Markovian framework account for the apparent long-range dependencies and global coherence we observe in LLM outputs, particularly in tasks requiring multi-step reasoning or maintaining context across many tokens? While the analysis elegantly explains local statistical patterns, it seems to conflict with transformer's ability to maintain attention across arbitrary distances... or am I missing something?
- ogogmad 2y agoThey use O(T^K) states where T is the size of the "vocabulary" (which would in fact be tokens in the case of an LLM) and K is the size of the context window. In practice, this would be an enormous number of states, and would be able to account for all possible long-range dependencies. This bit is trivial. The paper seems to be more about the stationary distributions of such Markov chains.
- magicalhippo 2y ago> In practice, this would be an enormous number of states The paper mentions this explicitly: For GPT-3 (Brown et al., 2020), this represents 5 × 10^11 training tokens, which pales in comparison with the number of non-zero elements in Qf, given by T^K+1 ≈ 10^9632.
- jampekka 2y agoTransformers (and other feedforward autoregressive models) are obviously Markovian. Here it's shown that these are equivalent to a Markov Chain with a transition matrix. There's in principle no reason why such can't have arbitrary (but not infinite) range dependencies. The major difference is that these Markov Chains can be too huge to compute and don't have a feasible training algorithm.
- fragmede 2y agoThat's not especially true. Markov models only care about the last N words of context. In practice, this leads to N being small. Transformers consider more/larger state using self-attention, a more complex mathematical mechanism, which means they integrate all across the entire context window in a richer attention model.
- jampekka 2y agoTransformers only care about N words/tokens of context. That was the primary reason they became the dominant language model architecture. A system being a Markov chain doesn't say anything more about the function mapping inputs to states than that the function has no memory and has access to the full state that affects the transition probabilities to the next state.
- ogogmad 2y agoMaybe this is at cross-purposes: Functions don't have memory. Nothing real ever "is" a Markov chain. It's simply a (slightly unusual) modelling choice which you can use whenever you've listed all possible inputs that can affect a transition probability - once you've done that, you can apply it to pretty much anything.
- Jensson 2y ago> Nothing real ever "is" a Markov chain. What do you mean? Quantum states over time are very much markov chains as well. > It's simply a (slightly unusual) modelling choice which you can use whenever you've listed all possible inputs that can affect a transition probability That has nothing to do with markov chains, do you mean a specific implementation? Markov chains is a probability theory concept and has nothing to do with any implementation.
- jampekka 2y agoIn many cases the Markov assumption is made to simplify things even though it's not realistic. But nevertheless it does describe what the model itself does even if in the real environment it doesn't act as a real Markov process (e.g. doesn't necessarily have a stationary distribution). It applies here too in a case where the model interacts with humans, as the LLM state doesn't include the human. While Markov chains/processes are a powerful and general abstraction, the assumptions are quite strict. E.g. even the simplest recurrent systems aren't Markov, and this leads to things like having to assume exponential distribition of event durations in time series contexts, which is quite limiting. In language models it draws a clear demarkation between recurrent and feedforward neural networks and brings some crucial understanding into benefits and limitations of e.g. transformers vs recurrent models.
- ComputerGuru 2y agoI’ve found that it’s entirely too easy to get inconsistent (i.e. self-conflicting) output in the same response from when the likes of gpt4o when you require reasoning and logical deduction on topics completely or mostly alien to it and not found in the massive input corpus. So I’m not sure this is the difference between a markov chain and a transformer llm.