3 ms·
LLMs definitely also have finite context length. And if we consider padding, it is also constant. The k is huge compared to most Markov chains used historically
by matusp 10mo ago
LLMs definitely also have finite context length. And if we consider padding, it is also constant. The k is huge compared to most Markov chains used historically, but it doesn't make it less finite.
- srean 10mo agoThat's not correct. Even a toy like an exponential weighted moving averaging produces unbounded context (of diminishing influence).
- matusp 10mo agoWhat do you mean? I can only input k tokens into my LLM to calculate the probs. That is the definition of my state. In the exact way that N-gram LMs use N tokens, but instead of using ML models, they calculate the probabilities based on observed frequencies. There is no unbounded context anywhere.
- srean 10mo agoThat's different. You can certainly feed k-grams one at a time to, estimate the the probability distribution over next token and use that to simulate a Markov Chain and reinitialize the LLM (drop context). In this process the LLM is just a look up table to simulate your MC. But an LLM on its own doesn't drop context to generate, it's transition probabilities change depending on the tokens.