3 ms·
I actually don’t, can you explain it? Doesn’t the AI model basically take an input of context tokens and return a list of probabilities for the next token, whi
by echoangle 2y ago
I actually don’t, can you explain it?
Doesn’t the AI model basically take an input of context tokens and return a list of probabilities for the next token, which is then chosen randomly, weighted by the probability? Isn’t that exactly the definition of a Markov chain?
- fooker 2y agoYou have almost answered the question. LLMs basically return a Markov chain every single time. Think of it as a function returning a value vs returning a function. Now, I'm sure a sufficiently large Markov chain can simulate an LLM but the exponentials involved here would make the number of atoms in the universe a small number. The mechanism that compresses this down into a manageable size is famously 'attention is a you need.!'
- JohnKemeny 2y ago> Now, I'm sure a sufficiently large Markov chain can simulate an LLM but the exponentials involved here would make the number of atoms in the universe a small number. No, LLMs are a Markov chain. Our brain, and other mammalian brains, have feedback, strange loops, that a Markov chain doesn't. In order to reach reasoning, we need to have some loops. In that way, RNNs where much more on the right track towards achieving intelligence than the current architecture.
- fooker 2y agoA modern reasoning model is exactly a feedback loop.
- seanhunter 2y agoThey have the Markov property that the next state depends only on the current state (context window plus the weights of the model) do they not? Any stochastic process which possesses that property is a Markov process/chain I believe.
- fooker 2y agoThe universe has that property too. But no, most LLMs have a tweakable 'temperature' parameter that introduces some randomness and sometimes have very interesting results.
- seanhunter 2y agoTemperature just changes how an llm samples from the token distribution so I wouldn’t think that constitutes “memory” that would contradict the Markov property