8 ms·
I usually have this technical hypothetical discussions with ChatGpt, I can share if you like, me asking him this: aren't LLMs just huge Markov Chains?! And now
by atum47 10mo ago
I usually have this technical hypothetical discussions with ChatGpt, I can share if you like, me asking him this: aren't LLMs just huge Markov Chains?! And now I see your project... Funny
- pavel_lishin 10mo ago> I can share if you like Respectfully, absolutely nobody wants to read a copy-and-paste of a chat session with ChatGPT.
- atum47 10mo agoWhen you say nobody you mean you, right? You can't possible be answering for every single person in the world. I was having a discussion about similarities between Markov Chains and LLMs and short after I found this topic on HN, when I wrote "I can share if you like" was as a proof about the coincidence.
- deleted 10mo ago[deleted]
- matusp 10mo agoLLMs are indeed Markov chains. The breakthrough is that we are able to efficiently compute well performing probabilities for many states using ML.
- famouswaffles 10mo agoLLMs are not Markov Chains unless you contort the meaning of a Markov Model State so much you could even include the human brain.
- sophrosyne42 10mo agoWell LLMs aren't human brains, unless you contort the definition of matrix algebra so much you could even include them.
- ben_w 10mo agoQM and GR can be written as matrix algebra, atoms and electrons are QM, chemistry is atoms and electrons, biology is chemistry, brains are biology. An LLM could be implemented with a Markov chain, but the naïve matrix is ((vocab size)^(context length))^2, which is far too big to fit in this universe. Like, the Bekenstein bound means writing the transition matrix for an LLM with just 4k context (and 50k vocabulary) at just one bit resolution, the first row (out of a bit more than 10^18795 rows) ends up with a black hole >10^9800 times larger than the observable universe.
- sophrosyne42 10mo agoYes, sure enough, but brains are not ideas, and there is no empirical or theoretical model for ideas in terms of brain states. The idea of unified science all stemming from a single ultimate cause is beautiful, but it is not how science works in practice, nor is it supported by scientific theories today. Case in point: QM models do not explain the behavior of larger things, and there is no model which gives a method to transform from quantum to massive states. The case for brain states and ideas is similar to QM and massive objects. While certain metaphysical presuppositions might hold that everything must be physical and describable by models for physical things, science, which should eschew metaphysical assumptions, has not shown that to be the case.
- chpatrick 10mo agoNot sure why that's contorting, a markov model is anything where you know the probability of going from state A to state B. The state can be anything. When it's text generation the state is previous text to text with an extra character, which is true for both LLMs and oldschool n-gram markov models.
- wizzwizz4 10mo agoA GPT model would be modelled as an n-gram Markov model where n is the size of the context window. This is slightly useful for getting some crude bounds on the behaviour of GPT models in general, but is not a very efficient way to store a GPT model.
- chpatrick 10mo agoI'm not saying it's an n-gram Markov model or that you should store them as a lookup table. Markov models are just a mathematical concept that don't say anything about storage, just that the state change probabilities are a pure function of the current state.
- srean 10mo agoYou say state can be anything, no restrictions at all. Let me sell you a perfect predictor then :) The state is the next token.
- famouswaffles 10mo agoYes, technically you can frame an LLM as a Markov chain by defining the "state" as the entire sequence of previous tokens. But this is a vacuous observation under that definition, literally any deterministic or stochastic process becomes a Markov chain if you make the state space flexible enough. A chess game is a "Markov chain" if the state includes the full board position and move history. The weather is a "Markov chain" if the state includes all relevant atmospheric variables. The problem is that this definition strips away what makes Markov models useful and interesting as a modeling framework. A “Markov text model” is a low-order Markov model (e.g., n-grams) with a fixed, tractable state and transitions based only on the last k tokens. LLMs aren’t that: they model using un-fixed long-range context (up to the window). For Markov chains, k is non-negotiable. It's a constant, not a variable. Once you make it a variable, near any process can be described as markovian, and the word is useless.
- deleted 10mo ago[deleted]
- cwyers 10mo agoYeah, there's only two differences between using Markov chains to predict words and LLMs: * LLMs don't use Markov chains, * LLMs don't predict words.
- arboles 10mo ago* Markov chains have been used to predict syllables or letters since the beginning, and an LLMs tokenizer could be used for Markov chains * The R package markovchain[1] may look like it's using Markov chains, but it's actually using the R programming language, zeros and ones. [1] https://cran.r-project.org/web/packages/markovchain/index.html https://cran.r-project.org/web/packages/markovchain/index.ht...
- arboles 10mo agoMarkov models with more than 3 words as "context window" produce very unoriginal text in my experience (corpus used had almost 200k sentences, almost 3 million words), matching the OP's experience. These are by no means large corpuses, but I know it isn't going away with a larger corpus.[1] The Markov chain will wander into "valleys" of reproducing paragraphs of its corpus one for one because it will stumble upon 4-word sequences that it has only seen once. This is because 4 words form a token, not a context window. Markov chains don't have what LLMs have. If you use a syllable-level token in Markov models the model can't form real words much beyond the second syllable, and you have no way of making it make more sense other than increasing the token size, which exponentially decreases originality. This is the simplest way I can explain it, though I had to address why scaling doesn't work. [1] There are 4^400000 possible 4-word sequences in English (barring grammar) meaning only a corpus with 8 times that amount of words and with no repetition could offer two ways to chain each possible 4 word sequence.
- arboles 10mo agoI was really sleepy when I wrote this
- srean 10mo agoThey are definitely not Markov Chains they may, however, be Markov Models. There's a difference between MC and MM.
- matusp 10mo agoWhat do you mean? The states are fully observable (current array of tokens), and using an LLM we calculate the probabilities of moving between them. What is not MC about this?
- srean 10mo agoI suggest getting familiar with or brushing up on the differences between a Markov Chain and a Markov Model. The former is a substantial restriction of the latter. The classic by Kemeny and Snell is a good readable reference. MC have constant and finite context length, their state is the most recent k tuple of emitted alphabets and transition probabilities are invariant (to time and tokens emitted)
- matusp 10mo agoLLMs definitely also have finite context length. And if we consider padding, it is also constant. The k is huge compared to most Markov chains used historically, but it doesn't make it less finite.
- srean 10mo agoThat's not correct. Even a toy like an exponential weighted moving averaging produces unbounded context (of diminishing influence).
- matusp 10mo agoWhat do you mean? I can only input k tokens into my LLM to calculate the probs. That is the definition of my state. In the exact way that N-gram LMs use N tokens, but instead of using ML models, they calculate the probabilities based on observed frequencies. There is no unbounded context anywhere.
- roarcher 10mo ago...are you under the impression that you have an exclusive relationship with "him"? Everyone else has access to ChatGPT too.
- atum47 10mo agoYes. Yes I was. Thank you for the wake up call. I was under the impression that he was talking only to me.
- deleted 10mo ago[deleted]
- atum47 10mo agoDon't know what happened. I stumbled onto a funny coincidence - me talking to a LLM about its similarities with MC - decided to share on a post about using MC to generate text. Got some nasty comments and a lot of down votes. Even though my comment sparked a pretty interesting discussion. Hate to be that guy, but I remember this place being nicer.
- roarcher 10mo agoEver since LLMS became popular, there's been an epidemic of people pasting ChatGPT output onto forums (or in your case, offering to). These posts are always received similarly to yours, so I'm skeptical that you're genuinely surprised by the reaction. Everyone has access to ChatGPT. If we wanted its "opinion" we could ask it ourselves. Your offer is akin to "Hey everyone, want me to Google this and paste the results page here?". You would never offer to do that. Ask yourself why. These posts are low-effort and add nothing to the conversation, yet the people who write them seem to expect everyone to be impressed by their contribution. If you can't understand why people find this irritating, I'm not sure what to tell you.
- pavel_lishin 10mo agoNobody was being nasty. roarcher explained why people reacted the way they did.
- deleted 10mo ago[deleted]