3 ms·
Even if LLMs are thought of as Markov chains, there are still substantial unsolved scientific questions. For a given prompt/"state", an LLM essentially compute
by psyklic 3y ago
Even if LLMs are thought of as Markov chains, there are still substantial unsolved scientific questions.
For a given prompt/"state", an LLM essentially computes the next-state probabilities. This is done by compressing language and storing particular patterns/distributions. To date, we still don't understand what statistics are stored by LLMs. Or even what stats might be necessary to produce natural-sounding language. (A canonical Markov chain only works with n-gram statistics, but we know these are insufficient.)
IMO figuring this out to a human-level understanding could be a major breakthrough in science. It would reveal a deeper structure to language. It may even help understand how our brains process language; we'll at least know one plausible algorithm. Prior to LLMs, many scholars assumed language was unique to the brain and could only poorly be "computed".