3 ms·
Imagine the nodes in those visualizations are not just a static state, but each have metadata attached to them used to infer the next state in the "chain", whic
by devmor 2y ago
Imagine the nodes in those visualizations are not just a static state, but each have metadata attached to them used to infer the next state in the "chain", which is keyed on a value assigned from a mysterious lookup table based on the current state - so each time the state shifts the metadata on all states can also shift.
(There are also types of LLMs where that metadata is limited in access, one such type is when the current state can only check metadata on previous states and weigh it against the base value of the next states in the chain.)
Then, imagine each chain of states in a Markov Chain is a 2D hash map, like a grid plot. Our current LLMs are like an Nth-dimensional hash map instead, and can have a finite, but extremely large depth. This is pretty near impossible to visualize as a human, but if you're familiar with array/map operations, you should get the idea.
This is a very "base level" understanding, as my learning on LLMs stopped around the time Tensorflow stopped being the new hotness, but hopefully that gives you an idea.