3 ms·
Could you help me understand how decoder-only LLMs maintain the Markov property? If you used the same random seed, the input to the model "The cow jumped over t
by eigencoder 1y ago
Could you help me understand how decoder-only LLMs maintain the Markov property? If you used the same random seed, the input to the model "The cow jumped over the" would not give the same output as just "the", right? So isn't that violating the Markov property?
- OJFord 1y agoState (in this sense at least) isn't word/token parsing progress, it's comprising all the input and any context (which may include the entire chat history for example).
- AnotherGoodName 1y agoThere would be need to a state specifically for “the cow jumped over the” (and any other relevant context) and states for all the other times ‘the’ is proceeded by something. This is the limitation i was getting at btw. In the example i wad getting at, if you have an image with solid vertical columns, followed by columns of random static, followed again by solid vertical colors a markov chain could eventually learn all patterns that go solid->32 random bits->different solid color And eventually it would start predicting the different color correctly based on the solid color before the randomness. It ‘just’ needs a state for every possible random color between. This is ridiculous in practice however since you’d need to learn 2^32 states just for relation ship between those two solid colors alone.
- thesz 1y ago> It ‘just’ needs a state for every possible random color between. You can use skipgrams - prefixes with holes in them. Sparse Non-negative Matrix Language Model [1] uses them with great success. [1] https://aclanthology.org/Q16-1024/ https://aclanthology.org/Q16-1024/ The pure n-gram language models would have hard time computing escape weights for such contexts, but mixture of probabilities that is used in SNMLM does not need to do that. If I may, I've implemented an online per-byte version of SNMLM [2], which allows skipgrams' use. They make performance worse, but they can be used. SNMLM's predictive performance for my implementation is within percents to performance of LSTM on enwik8. [2] https://github.com/thesz/snmlm-per-byte https://github.com/thesz/snmlm-per-byte