3 ms·
RNNs (recurrent neural networks) can be seen as generalizations of Markov chains - or as I prefer think of it, Markov chains are very limited RNNs. There is an
by kastnerkyle 11y ago
RNNs (recurrent neural networks) can be seen as generalizations of Markov chains - or as I prefer think of it, Markov chains are very limited RNNs. There is an enormous amount of research happening around these sequential models (and differentiable DAGs (directed acyclic graphs) aka neural networks in general) these days.
Markov chains have never (in the work I have seen at least) tried to imply causation - purely correlation/probability. There is other recent research on causality (see the work of Pearl), but just because one thing very likely follows another doesn't mean the preceding event "caused" the next one, as both events could be side effects from the same (unobserved) root cause.
- justifier 11y agobut the difference between neural nets and markov chains is the probabilistic element of nn is the optimal traversal and could be replaced with a perfect algo if one were found to be computationally feasable whereas markov chains start from a place of probability and require it the construction stage of the markov model agrees with your statement: > just because one thing very likely follows another doesn't mean the preceding event "caused" the next one but when in use the current state is the cause of the following state informed by your model nn can simulate a system whereas an mc can only simulate data from a system
- kastnerkyle 11y agoRNNs are optimizing a probability, at least with standard training to maximize likelihood. As you say, with a different training criterion sure it could do something different - but so could a Markov chain if it was built to optimize a different objective (though it wouldn't be a Markov chain anymore I don't think). Look at reinforcement learning - loads of stuff making the Markov assumption and doing quite well on complex tasks, though once again this is quite different than a Markov chain. In theory (at least my opinion on theory) if we had an infinite size "state" it would be possible to just encode all past information in every state, and use Markov rules. It is just not practical, and much easier to extend lookup to multiple timesteps in the past especially with better understanding of optimizing things like RNNs. Choosing the maximum p(X_t+1 | X_t) does not infer causality in any way in my mind - only saying that given the data when X_t happens, some event at X_t+1 is likely to happen. Classic correlation. I don't see where you are seeing causality in that part. I also don't see a distinction in simulating a system and data from a system - to me it is just that NNs can model much more complex systems than typical Markov chains. In general I agree with you, but I find these two concepts are much, much more similar than they are different.
- justifier 11y ago> but so could a Markov chain if it was built to optimize a different objective (though it wouldn't be a Markov chain anymore I don't think). i wanted to think about this, and i thought maybe you could have a purely markov model if it were a self constructing self similar model a markov fractal at each event the chain reconstructs itself from its new perspective in the fractal but even done elegantly i think it would be more undue overhead to constantly reconstruct the model than just feeding a self similar model constantly reconstructed inputs, which i see a more clear affinity to neural networks or something wholly different than either