3 ms·
RNNs are optimizing a probability, at least with standard training to maximize likelihood. As you say, with a different training criterion sure it could do some
by kastnerkyle 11y ago
RNNs are optimizing a probability, at least with standard training to maximize likelihood. As you say, with a different training criterion sure it could do something different - but so could a Markov chain if it was built to optimize a different objective (though it wouldn't be a Markov chain anymore I don't think).
Look at reinforcement learning - loads of stuff making the Markov assumption and doing quite well on complex tasks, though once again this is quite different than a Markov chain.
In theory (at least my opinion on theory) if we had an infinite size "state" it would be possible to just encode all past information in every state, and use Markov rules. It is just not practical, and much easier to extend lookup to multiple timesteps in the past especially with better understanding of optimizing things like RNNs.
Choosing the maximum p(X_t+1 | X_t) does not infer causality in any way in my mind - only saying that given the data when X_t happens, some event at X_t+1 is likely to happen. Classic correlation. I don't see where you are seeing causality in that part. I also don't see a distinction in simulating a system and data from a system - to me it is just that NNs can model much more complex systems than typical Markov chains.
In general I agree with you, but I find these two concepts are much, much more similar than they are different.
- justifier 11y ago> but so could a Markov chain if it was built to optimize a different objective (though it wouldn't be a Markov chain anymore I don't think). i wanted to think about this, and i thought maybe you could have a purely markov model if it were a self constructing self similar model a markov fractal at each event the chain reconstructs itself from its new perspective in the fractal but even done elegantly i think it would be more undue overhead to constantly reconstruct the model than just feeding a self similar model constantly reconstructed inputs, which i see a more clear affinity to neural networks or something wholly different than either