4 ms·
I have done some work with RAAMs and other simple recurrent network based methods. It's terribly tedious, you can hardly represent more than five symbols or so.
by werg 15y ago
I have done some work with RAAMs and other simple recurrent network based methods. It's terribly tedious, you can hardly represent more than five symbols or so.
I did my Bachelor's thesis on related techniques (pressing variable length strings into fixed sized vectors to deal with neural nets - URL below) - I can only say pertaining to neural networks that the decoding capabilities of NNs are severly limited. I actually developed a Jordan Network with a conventional additive sigmoid NN that could encode/decode more than 40 symbols! The technique was based on Cantor coding, but I had to set the weights by hand and could not retrain without losing performance (unstable fixed point attractor of Backpropagation through Time) - i.e. it's kind of a dirty hack.
It's also really hard to adapt parameters for sufficiently large nets (so my idea is that you need huge nets to represent language). In order to deal with this complexity limitation I've also looked into reservoir computing style networks (Echo State Networks) which use a large randomly initialized network paired with a linear learned network. ESNs may be good at modelling many kinds of temporal dynamics, but their capacity to represent relations seems rather limited as well.
So making an arbitrary length string walk into a vector may sound attractive, but either (a) don't expect to be able to decode it using conventional NNs or (b) don't expect to acheive compression, i.e. blowing up your representation may help (cf LVQ nets) and using compressing techniques such as Hinton's deep belief networks won't.
http://gpickard.wordpress.com/2008/11/11/well-heres-my-actual-thesis/ http://gpickard.wordpress.com/2008/11/11/well-heres-my-actua...
- bravura 15y agoI believe the problem is the training algorithm, backprop, not the model (NNs). "unstable fixed point attractor of Backpropagation through Time", as you said. Ilya Sutskever in Geoff Hinton's lab has had great success recently using hessian-free optimization to train recurrent networks. He has trained character-level RNNs on Wikipedia, and they can generate very long sequences of quasi-plausible text. In particular, it seems like they can remember many symbols. See his most recent pubs: http://www.cs.toronto.edu/~ilya/pubs/ http://www.cs.toronto.edu/~ilya/pubs/