4 ms·
These things you mention are conditional variables - from a modeling perspective it would be perfectly plausible to build - think something like text-to-speech
by kastnerkyle 11y ago
These things you mention are conditional variables - from a modeling perspective it would be perfectly plausible to build - think something like text-to-speech but instead of the sentence you input some characters representing the dynamics/feel of the desired song. You also have the issue of RNNs being deterministic - even sampling from the output softmax probably doesn't give enough to have interesting variations if you run the network multiple times from the same seed. Also choosing the argmax(prob(y | x)) at each step does not guarantee that the path generated is maximally likely. For that you need a beam search or something like it. But I don't think any of that is the key problem here.
In this case, an RNN meant for 1-of-K output is not well suited to outputting chords. It worked fine for single notes however! Check out this link from Joao Felipe Santos https://soundcloud.com/seaandsailor/sets/char-rnn-composes-irish-folk-music https://soundcloud.com/seaandsailor/sets/char-rnn-composes-i... . These are pretty cool - and he generated the titles to boot.
For chords you really can't model all possible combinations (2^88 for midi) naively, and you also can't really model notes independently - chords are highly structured and follow specific rules! Even bounding to only chords of up to 3 or 4 notes still makes a pretty large output space, which means more training data is needed, it is harder to optimize, etc. etc.
You should be much better off with some kind of conditional/factorized model strapped to the output of an RNN - this is the idea for RNN-RBM, RNN-NADE, LSTM-DBN, etc. You could also just try to model the audio representation directly using LSTM-GMM or VRNN, but this is pretty hard and an active area of research.