3 ms·
I haven't tried it, so I can't say for sure, but my intuition is that learning to predict the latent factors is a much less 'complex' task than learning good fe
by benanne 12y ago
I haven't tried it, so I can't say for sure, but my intuition is that learning to predict the latent factors is a much less 'complex' task than learning good features to reconstruct the input (i.e. the spectrograms), in terms of the required capacity of the model.
With a purely unsupervised approach, you are basically wasting capacity on modeling aspects of the data that are relevant for reconstructing the input, but not for solving the task at hand. For example, the model doesn't have to care about precise pitches and timing, because those are not relevant for recommendation (and latent factor prediction) anyway. That means no model capacity is wasted on these things. With a fairly complex task such as this one, I think that probably makes a big difference.