5 ms·
I'm the author -- this is a technique to use a trained random forest as a kernel function for similar data. It can be used as a sort of primitive transfer learn
by RMarcus 9y ago
I'm the author -- this is a technique to use a trained random forest as a kernel function for similar data. It can be used as a sort of primitive transfer learning, or just as a way to get a very high quality kernel for a (labeled) dataset.
Happy to answer any questions!
- danharaj 9y agoIn your mnist example, what is the relationship between all of the kernels you get for training to distinguish each pair of digit classes? Is there a natural way to intersect random forests or their kernels to get at their "common knowledge"? Is there a way to union them to get their "collective knowledge", and if so, how does that fare against an algorithm that learns all the classes simultaneously?
- RMarcus 9y agoThese are all great questions. The simplicity of the kernels produced might make them very amenable to analysis. I'm not sure I have the technical skills to do it though...
- stochastic_monk 9y agoVery interesting; I've ignored random forests more than I should. Thank you! I won't comment on your methods except to say that comparing classification over linear PCA against a kernel PCA in linearly unseparable data isn't exactly fair, and I think that providing an SVM performed on a different kernel PCA decomposition or a kernel SVM itself would be more illustrative. (Is your code available?) I generally think of neural networks as enormous meta-kernels. (Composites of <composites of...> kernels) This generally leads me to think of ways that kernels can be turned into neural network layers. Great work has been done turning powerful tools like Random Fourier Features/Kitchen Sinks into layers in neural networks (e.g., Alex Smola's Deep-Fried ConvNets [https://arxiv.org/abs/1412.7149 https://arxiv.org/abs/1412.7149] and Choromanski's Structured Adaptive/Random Spinners[https://arxiv.org/abs/1610.06209] https://arxiv.org/abs/1610.06209]). Deep Forest [https://arxiv.org/abs/1702.08835 https://arxiv.org/abs/1702.08835] is a method which claims to work well, but it's somewhat odd; It's not quite what I would have imagined. My biggest criticism of random forests is that the more expressive models are more memory-hungry and expensive at both training-time and run-time than many comparable methods. But bounded-complexity trees with smart implementations seem to be a lot more useful and have broader applications than I give them credit for.
- RMarcus 9y agoI agree the comparison could be a lot more fair. My goal was to show how adding in labels can give you better principle components, but I should've used an RBF kernel, or some kind of unsupervised kernel, instead of linear PCA. I really should go and make the code available. I spent a lot of time writing the RF kernel in Cython to get it to be fast. But it is so fragile and incomplete, I'm not sure I am ready to release and maintain it. Maybe I'll just post the scripts. Thanks for the links and insights about ANNs as combined kernel learners + classifiers. I'll check out the papers (I'm also surprised at what Deep Forest is. Not what I was expecting.)
- theSage 9y agoWhat were you expecting deep forests to be like?
- PaulHoule 9y agoDo you see this as being competitive with Siamese networks for one-shot learning?
- RMarcus 9y agoI have no idea :). I don't think this approach is state of the art, but I think the advantages are that RFs are lightweight when compared to ANNs.
- j2kun 9y agoDo you have any insights as to why random forests are better (are they better?) in this context (or any context?) than boosted decision trees?
- RMarcus 9y agoI've done no experiment, but I predict a boosted forest would perform poorly here. Since the first tree learns the majority of the information (and each of the following trees only fits the errors, and are therefore less "meaningful"). Since the kernel approach suggested here weights all the trees equally, the results probably won't be as good. Maybe if you gave the first trees more weight, though...
- j2kun 9y ago> Maybe if you gave the first trees more weight, though That is exactly what boosting does...and one would weight the kernel computation in the same way.