3 ms·
I find it interesting that no group (to my knowledge) has tried something similar to [Do Deep Networks Need to Be Deep?](http://arxiv.org/abs/1312.6184 http://a
by kastnerkyle 12y ago
I find it interesting that no group (to my knowledge) has tried something similar to [Do Deep Networks Need to Be Deep?](http://arxiv.org/abs/1312.6184 http://arxiv.org/abs/1312.6184) for ImageNet scale networks. There have been several results which show that the knowledge learned in larger networks can be compressed and approximated using small or even single layer nets. Extreme learning machines (ELM) can be seen as another aspect of this. There have also been interesting results in the "kernelization" of convnets [from Julian Mairal and co.](http://arxiv.org/abs/1406.3332 http://arxiv.org/abs/1406.3332) that, accompanied by the stong crossover between Gaussian processes and neural networks from back in late 90s, point to the possibility of needing different "representation power" for learning vs. predicting which may lead to the ability to kernelize the knowledge of a trained net, ideally in closed form.
I am doing some experiments in this area, and would encourage anyone thinking of doing hardware to look at this aspect before investing the R&D to do hardware! If this knowledge can really be compressed it could be a massive reduction in complexity to implement in hardware...
I am a bit biased on this topic (finishing a talk about this exact topic for EuroScipy now) but I find the connections interesting at least.