4 ms·
Technically what is visualized here isn't "the kernel trick". This is the general idea of how nonlinearly projecting some points into a higher-dimensional feat
by lliiffee 16y ago
Technically what is visualized here isn't "the kernel trick". This is the general idea of how nonlinearly projecting some points into a higher-dimensional feature space makes linear classifiers more powerful. You can do this with out SVMs. Just compute the high-dimensional features corresponding to your data, then use logistic regression or whatever. Trouble is, if the higher-dimensional space is really big, this could be expensive. The "kernel trick" is computational trick that SVMs use to compute the inner product between the high-dimensional features corresponding to two points with out explicitly computing the high-dimensional features. (For certain special feature spaces.)
But this is definitely a cool visualization of the value of feature spaces!
- moultano 16y agoCan you explain more about how the kernel trick specifically works?
- lliiffee 16y agoBasically, the idea of feature spaces is to blow up the data into high dimensions. So, we use x' = f(x) as our data, instead of x. It turns out that in lots of machine learning algorithms (notably SVMs), you end up only needing inner products between different data elements. That is, we need to compute x^T y for two data elements x and y. In feature space, we need could compute this by doing f(x)^T f(y). However, it turns out that for certain feature spaces (like polynomials) one can compute the number f(x)^T f(y) quite quickly with out ever explicitly forming the big vectors x' or y'.
- moultano 16y agoYou stopped just short of the explanation I was hoping for. :)
- deleted 16y ago[deleted]
- lliiffee 16y agoTry section 7 of these notes: http://see.stanford.edu/materials/aimlcs229/cs229-notes3.pdf http://see.stanford.edu/materials/aimlcs229/cs229-notes3.pdf
- gaika 16y agosee http://videolectures.net/mlss09uk_schoelkopf_km/ http://videolectures.net/mlss09uk_schoelkopf_km/ - kernel methods in general, not limited to SVM
- iskander 16y agoWhile the computational benefits are certainly nice, I think that Mercer's theorem is much deeper than that. It says that any [positive definite](http://en.wikipedia.org/wiki/Positive-definite http://en.wikipedia.org/wiki/Positive-definite) function is an inner product, though the space of that inner product might be infinite. The consequences of this go much further than just speeding up algorithms-- we can take plain old linear methods like SVMs and apply them to any data type on which we can define a kernel. This is huge, because it allows machine learning to break free of the vector world in which the algorithms are classically defined. Want to classify gene sequences? You don't have to hack them into a fixed-width vector, just define a string kernel for genes. Want to classify fragments of lambda calculus? Just define an appropriate tree kernel. It's wonderfully powerful.