15 ms·
An Introduction to Support Vector Machines
- aaron-lebo 9y agoFor a recent practical example of their usefulness: This paper presents the Militarized Interstate Dispute (MID) 4.0 research design for updating the database from 2002-2010. By using global search parameters and fifteen international news sources, we collected a set of over 1.74 million documents from LexisNexis. Care was taken to create an all-inclusive set of search parameters as well as a sufficient and unbiased list of news sources. We classify these documents with two types of support vector machines (SVMs). Using inductive SVMs and a single training set, we remove 90.2% of documents from our initial set. Then, using year-specific training sets and transductive SVMs, we further reduce the number of human-coded stories by an additional 21.6%. The resulting classifications contain anywhere from 10,215 to 19,834 documents per year. http://steventlandis.weebly.com/uploads/1/2/1/4/12144932/dorazio_et_al_2012_here.pdf http://steventlandis.weebly.com/uploads/1/2/1/4/12144932/dor...
- ice109 9y ago5 years isn't a recent example. my impression is that deep nets ate everyone's lunch (including svm).
- theikkila 9y agoNot at all, deep nets are difficult to train and they need lot's of processing before they learn same kind of classifying features than ie. SVM has. So yes you can simulate SVM with deep networks but usually it's not very good solution. SVM can be also used as part of the neural network such as in classifying layer
- dmreedy 9y agoDeep nets ate everyone's hype. The lunch is still there. SVMs have many advantages over ANNs that recommend themselves to practical applications still.
- ska 9y agoThey really aren't appropriate for the same set of problems, even if there is some overlap. There is a lot of buzz around deep learning right now, but the SNR isn't great.
- deleted 9y ago[deleted]
- nilkn 9y agoOn the research frontier, this is pretty much true. But, for what it's worth, I'm spearheading some machine learning efforts at my current company, and most of my initial production models have not been deep networks but rather classical approaches like boosted trees or SVMs. Actually, gradient boosted trees in particular are one of the most powerful general-purpose models out there, and there are some really fantastic distributed implementations available now that routinely win Kaggle competitions (xgboost, Spark MLLib's version). I will note though that the problems I'm tackling do not involve any image processing or recognition. Convolutional networks really have completely dominated that area both in research and practice in the last few years.
- deepnotderp 9y agoI'd just like to note that instead of creating additional animosity between SVMs and deep nets, you could use both together. SVMs with hinge loss can be Yet-another-layer (tm) in your deep net, to be used when it provides better performance.
- strebler 9y agoThat's a great point. Fundamentally, if you look at something like a CNN, what it's really doing is producing a feature descriptor based on the input image. One can easily use that feature descriptor in a classic SVM, alongside (or instead of) SoftMax.
- deepnotderp 9y agoYup, in fact, the universal feature extraction is what allows imagenet pretraining to work well on lung cancer images. One nitpick though, ConvNets can absolutely be used to do "thinking" and more than just feature extraction. For example, fully convolutional networks can be extremely competitive with FC-layer based nets.
- JustFinishedBSG 9y agoYou may be interested by https://arxiv.org/abs/1605.06265 https://arxiv.org/abs/1605.06265 http://papers.nips.cc/paper/5348-convolutional-kernel-networks.pdf http://papers.nips.cc/paper/5348-convolutional-kernel-networ...
- sivvy 9y agoCould you explain in a bit more detail how you would integrate an SVM layer into a DNN? The kernel matrix depends on all samples, while at training time you would only have access to those in the minibatch.
- IanCal 9y agoThe simplest is to pop it on the top. Run you DNN to reduce your input down to a nicer cleaner smaller dimensional output, then plop an SVM on top for classification.
- chrischen 9y agoCan someone explain this part: Imagine the new space we want: z = x² + y² Figure out what the dot product in that space looks like: a · b = xa · xb + ya · yb + za · zb a · b = xa · xb + ya · yb + (xa² + ya²) · (xb² + yb²)
- nerdponx 9y agoTake the 2nd line, and drop in the definition of "z" from the first line. You get the 3rd line as a result.
- tgeery 9y agoI have the same problem. Where did a & b come from? Which two vectors are we taking the dot product of? And how is this less expensive?
- Longwelwind 9y agoIn the decision function of an SVM, you compute the scalar products of the support vectors (points that are on the margin of your hyperplane, or more precisely, the points that constrain your hyperplane) and your new sample point: x· sv The "z" the article defines is a new component that will be taken into account in the scalar product. A more mathematical way of seeing that is that you define a function phi that takes an original sample of your dataset, and transform it into a new vector. In our case, we simply add a new dimension (x3) based on the two original dimensions (x1, x2) that we add as a third component in our vector: phi(x) = [x1, x2, x1² + x2²] The scalar product we will have to compute in our decision function can then be expressed as (this is the a and b in the article, i.e. the sample and the support vector in our new space): phi(x)· phi(sv) The SVM doesn't need phi(x) or phi(sv), but the scalar product of those two numbers. The kernel trick is to find a function k that satisfies k(x, sv) = phi(x)· phi(sv) and that satisfies the Mercer's condition (I'll let Google explain what it is). Your SVM will compute this (simpler) k function, instead of the full scalar product. There are multiple "common" kernel functions used (Wikipedia has examples of them[1]), and choosing one is a parameter of your model (ideally, you would then setup a testing protocol to find the best one). [1] https://en.wikipedia.org/wiki/Positive-definite_kernel#Examples_of_p.d._kernels https://en.wikipedia.org/wiki/Positive-definite_kernel#Examp...
- Asdfbla 9y agoI remember that only a few years ago, in a computational statistics class I took the lecturer mentioned how SVMs (and Random Forests) have largely replaced neural networks. How things can change so quickly... I always liked SVMs for the elegance of the kernel trick, but I guess choosing the right kernel functions and parameters for them wasn't that much easier than training a neural net either.
- joe_the_user 9y ago(All this from my rough, amateur understanding), SVMs are more or less equivalent to linear regression in a "feature space" and also equivalent to shallow neural network (~2-3). This means their size more or less increases with the amount of data they are attempting to approximate. And this means they don't do well scaling to truly huge data sets. Deep nets pulled ahead of SVMs at the point people figured out how to train them on truly huge data sets using GPUs, gradient descent (and an ever increasing arsenal of further tricks - all the schemes together are mindboggling to read about). This was basically because the deepness of a deep neural net means that it's size isn't as prone to increase with the size of data. I don't really know why SVMs haven't been able to scale to a multi-layer approach though I know people have tried (someone has tried just about everything these days). Part of the situation is leveraging simple code with GPUs still may be the most effective approach.
- genericpseudo 9y agoClose but not quite. The difference between (soft) SVM and a kernel linear classifier is choice of loss function; SVM minimizes hinge loss, linear regression minimizes squared loss. (Choice of different loss functions will also give you Elastic Net, LASSO, logistic regression. From an engineering point of view I tend to think of the entire class as being different flavors of "stochastic gradient descent", in the spirit of Vowpal Wabbit etc.)
- CuriouslyC 9y agoIf you like SVMs, you should check out gaussian processes (GP). They work with covariance kernels similar to SVM, but the result is fully Bayesian. With most modern GP packages you can even set priors on your kernel and mean functions, then use either optimization or markov chain monte carlo to select optimal values. The only downside to GPs is that they are O(N^3) in time, so not applicable to big data. There are stochastic GPs that approximate using batch learning, but they're not as polished.
- idrism 9y agoFor anyone interested in SVMs (and other introductory Machine Learning concepts), Udacity's intro course is really good: https://www.udacity.com/course/intro-to-machine-learning--ud120 https://www.udacity.com/course/intro-to-machine-learning--ud...
- yamaneko 9y agoThis class from MIT taught by Patrick Winston is also a great resource: https://www.youtube.com/watch?v=_PwhiWxHK8o https://www.youtube.com/watch?v=_PwhiWxHK8o At the end of the class, he also gives some historical perspectives, like how Vapnik came up with SVMs.
- chestervonwinch 9y agoWith SVM, you often must perform a rather larger grid search over kernels and kernel parameters. It seems like no matter the model, we can't avoid the hyperparameter problem -- although boosting and bagging meta-methods come close. It would be nice if we could quantify the complexity of a dataset and match this to a model with similar complexity. I imagine that it's hard (or impossible) to decouple these two complexity quantifiers, however.
- currymj 9y agoSince neural nets are winning at the moment, it's easy to see SVMs as an underdog, being ignored due to deep learning hype and PR. This is kind of true, but it's worth noting that 10-15 years ago we had the exact opposite situation. Neural nets were a once promising technique that had stagnated/hit their limits, while SVMs were the new state of the art. People were coming up with dozens of unnecessary variations on them, everybody in the world was trying to shoehorn the word "kernel" into their paper titles, using some kind of kernel method was a surefire way to get published. I wish machine learning research didn't respond so strongly to trends and hype, and I also wish the economics of academic research didn't force people into cliques fighting over scarce resources. I'm still wondering what, if anything, is going to supplant deep learning. It's probably an existing technique that will suddenly become much more usable due to some small improvement.
- rm999 9y agoThis is true, but only in the academic research world. SVMs had relatively little success on practical problems and in industry, so they never built up the kind of standing that neural networks did. Even in 2003-2005 - arguably the peak time for SVMs - neural networks were much better known to almost everyone (industry practitioners, researchers, and laypeople) than SVMs. What frustrates me is that people who are starting out in machine learning often never learn that linear/logistic regression dominates the practical applications of ML. I've spoken to people who know the in-and-outs of various deep network architectures who don't even know how to start with building a baseline logistic regression model.
- stared 9y agoFor SVMs I really like this intro: https://generalabstractnonsense.com/2017/03/A-quick-look-at-Support-Vector-Machines/ https://generalabstractnonsense.com/2017/03/A-quick-look-at-... (with hand drawings!)