3 ms·
Are there any linear algebra books which use machine learning as the motivating example? Or a book which teaches linear algebra and machine learning together?
by paperwork 14y ago
Are there any linear algebra books which use machine learning as the motivating example? Or a book which teaches linear algebra and machine learning together?
- j2kun 14y agoSadly you need a hefty amount of it to get to machine learning that uses linear algebra in a nontrivial way.
- kylebrown 14y agoFunny, I just spent last night implementing a "principal compnents analysis" using numeric.js, for a character recognition project. PCA / SVD seems like a great motivating topic. btw, I couldn't use Singular Value Decomposition in numeric.js for PCA because the method, numeric.svd, uses the "thin" algorithm, and throws an error if there are more columns than rows. I calculate way more features (50-200+ columns) than I have training samples (rows, 10-30 written manually). without svd I had to use the "covariance method", which I guess can sometimes present approximation issues, but seems to be working well for me. The purpose of the PCA is dimensionality reduction (google "curse of dimensionality"). I had used Mahalanobis Distance as a p-value score to detect outliers (p-val < 0.05), and it worked well when there were only 6 features. Curse of dimensionality makes MD useless when there are 50 or 100 features, and PCA reduces them to 3-10 features which carry the most information with the rest approaching zero. So I do MD on the projected (reduced) features, and its working great. If I do it all over again I might try a "one-class SVM" (which, sadly, I had not heard of until only recently and very late in the project). SVM's are non-linear like most machine learning algos, but the linear PCA can still be used to complement other methods, eg to do pre-processing before feeding to a neural network.[1] For deeper linear methods, check out the PCA-related extensions like Fisher Linear Discriminants or Projection Pursuit. 1. "Many neural networks are susceptible to the Curse of Dimensionality though less so than the statistical techniques. The neural networks attempt to fit a surface over the data and there must be sufficient data density to discern the surface. Most neural networks automatically reduce the input features to focus on the key attributes. But nevertheless, they still benefit from feature selection or lower dimensionality data projections."
- kylebrown 14y agoJust wanted to share something I ran into the other day in the course of research on PCA, a very cool motivating example, sports-related.[1] On page 13, they take some NBA player stats and graph the projection of the third PC (y-axis) against the second PC (x-axis). The third PC is a negative correlation of rebounds with assists and steals. Turns out to sort players roughly by height, placing Karl Malone (6'8") and Mugsy Bogues (5'3") at opposite extremes. 1. Faloutsos et al, Quantifiable Data Mining Using Principal Component Analysis, 1997