7 ms·
Machine Learning from scratch: Bare bones implementations in Python
- fnl 10y agoThis could become a fantastic resource for anybody who is teaching machine learning. One vital improvement suggestion to make that path attractive would be if the Jupyter notebook format were used. It would be easier to add more documentation and references. But in any case, thanks for sharing!
- bayonetz 10y agoBetter yet, write your own simple versions. That's what we did in undergrad and grad classes in AI, ML, Neural Nets, etc. There is nothing like building one yourself, even a simple model like kNN!
- victor106 10y agoWould you suggest any books/resources to learn the theory behind these implementations so a newbie can follow along?
- jumpCastle 10y agoIntroduction to statistical learning http://www-bcf.usc.edu/~gareth/ISL/ http://www-bcf.usc.edu/~gareth/ISL/
- make3 10y agoprobably this https://www.amazon.ca/Python-Machine-Learning-Sebastian-Raschka/dp/1783555130 https://www.amazon.ca/Python-Machine-Learning-Sebastian-Rasc...
- nkozyra 10y agoWhile a great book, most of this is just implementations of sklearn.
- asteinbr 10y agoI can recommend Applied Predictive Modeling by Max Kuhn: http://appliedpredictivemodeling.com/ http://appliedpredictivemodeling.com/
- saboot 10y agoPattern recognition and machine learning by Bishop is one of the canonical text books. It helps to have a linear algebra background, it includes a refresher though
- tnecniv 10y agoBishop is good but reads a little too much like a literature review sometimes. That may or may not be a problem depending on what you are looking for.
- victor106 10y agoThanks everyone for the suggestions...will check these books out
- capkutay 10y agoHow does that compare to Andrew Ng's course?
- schmit 10y agoOne quick comment: in general it is a bad idea to compute the inverse of a matrix (to solve a linear system). It's much better to compute the QR factorization or SVD instead (or simply call least square solver). See for example: https://www.johndcook.com/blog/2010/01/19/dont-invert-that-matrix/ https://www.johndcook.com/blog/2010/01/19/dont-invert-that-m...
- eriklindernoren 10y agoThank you for the feedback. :) I plan on fixing this soon!
- JustFinishedBSG 10y agoNot very important and for a learning project I find using cvxpy a better idea as it's more readable ( like you did ) but: Solving the full quadratic optimization problem for SVMs in basically impossible to do. You are forming an n^2 matrix, so I'm going to let you imagine what happens when n = 100 000. Using people use either approximation methods ( Incomplete Cholesky, Nystrom ) or do it exactly but iteratively ( SMO, Pegasos... ) I'm implementing them for class right now so it's still fresh in my head haha
- JustFinishedBSG 10y agoIt's usually even better to use iterative methods
- thearn4 10y agoGMRES is definitely my go-to these days. Though it is worth noting that direct methods do have a benefit of letting you quickly solve many successive linear problems involving the same matrix, but different right-hand sides. But iterative methods scale very well for large sparse problems that they are very often the only tool to consider. Block Krylov methods are a thing , but I haven't experimented with them yet.
- JustFinishedBSG 10y ago
- jogundas 10y agoVery cool! I have actually been planning to do exactly what you did, sir :)
- eriklindernoren 10y agoThanks! :)
- natch 10y agoNot sure what you had in mind, but a Python 3 version of this would be great!
- jogundas 10y agoThanks for the comment, will have this in mind!
- edshiro 10y agoNice! I have started brushing up my maths and reading about machine learning in general. Next step is to get my feet wet in the implementation. I think looking at your project can give me a good idea as to how to implement some of the most basic algorithms. Good luck!
- eriklindernoren 10y agoAwesome! :) Doing the nitty-gritty has been a great way of learning the limitations and benefits of using certain models for a given task. Good luck to you as well!
- ussser 10y agoCool! How long did it take to learn and implement these models?
- eriklindernoren 10y agoI have taken some ML courses during my university studies and have also done some model implementations in other programming languages. So I didn't have to start from scratch. But I have been working on this project for about three weeks now.
- onlyrealcuzzo 10y agoThis is awesome! I'm working on something similar for JavaScript. Definitely will be using yours for reference. Thanks, dude!
- eriklindernoren 10y agoThank you!
- mrcactu5 10y agosci-kit learn is excellent, but their implementations are a bit to complicated to learn from. this is for people who don't just want to tune parameters but build the whole thing from scratch I can buy buy a pie all the fix-ins from a bakery, or I can buy the ingredients myself, and make it to exactly my liking. it may not be a professional.
- eriklindernoren 10y agoGlad that you liked it!
- onvalleysilic 10y agoJust tried it with an equities dataset http://54.174.116.134/recommend/datasets http://54.174.116.134/recommend/datasets and it seems to have performed nicely. Great work!
- eriklindernoren 10y agoAwesome. Thank you!
- Winterflow3r 10y agoThis is really cool and inspiring!
- eriklindernoren 10y agoThank you! The response here has been amazing. :)
- SvenDowideit 10y agoDeliver and release stuf that people actually use. Or work on projects that do. Delivering value trumps painting every day
- mmrr88 10y agohttp://vschool.io/en/apply/ http://vschool.io/en/apply/
- mmrr88 10y agohttp://vschool.io/en/apply/ http://vschool.io/en/apply/
- thinkr42 10y agoThis is awesome!
- eriklindernoren 10y agoGlad you like it!
- compactmani 10y agoThis is a nice project. I think it would be great to add references used for the implementations and some tests that demonstrate they return what is expected (or perhaps the same result of sklearn maybe).
- eriklindernoren 10y agoThose are great suggestions. I will look into adding that.
- Jasamba 10y agoThis is impressive, and kindof exactly what I am in the process of doing. It's certainly the best way to get familiar with the internal workings of these methods than just tune parameters like an oblivious albeit theoretically informed monkey. How long did it take you to do them?
- eriklindernoren 10y agoIt certainly is! Thanks. :) I have been working on it for about three weeks now.
- JustFinishedBSG 10y ago> I have been working on it for about three weeks now. That's not a whole lot, you're quick
- dnautics 10y agoNice project! I'm doing something similar in julia, with the added advantage that as I build it the numerical types are variadic so I can play around with numbers that aren't IEEE FPs.
- f311a 10y agoSimilar project: https://github.com/rushter/MLAlgorithms https://github.com/rushter/MLAlgorithms
- joelberman 10y agoVery nice project! Learning stuff makes me happy.
- peter_retief 10y agoI feel happy to see your wonderful work you share so freely
- eriklindernoren 10y agoGlad you liked it!
- searchfaster 10y agoVery nice project! Very very useful for a ML beginner like myself. Thank you very much !
- sp4ke 10y agoAmazing, thanks for sharing :)
- opoooopopooo 10y agoFuck you
- imdsm 10y agoGreat resource, but it could be a phenomenal resource if you documented each method and explained how and why it does what it does. Don't get me wrong, having working code to play with is key, but when you don't fully grasp the concepts behind it, an explanation can become so valuable. That being said, you've included names, so research can be done. Great work and I hope you're enjoying it!
- eriklindernoren 10y agoGiven the amount of publicity this repository has gained I will make sure do put more work into documenting the code. And also add references. Thank you for your feedback!
- metaobject 10y agoIn your RandomForest implementation, on the line in fit() where you're building the training subsets to give to each tree, it appears that your bagging approach doesn't use 'sampling with replacement' strategy. idx = np.random.choice(range(n_features), size=self.max_features, replace=False) It would appear that the replace=False prevents the 'sampling with replacement' behavior usually implemented by bagging algorithms. Should the replace=False be changed to replace=True?
- eriklindernoren 10y agoThank you for your feedback! I have read up on the feature bagging part of the algorithm and I believe that you are correct. This is fixed in the latest commit.
- grzm 10y agoFor the future, if it meets the guidelines, this likely should have been a Show HN: https://news.ycombinator.com/showhn.html https://news.ycombinator.com/showhn.html
- jbrambleDC 10y agoThis is awesome. I am currently building a decision tree from scratch in Java and will use yours as a reference. One comment I have. in kNN, it is best to ensure that the neighbors list occupies O(k) space.
- ieils_lese_ 10y agoThe biggest entertainment award show of the year, https://www.linkedin.com/pulse/watchoscar-award-2017-live-stream-free-f1-500-nascar-frances-xiong https://www.linkedin.com/pulse/watchoscar-award-2017-live-st...
- screwston 10y agoA friend sent me a link to this - nice work, and I happen to be intermittently working on a very similar (and unfortunately similarly named) project - https://github.com/jarfa/ML_from_scratch/ https://github.com/jarfa/ML_from_scratch/. Check my commit history if you suspect me of copying you ;) I don't think I'll be implementing as many algorithms as you though, I should force myself to work on more projects outside my comfort zone.
- eriklindernoren 10y agoOh, cool! Haha, that's unfortunate. Maybe I should have done a better job finding an original name. Good luck to you anyways!