9 ms·
Show HN: Igel – A CLI tool to run machine learning without writing code
- ericpts 6y agoUsually the hardest part of a learning pipeline is data gather and cleaning; once it is in a suitable format (such that it is easy to create a structured CSV file), the training part is probably the easiest part: just a few lines of Python code.
- devaler 6y agoAnd, arguably, data cleaning is the most overlooked part.
- kthejoker2 6y agohttps://images.app.goo.gl/ZrvQDrMtKxbnMo2C9 https://images.app.goo.gl/ZrvQDrMtKxbnMo2C9
- nidhaloff 6y agoI agree. That's why some usually used pre-processing methods were implemented in the stable release.. and more is yet to come
- TheRealPomax 6y agoAll parts of a learning pipeline are hard if you want to do it right. Gathering, weeding, and binning your data is meticulous and hard work, and while "a single run" is trivial, rerunning it over and over with new parameters or even a completely different model because the outcome made no sense whatsoever is not. If updating a YAML file and hitting "run" makes that other "hardest part of learning" easier: hurray!
- crehn 6y agoFrom a purely UX perspective, there’s a huge difference between “no lines of code” and “a few lines of code”.
- master_yoda_1 6y agoThis is going too far, looks like nobody understand machine learning at hacker rank
- henvic 6y agoWell, it's probably almost the truth.
- lioeters 6y agoOff topic, but I recently found myself using the phrase "pretty much exactly". I realized it's nonsense, because the first part contradicts the meaning of "exact". I vowed to never use that phrase again. I feel the same about "probably almost the truth" (not a criticism, just a thought) - unless truth is a range (100% true to 100% false) rather than a binary (either true or false).
- codetrotter 6y agoTruth is a range though, for the vast majority of things.
- st1x7 6y ago"Automate everything" is a disingenuous claim. You're simply replacing a couple of lines of scikit-learn with a couple of lines of your CLI tool. There is pretty much no benefit to using this.
- nidhaloff 6y agoWell. The user is writing a description in a human readable format. Then, the tool will take that description and start running the pipeline. From data reading, preprocessing until creating and evaluating the model. If this isn't automation, please define automation for me. Also there are new features that I'm working on. The stable release was done this week.
- hobofan 6y agoFrom what I can tell it's a declarative framework (while most other common ones are imperative). And as generally with the tradeoffs of a declarative approach, if the data/model is easy the baked in assumption require less input, while if it's more complex your config files will be equally verbose and/or you will run into a wall. I don't see much automation there, just abstraction.
- nidhaloff 6y agoInteresting opinion. Well I must disagree in some points. First, yes sure the tool uses declarative paradigm, which is the goal of the project. If you want to use ML without writing code, then you will certainly not want an imperative framework. Second, I must disagree that most other common frameworks are imperative. I would say it's a mix of declarative & imperative but certainly not imperative. Finally, it's interesting how you see this as a just abstraction tool. I find other ML frameworks are more about abstraction since you are focusing on building your model but all details are hidden from you using the framework. Sure, igel is also about abstraction but to say it's JUST abstraction? mmm I find it not quite right, instead it's more about automating the stuff that you would write yourself using other frameworks. At the end of the day, we all have different opinions and feedback is important ;)
- 6y ago
- mk_chan 6y agoI find it difficult to believe anyone who can use the models listed on the repository effectively would have any difficulty using scikit themselves. Abstracting scikit out into a configuration file only very slightly simplifies the actual code involved but I can see this being useful for some non technical users who don't care about the code and just know the ML terms.
- nidhaloff 6y agoIt's not about that someone will have difficulty using sklearn. It's more about how clean the approach is if you have all your configs in a yaml file and you can change things very easily/quickly and rerun an experiment. I'm working with data & ML models everyday and it became overwhelming when my codebase is large and I want to change small things and re-run an experiment. Also It would be great to not lose much time writing that code in the first place (although it's easy to do), if you want a quick and dirty draft. The thing is, it is much cleaner if you have your preprocessing methods and model definition in one file. However, there are other features that will be integrated soon, like a simple gui built in python
- iamflimflam1 6y agoThis is a good point, something that I've been struggling with in my own personal projects is keeping track of parameters as I tweak and play with hyper-parameters and model structures. A few parameters are fine, you can pull them out into constants, but you quickly end up with a lot of variables to keep track of.
- mxscho 6y ago> A machine learning tool that allows you to train/fit, test and use models without writing code I recently had a discussion about the requirements that a text file format (like YAML) has to fulfill to be considered "code". :)
- nidhaloff 6y agoHi, and what was the result/conclusion of the discussion? I'm interested in your finding, is it considered code or not :D
- mxscho 6y agoWell, we came to the conclusion that there is no hard border and therefore good answer to that question. But we also agreed that it's not the most important factor whether it's code, a graphical user interface or a command line interface to make a tool usable for a lay person. What's more important is that the entry point is easy, and that the complexity and flexibility is abstracted away in layers that do not have to be fully understood from the beginning, so that the learning curve is not too steep. Of course, my first post was not meant to be criticism of the project, just some pseudo philosophical thoughts that crossed my mind when reading that sentence. Sorry for being too off topic with that. :)
- marcinzm 6y agoI feel like the places where a non-technical user would be building non-trivial models also have the money to pay for one of those commercial GUI drag boxes around tools.
- eyeball 6y agohttps://pycaret.org/ https://pycaret.org/
- somurzakov 6y agothank you for the link!
- kamhh94 6y ago"non-technical" and "cli tool" sound like an oxymoron. But if you hide the yaml config behind a UI i guess it can pass for "non technical".
- nidhaloff 6y agoalready working on a gui that users can launch using a command. You can check the issues list
- TheRealPomax 6y agoMake it (double) clicking an icon on a desktop/app list, and you have a winner. The moment a terminal is needed, you've lost the non-technical crowd (and some of the technical crowd, even)
- joshspankit 6y agoWho else thought this was something that would turn AI loose on your bash commands, and automate everything in your CLI?
- _frkl 6y agoIf you are disappointed, there's mcfly, which does things with your shell history and ML: https://github.com/cantino/mcfly https://github.com/cantino/mcfly :-)
- djhaskin987 6y agoSo a lot like weka then. https://www.cs.waikato.ac.nz/ml/weka/ https://www.cs.waikato.ac.nz/ml/weka/
- jeroenjanssens 6y agoReminds me of SKLL: https://github.com/EducationalTestingService/skll https://github.com/EducationalTestingService/skll
- tpetry 6y agoTo be really automatic the only thing i should need to do is feed an csv, correct the suggested data types and then run all algorithms on the data with the information at the end which has been the most effective and which i should further optimize.
- nidhaloff 6y agothis is a great feature! Thanks for the feedback
- fractionalhare 6y agoNo, it's emphatically not a great feature, and it's not clear to me the commenter was recommending that so much as making a nit. Please don't automate the process of choosing and running algorithms on a single sample of data, it's unsound experimental design that undermines your results. If you insist on doing it anyway, at minimum you will need to automate an initial assessment of the sample data to determine if it has a suitable size and distribution to allow you to adjust the significance of results for the number of tests you're running, and partition the data into smaller subsamples.
- nidhaloff 6y agoHi, thanks for your comment. I actually understood that he meant something like a hyperparameter search/tuning using cross validation (at least that what came in my mind).
- fractionalhare 6y agoCross validation would be good! I think if you build this in you could automatically run a few heuristics to see if the data can be partitioned, or maybe just prompt the user for another sample of the data with the same distribution.
- tpetry 6y agoParameter tuning and algorithm selection! I just don’t want to manually start 5 different runs of algorithms i believe which could work good on the data and manually compare the results. And maybe i was too lazy to run the 6th algorithm which now performs much better. But to be sure, every test should be done with k-fold cross validation. The decision whether to split the training set should not be chosen by the user. It‘s crucial that this is a must!
- alphachloride 6y agoI think it can be a useful tool for automation of very standardized ML tasks. However: It's a command line tool that is also intended for non-technical folks. I sense a contradiction. That doesn't even speak to the requirement of understanding all these ML algorithms so I can specify them in the config file, or understanding YAML format, or data curation. At this point it would be easier to write the python code - especially scikit-learn which is a very well-documented library.
- nidhaloff 6y agoHi, I want to clear up some points. First, it is not intended for non technical folks, this was never claimed! However, even if it was, we are currently working on a gui, where (non technical)users can run it by writing a simple cmd in the terminal. Second, I'm a technical user, in fact this is my daily work and we build this tool for reasons that were mentioned in the docs/readme, so you can check it out. Third, you mentioned understanding YAML Format. Really? I mean yaml is the most understandable format any person can understand. I can never imagine that a person cannot learn yaml in 30 min at most. Finally, yes sklearn is great and well documented but did you checked how many libraries are out there that represent basically a wrapper to make it easier/abstracter to write sklearn code? you ll be surprised. As discussed in the official repo & docs, it is a much cleaner approach to gather your preprocessing & model definition parameters/configs in one human readable file/place, where you can manipulate it easily. Re-run experiments, generate drafts, building proof of concepts as fast as possible, than to write code. At the end of the day, we all have different opinions, you can still write code of course. The tools are there to help.
- howmayiannoyyou 6y agoTerrific! Keep pursuing this and ignore critics. What you're doing is important b/c ML is just out of reach of a big percentage of developers and technical lay people. It will take time to get your approach right, but it will make a difference. As a suggestion - provide more real-world examples (eg. business, sports, etc) so that users can tinker with your samples as pathway toward learning. Please don't give up on this. Great job.
- nidhaloff 6y agoHi thanks a lot. I received positive interactions on github from the community, however, your comment is the first encouraging feedback I ve got here :D so, I appreciate it. I will take your suggestion into consideration. You are right, there should be more real-world examples that will help users get started and see how this can be useful. The thing is, I started the project two weeks ago, so it still relatively new. I ve been coding day n night because the idea got me excited. I published the first stable release this week. However, there are new features that will be implemented in the next releases.
- musingsole 6y agoIf the project is only 2 weeks old, all the more reason to ignore any critics. Particularly here where people are likely to criticize a baby in the crib for not working on coding projects outside of naptime.
- reagent_finder 6y agoWell, I mean that baby doesn't have a functional colon yet so putting semicolons everywhere just makes perfect sense.
- toxik 6y agoI feel like if there is one thing that works on a baby, it is the colon...
- 6y ago
- foolfoolz 6y agoi think long term this is the future of ML. it’s like a database. every engineer needs to know when to use one and how. not every engineer needs to be able to write a database
- toxik 6y agoCoincidentally, poor understanding of your tools (especially databases) seems like a huge source of frustration and pain for anyone involved in software development.
- IdiocyInAction 6y agoThis is already how most ML in production works though. Noone writes their own NN, optimizer or even linear regression and for good reasons.
- mcint 6y agoThank you for sharing! I was thinking about starting something like this, and had reached out to datasette [1] / Simon Willison for advice on starting and maintaining a project. [1]: https://simonwillison.net/2017/Nov/13/datasette/ https://simonwillison.net/2017/Nov/13/datasette/
- asimjalis 6y agoWhat does “IGEL” stand for? I couldn’t find it in the documentation.
- phil294 6y agoNot sure if there is any deeper meaning hidden behind it, but it is the German word for Hedgehog (pronounce as: "Eagle")
- asimjalis 6y agoInteresting. That makes sense since the logo is a hedgehog. What’s the connection though I wonder.
- hashmush 6y agoInteresting, igel is Swedish for leech (Egel in German). The Swedish word for hedgehog is instead igelkott. In short, Egel = igel and Igel != igel... TIL
- nidhaloff 6y agoIndeed interesting :D
- nidhaloff 6y agoIt's a german word and means Hedgehog. It's funny we were discussing a name for the project and we wanted to make an abbreviation from some words that make sense, so we started throwing ideas spontaneously. At the end we wanted to make an abbr for these words: "Init, Generate, Evaluate Machine Learning". IGEL made sense for us then since it's a german word too. Easy to say, type and remember ;)
- tommica 6y agoThis is really cool!
- iamflimflam1 6y agoKeep going with this, I think you are onto something.
- jkmcf 6y agoThis is the giant's shoulders I like to stand on! Need to see more projects abstracting away the hard stuff (I'm looking at you, GUI libraries!)
- nidhaloff 6y agoThanks for your feedback. Stay tuned, we are working on an integrated gui tool written in python too.
- xtracto 6y agoThe great thing about this is that it is directly usable in a gui . Someone will build a gui and make it even more accessible. I love it and plan to use it in my data as it is.
- nidhaloff 6y agoWe are already working on a gui ;) stay tuned
- mjgs 6y agoGreat idea - I’ve been waiting for a good cli tool for machine learning, saves the hassle to have to learn python and also can use with other existing shell tools.
- zatel 6y agoThis is so cool! I know the answer is to just write what I'm describing myself but does anyone know of an existing way to find the best SciKitLearn algorithm for a particular problem. Like if I want to find the regression fit is there a way to just pass in the data and have it trained,tested on all of the regression algorithms in SKLearn? My current workflow is to just pick a handful of algorithms that sound like they should be good for the problem at hand and try each one of them manually. Igel seems like a step towards making this sort of thing possible if another tool doesn't exist already.
- nidhaloff 6y agoHi, we should be careful with the feature you are talking about. The results from all machine learning algorithm can be very misleading and probably some models will overfit the data. So, if you throw some data and fit all machine learning models on it and then compare the performance. You will probably receive misleading values since different models require different tuning approaches. It's not as easy as you said it, you can't just feed data (also depends on the data) to models and expect to get the best model at the output. One approach I can think of here is to integrate cross validation and hyperparameter tuning with your suggestion. However, I can imagine that this can be computationally expensive. I will take it into consideration as an enhancement for the tool. Thanks for your feedback
- craftinator 6y agoHey, I really appreciate your answer to this question. As I was reading the question, red flags started popping up in my mind about the risk of overfitting when using the ensemble approach, and I think your response was spot on for how an ML researcher would go about it! Most ML professionals I've talked to have been really against making a user friendly ML suite because of how easy it is to misuse these algorithms.
- zatel 6y agoThank you for explaining this more indepth. I should have been more specific with my original comment, I did intend cross validation and hyper parameter tuning as inclusing to the automatic feature I was describing. These operations certainly are computationally expensive, a recent hyperparameter tuning operation locked up my laptop for 3 days but this seems to be the case for any similar operation. The only approaches I've come across so far to overcome it are things like converting the data to smaller sizes (which seems outside the scope of this tool) and some way to batch the data so that it can be "paused" and resumed as needed. Thank you again for creating Igel.
- deleted 6y ago[deleted]
- sthatipamala 6y agoAwesome! How do you compare this to https://github.com/uber/ludwig https://github.com/uber/ludwig, which also has a YAML-based cli for ML?
- nidhaloff 6y agoWow! this is great! I didn't know that such a tool exists, thanks for posting it here. The python & AI community are moving really fast, it's crazy! However, It looks like the ludwig tool is about deep learning and not ML, or am I wrong? It looks like there is no support for ML models or am I missing something I didn't try it yet, I just read the get started section but looks really great for training deep neural networks.
- JshWright 6y agoI'm more familiar with a different "i-gel" (which does a similar thing for emergency airway management as this does for ML, allowing less trained users to still achieve "advanced" results) https://www.intersurgical.com/info/igel https://www.intersurgical.com/info/igel
- btach 6y agoGlad I wasn't the only one whose first thought went to protecting someone's airway when some else fails!
- fakedang 6y agoI remember how I first got interested in ML and DL. I did not know the a lick of programming ML in Python or whatever language was out there. I simply began by using Matlab's Neural Network and Machine Learning toolboxes and playing around on them. That turned into real coding interest on Matlab, which carried on to Python, so on and so forth. In a sense, I rediscovered programming because of those toolboxes. What you're doing is great stuff and I hope it encourages a lot of folks to play around just as I had, just to get started.
- desilinguist 6y agoGreat idea! We had a similar idea back in 2015 with SKLL[1]. We are still actively maintaining it and it’s definitely been helpful to many folks, including many outside our organization, over the years! Wishing you the best! [1] https://github.com/EducationalTestingService/skll https://github.com/EducationalTestingService/skll
- tyteen4a03 6y agoNot sure if you're aware but https://igel.com https://igel.com (thin client manufacturer) is a thing.
- nidhaloff 6y agoOMG this is funny :D I wasn't aware about this, thnx. Looks like it's a thing and the company can even skyrocket in the near future! I only read the about us section there. I also noticed they pronounce the name differently. Actually, igel is a german word and is pronounced "Eagle" but they pronounce it as I-jeel
- czhu12 6y agoI really love seeing projects like this. My humble two cents is that in past projects I've often found the process to create the CSV that goes into a model is often much more time consuming and error prone than actually training a model. I really love the dataset operations that are provided. Do you have any thoughts about what a gold standard library for data preprocessing would look like? If you have any plans to move further in that direction? Any projects that you find compelling in that space?
- nidhaloff 6y agoHi, thanks for your feedback. hmm any machine learning project has to start with a dataset. Most of the time you will have to construct it (or take an existing one and update it) manually. Sure there are tool that generate a dummy dataset for you but that would be just for playing a world and certainly not for a real world use/production. We are actually working on adding support for text, excel and json format in igel. We already implemented some of the famous preprocessing methods in igel, which you can use by providing them in the yaml file. Now about preprocessing libraries, I personally use numpy, pandas and some of sklearn functionality to preprocess data. Furthermore, I use matplotlib and seaborn for some visualisation & further analysis.
- tinyhouse 6y agoNice work! This is def gonna be useful for a lot of people. Companies can benefit from using something like this too. It helps when all your people use the same tool to build and run models and just need to share yaml/json files. The alternative is each group has its own scripts and sharing is harder. The simplicity of such tools has a tradeoff though. It takes away some of the flexibility. Also, actually writing code helps people learn about ML so there's benefit of doing the hard code when building ML models. But that's not the target audience of this project so that's OK. Good luck!