9 ms·
The image for logistic regression is hilariously wrong. It shows the sigmoid as a decision boundary. Also dont get hung up about no free lunch theorem. That is
by polkapolka 8y ago
The image for logistic regression is hilariously wrong. It shows the sigmoid as a decision boundary.
Also dont get hung up about no free lunch theorem. That is a great result in computer science theory with little practical impact: just pick neural networks for unstructured and GBDT for structured data. For the vast majority of real-life problems (not all possible problems) these are the single best algorithms.
- deleted 8y ago[deleted]
- 1024core 8y ago> just pick neural networks for unstructured data ... except that NNs require a large amount of training data.
- polkapolka 8y agoSure there are constraint and pros and cons, but still: pick a neural network for unstructured data. Can always unsupervised pretrain and fine tune on a tiny dataset.
- swsieber 8y agoWhat qualifies as unstructured data? Would you consider text content to be structured or unstructured? (e.g. for classification of documents)
- polkapolka 8y agoText data traditionally seen as unstructured. Try a simple MLP.
- yorwba 8y agoGeneral-purpose machine learning algorithms require a large amount of training data. There's no magic that can make accurate predictions without any data to base them on. If you don't have enough data, the information needs to come from somewhere else. If you have a large amount of data on a similar problem, you can try transfer learning to learn shared properties and only fine-tune the domain-specific stuff on a smaller data set. If you have no quantitative data, but know domain experts, you can build a custom model based on their advice, with fewer parameters that need to be fit to the data you do have. But if you have so little data that you can't train a neural network, you can't be getting new data very frequently. It might be cheaper to just pay a human to look at it. If you don't even have enough data for humans to work with, fancy machine learning isn't going to help you.
- bunderbunder 8y agoLogistic regression isn't sexy, but it can still achieve near state-of-the-art results, is reasonably resistant to bias^H^H^H^H variance, and generates parameters that you can easily explain to someone with no background in math. There's a lot of value in all that. Especially if your deliverable is something that a business is going to use, and not just a Kaggle entry.
- polkapolka 8y agoYeah, it is a good first benchmark. But view interpretability as separate from accuracy. You can explain black box algorithms just fine these days. Logistic regression is high bias low variance. If you were talking about fairness bias, then resistance to bias comes from logreg being too dumb to recognize complex non-linear patterns. Not necessarily a pro.
- bunderbunder 8y agoSorry, I misspoke - will edit. Was talking about resistance to overfitting. Which largely comes from logistic regression's assumption of a linear decision boundary. It's true surprisingly often in classification tasks, and, when it's not, you can usually model it just fine with interaction variables. With an ANN, your easiest defense against overfitting is to have great big heaping piles of training data. That's something that's hard to come by in many interesting situations.
- polkapolka 8y agoAgreed. Logistic regression with poly kernel or good engineering interactions can equal or beat more complex models for a fraction of the budget. All the more power to you if a solid simple logreg model (or even no ML at all) is your first deliverable.
- tomrod 8y agoWould you mind talking to how black box interpretability is becoming well known? I've seen Shapley values used for feature interpretation, but not sure what else is being done.
- abhgh 8y agoHa! Surprisingly this is not the first time I have seen someone describe this to be the boundary for logistic regression. (Btw agree with your other comment.)
- utopcell 8y agoExactly. I stopped reading as soon as I saw that image.
- Topolomancer 8y agoI agree that the image is wrong, but I find your suggestion a little bit too plain: _what_ kind of neural networks? How do I choose the training strategy, the learning rate, the architecture, etc. This opens up a can of worms that can be overwhelming for beginners. The list is not too bad actually, but the phrasing can be improved. For example, Naive Bayes should rather be introduced before LDA, as this will make LDA much more understandable. Also, LVQ seems a little bit odd---I would rather discuss better strategies for neighbourhood enumeration (approximate kd-trees or something like that).