6 ms·
As a single batch job that won't be repeated this doesn't sound like a good candidate for ML. ML is more suited to on-going processes. Why would you use server
by asfdsfggtfd 8y ago
As a single batch job that won't be repeated this doesn't sound like a good candidate for ML. ML is more suited to on-going processes.
Why would you use server instances in Azure to do ML? Something like Google CloudML (I'm sure that the other major cloud providers do managed Tensorflow as well I've just never tried it on their platforms) would be a better fit to a project with only two technical staff. Your two staff probably spent a combined total of one-person-month working on infrastructure.
Your issue with small data is very real. People need to stop trying to do ML on small datasets. The results will be sub-optimal.
- a008t 8y agoWhen you say ML, you must mean deep learning applied to unstructured data (vision, audio). ML in general can absolutely be used with small datasets. ML is all about finding the right model complexity to fit to the data to maximize out-of-sample performance. If your dataset is small, all that means is that your model will have to be more crude. A simple cross-validated regularized linear regression or a shallow decision tree are ML models too, and you can usefully apply them to a dataset of just 100 samples.
- heavenlyblue 8y agoOr instead of hiring an ML employee, building a model for 100 samples I might as well apply my own business intuition.
- hikarudo 8y agoIf the number of dimensions is high, building a model by hand could be extremely difficult, and using ML makes sense.
- asfdsfggtfd 8y agoIf you have 100 examples and a high number of dimensions you will end up with a very over-fitted model.
- a008t 8y agoNot necessarily. Cross-validation can give you a valid estimate of out-of-sample performance even if you have more dimensions than samples, and even if some of the features you have are (sporadically) perfectly correlated with the target variable. See https://stats.stackexchange.com/questions/295626/does-cross-validation-really-work https://stats.stackexchange.com/questions/295626/does-cross-...
- heavenlyblue 8y agoThe terms you're using - they come from Machine Learning, which comes from the same division of mathematics as ordinary statistics. And that's how ML is educated at the universities anyway. So if you've got an analyst sitting on that issue - then your problem is solved anyway. So again, why hire somebody with a trendy specialty who is <probably> full of BS, when you can hire someone with some respect for classics?
- a008t 8y agoI never said anything about who to hire.
- a008t 8y agoIt depends on the application. You might want to use the data to verify your intuition, which may not be consistent with the data. Or your intuition may have been based on the very same dataset but overfit to it. I do think that people that think that ML = big data are mistaken. ML is about making the most of your data, however much of it you have.
- heavenlyblue 8y agoThen you're mistaking a trendy term "ML" for "hiring an analyst".
- a008t 8y agoML is ill-defined. But take most textbooks on ML, and you will find regularized linear regression, decision trees, cross-validation and bootstrapping all in there. In my view, the main difference between ML and plain statistics is that with the latter, you come up with the appropriate model apriori, and then make sure the data satisfies the assumptions so that you can draw conclusions from the in-sample fit of the model. You control for overfitting by choosing the simplest model that is reasonable - often univariate linear regression. Whereas with ML, you let the data dictate how complex a model you should use. You choose the appropriate model complexity using techniques like cross-validation, and verify the effectiveness of your model empirically. ML is often used interchangeably with ANNs which I think is a mistake. Take structured data problems on Kaggle and you would very rarely see ANNs as a major predictor in the winning models.
- asfdsfggtfd 8y agoIf your model doesn't beat a reasonable benchmark created using some business intuition then this is a bad idea.