7 ms·
Some thoughts on machine learning with small data
- andersource 4y agoThis article strongly resonates with me (thanks OP!). Models trained on huge datasets are truly very impressive, and it's easy to jump on the hype train of "BIG MODEL GO VROOM" and overlook the cool / useful things one can do with a little data and solid domain knowledge. Transfer learning can be immensely useful when applicable, but many times it's not (e.g. imagine a medical domain where you track a patient through some process, recording procedure choices, measurements and outcomes, etc., where it can be difficult to find relevant data elsewhere). Some approaches I've found useful: * Get to know the domain really well * It's not a lot of data - that's a potential for rich interactive visualizations that allow you to get to know the data quite well, and grok how it relates to the domain knowledge * Following the advice that ML models in production could/should start with simple heuristics, view your model more as an augmented heuristic than a powerful model to solve everything - that means also figuring out how to catch and handle cases where it's wrong (which is something one ought to do anyway) * Invest in tailoring priors suitable to the problem, based on your domain knowledge and understanding of the data. This can range from writing your own loss function to training an ad-hoc type of model, not based on DL, e.g. using metaheuristics (genetic algorithms, simulated annealing etc.). The advantage of small data is that evaluation on it can be relatively fast and ad-hoc models using nonstandard techniques can be realistically optimized (sometimes, depending on context of course).
- nri 4y agoGlad you like the post. I strongly agree with all of your points, especially custom loss functions can be a great tool. If the problem you are trying to solve has some grounding in e.g. physics you can even go a step further and let the model itself mirror the physical equations. It's like you say, of course these big models are really cool, but I feel like most of the popular machine learning online courses are too narrowly focused on them and many people discard useful techniques if they are not popular in kaggle competitions.
- andersource 4y agoI fully agree with your perspective, and I think there's a lot of cargo-culting in that area that explains your observations. Sure, if you're a FAANG collecting massive amounts of data comes almost for free, and it makes sense to find a way to properly utilize that data. But for many startups collecting that amount of data doesn't make sense (either because they don't have a lot of users yet or because their domain isn't high-freq digital activity or both), and in those cases it can be more worthwhile figuring out what to do with the little data they have than shoe-horning their way into Big Data (TM). There's much more money to be had convincing startups to use big data infrastructure and train huge nets on GPU clusters etc. than convincing them to iterate on small ad-hoc models developed in-house. Not saying that big data or models are necessarily wrong for small startups, just that they don't have to be the default.
- Pamar 4y agoOne question (please take in account that I am no more than a dabbler in this field) - you conclude your piece with ... the nature of the problem slowly changes over time and prediction quality deteriorates. Isn't this a problem anyway, even with models based on much larger datasets?
- nri 4y agoYou are right, data/concept drift affects both approaches. What I meant to say was that retraining your model might not fix that if you baked strong assumptions into it. I edited the blog post to make it clearer.
- crabbygrabby 4y agoWith small datasets you need to know how the concept drifts well enough to model it or model the failure modes. That's half the game in my opinion.
- workingon 4y agoI’ve found constraints in the loss functions are key to finding the correct solution space. With small amount of training data and SGD you can get a lot of mathematically “correct” answers, but a well informed constraint based on your problem space can eliminate 99.9% of the mathematically correct but practically incorrect answers.
- andersource 4y agoMakes sense, and designing proper constraints is an art on its own. Out of curiosity, did you apply the constraint as a term in the loss or did you use constrained optimization?
- workingon 4y agoI’ve only ever played with constraints in the loss at this point. Would love more time to also deal with constraining the optimization. Do you have any examples that come to mind that worked well when doing this? Thank you for your suggestion!
- andersource 4y agoThanks for your reply! I've only seen (and used) this for linear / convex models and constraints, so that's actually linear / convex programming (highly recommend the cvxpy library [0]). I was curious if you've integrated that into DL models. Here's a fun example I did for a uni project: say you have a small recipe dataset, where for each recipe you have the ingredient list ("3 tbsp sugar", ".5 kg flour" etc.) and macronutrients (carbs, proteins, fat), and you want to learn to predict the nutritional content given a list of ingredients. With a bit of text manipulation you can split each ingredient text to "amount", "measurement unit" and "ingredient type". Then you can decompose the nutritional values into nutritional density of ingredients and the conversion of measurement unit -> actual mass for each ingredient type. Then you can introduce constraints in the form of known conversions between units, e.g. 1 cup = 16 x tbsp, known nutritional densities of some simple ingredients, and known unit conversions for specific ingredients. It worked better than simple regression (that didn't take into account unit conversions) and a simple MLP, though not sure how it would compare if you actually tried to finetune a language model to the task. The overall predictions weren't too accurate, but for common ingredients it actually gave very close nutritional density values (without being constrained), which was cool. Edit: another constraint I initially forgot about is that each gram of an ingredient cannot contain more than a gram of macronutrients (it can contain less because water), and of course nonnegative nutritional density. [0] https://www.cvxpy.org/ https://www.cvxpy.org/
- smartmic 4y agoIn my experience, overfitting also helps in case of small data problems. After all, there is a reason that you do only have few training data samples and the chance that any new prediction sample is close to the existing training ones is higher. In other words, for small data its more about specialization than generalization. Of course this has to be evaluated case specific.
- nri 4y agoInteresting, that advice is exactly the opposite of the common wisdom. You mean overfitting as in aggressively maximizing your crossvalidation scores? How do you decide for which problems that is a good approach and for which a more conservative approach is better?
- smartmic 4y agoWell, I mean not actively maximizing your CV score, just accepting normally insufficient CV scores. For example, in a tree based regression with a very small number of training samples, the leaves will almost resemble some of the training samples. A good generalization might not be possible at all. But this is not too bad if you know that your prediction samples have high similarity with at least one of the training samples. In the end, this is an edge case for Machine Learning but goes more into the direction of Expert Systems.
- isoprophlex 4y agoIs this a complicated way of saying "just use k-nearest neighbors"?
- MrMan 4y agoI think its implying that ML often works because the manifold assumption is partially correct. so samples will lurk in the same neighborhood.
- crabbygrabby 4y ago
- rmnclmnt 4y agoThanks for the article, it resonates strongly! I teach Data / AI courses regularly and this topic is a recurring one with students. We the instructors like to make students start without even ML in the first place for them to construct a baseline model, usually using simple heuristics (e.g. conditions, regexes, etc.). At first, students don't like it and tend to skip it because it is not that impressive, but in the end they learn so much about the data and the problem at hand! Then they are allowed to try simple ML models to see if that's a performance improvement and so on with more complex models. In some cases, improving baseline models with complex ML/DL models is really hard (e.g. time-series). Another benefit of simple models is explainability, for which the industry demand is growing everyday.
- figurative 4y agoJust out of curiosity where do you see the demand? I've also heard from others that there's been a shift lately towards simpler models. Combined with domain knowledge linear/logistic regression can be really impressive!
- dbs 4y ago> Combined with domain knowledge linear/logistic regression can be really impressive! Can you elaborate?
- rmnclmnt 4y agoI can agree this the comment. Linear models combined with advanced feature engineering gathered from domain knowledge can achieve great results in a white-box fashion! A nice keynote by Vincent Warmerdam [1] talks about tips and tricking for advanced feature engineering combined with linear models. [1] https://www.youtube.com/watch?v=68ABAU_V8qI https://www.youtube.com/watch?v=68ABAU_V8qI
- beckingz 4y agoA significant portion of ML workloads involve predicting or classifying something. Linear/logistic regression of the right variables/features typically gets a significant portion of the data's ability to predict /classify correctly, while being significantly easier to build, train, deploy, and understand. Heck, in a large number of domains, simple ratios -- debt to income ratio in finance for example -- will dominate the feature weight for many models and can be used on their own as a pretty good heuristic.
- evrydayhustling 4y agoA powerful concept for thinking about small data is bias/variance tradeoff. The ELI5 is: data is how a learner selects confidently between one model and another. So if you have little data, you must choose between high bias (fewer models to choose from) or low confidence (can't be sure you picked the best one). Bias in a lot of contexts is Bad, and large data techniques like DNN are basically about how you can have biases that are so vague the learner has plenty of room to surprise you. But when you add bias that truly aligns with the structure your learner is going to be exposed to, you enable it to reach confidence sooner. Many important techniques for small data are about allowing you to express specific biases. For example: - PGMs and Probabilistic Programming are about giving you a specific, interpretable structure for how your data are related to each other. - Picking informed bayesian priors - Data augmentation lets you add bias by expressing what kinds of variation don't matter (I.e. permitting and fuzzing your dataset) - feature engineering is about selectively adding and removing biases in terms of what kinds of data transformations are informative With all that said, the large data ML community has produced an incredible tool for practical work with small data: transfer learning. If your dataset is related to a larger one (e.g. by including natural language), you can borrow informed biases from models that were trained on a much larger corpus.
- Hendrikto 4y ago> Data augmentation lets you add bias by expressing what kinds of variation don't matter (I.e. permitting and fuzzing your dataset) What does ”permitting your dataset“ mean? Is it a typo? Did you mean ”permuting“?
- evrydayhustling 4y agoYes, that was supposed to be "permuting"!
- xksteven 4y agoI'd be careful of over applying the "bias-variance tradeoff." How to define the variance of a model is not a simple task. I wouldn't say it is immediately obvious how bias-variance relates to small data scenarios. How much data is considered small? What is the complexity of the dataset itself? Even in Machine Learning it is possible to learn from small datasets without transfer learning. See meta-learning for instance.
- cyocum 4y agoI occasionally work with data in the Humanities. The data here is often very, very small. I talk to other Humanities researchers and I often find that they really want to get on the ML bandwagon but they do not realize the sheer amount of data that they need to make ML as practiced today work. I have not looked into small dataset techniques in a long time (I have a day job so I do not get much chance to do this often) but I hope that one day we can find a technique that will work. One side note, when I speak to other Humanities researchers about this, I always tell them that I have yet to find a technique that will give them novel insights. These techniques almost always tell the researchers things that they already know. I usually follow this up with a note that even formalizing Humanities knowledge in statistical or other computational terms is highly valuable and worth doing. Maybe someone else can take that formalism and build on top of it something truly new.
- LudwigNagasena 4y agoWhat is even small dataset machine learning and how is it different from statistical forecasting?
- mellavora 4y ago> how is it different from statistical forecasting In statistical learning/forecasting, the researcher typically specifies the statistical model. In machine learning, the statistical model is approximated by the algorithm. Since a ML model needs to learn both the model form and the model parameters, it takes more data and also it does not allow for understanding (since it does not output the form of the model it learned).
- mellavora 4y agoYou can fit a linear regression with just a few points, technically if you have one more data point than regressors, it works. Because you've assumed a linear relationship with normally distributed errors. And you can interpret the output, because the values of the regression coefficients tell you something. "having a high score on X doubles the odds of outcome Y", for example. Also, because you've assumed a structure to the data, you can more easily test if the data has deviated from that structure. This can be data drift or single outliers-- for example, GARCH models (a type of regression) allow the normal distribution of the error to have a varying variance, so you can detect different variance regimes. In short, they help a human understand and interpret data. From what little I know, ML is not so good at that. But it has other advantages, and you don't always need or want the understanding. If your want to i.e. detect ground cover in satellite images, then all you care about is valid outputs, not necessarily the importance of near-infrared vs red band. And ML (can) beat regression models by providing better interpolation, by better handling regions of the data space which violate the assumptions of the regression model, etc. So it is a tradeoff. Both approaches are highly performant, just at different tasks.
- zone411 4y ago1. Try to get more data 2. Augmentation 3. Heavy regularization 4. Cross-validation 5. Transfer learning/fine-tuning from related data 6. Sharpness aware optimization 7. Ensembles 8. Any domain knowledge/feature engineering
- rexreed 4y ago9. Synthetic data 10. Single-shot / low-shot learning methods 11. Reducing the scope of the problem so you don't even need that much data 12. Iterative model approaches that use submodels focused on different target problems (instead of trying to boil the ocean with a single model) A lot of these are not what you do, but how you do it, something popularized by CPMAI methodology. I would also add that I'm not sure the OP was tongue-in-cheek when saying that AGI is coming, so we don't need to worry about low data requirements. But needless to say, AGI is not coming any time soon, and if OpenAI's strategy for AGI is to be believed, it's anything but a low-data method. Unless you're counting on them to build the super-sized model that meets all needs.
- mupuff1234 4y agoAs someone who doesn't deal with ML, what would be considered a small data set? Is there a good way for a given problem to estimate how many examples would be required to produce a decent model?
- bavell 4y agoI have very limited experience with ML but as long as your data is representative of the full problem domain, then you should be good.
- rendang 4y agoA common rule of thumb is 10 observations per variable. Small data depends on the domain/context. It might mean < 30 observations or it might mean < gigabyte scale.
- borroka 4y agoI first read about 10-20 observations per feature in a statistical model (not ML), including interactions between features, in Frank Harrell's "Regression Modeling Strategies," an excellent book that many coming from the CS side of modeling would benefit greatly from studying.
- ChrisRackauckas 4y agoThere's also data in another source: domain knowledge and physical laws. This is the core tenant of scientific machine learning techniques (SciML). I have a talk that walks through how you can go from something without domain knowledge (like a neural ODE) and how as you add more prior model knowledge to the system the extrapolation accuracy improves even for small amounts of data (https://www.youtube.com/watch?v=FihLyzdjN_8 https://www.youtube.com/watch?v=FihLyzdjN_8). People have used these techniques in all kinds of places, like extrapolating black hole trajectories and building earthquake-safe buildings, off of like 20-40 data points (https://www.youtube.com/watch?v=eSeY4K4bITI https://www.youtube.com/watch?v=eSeY4K4bITI). While there is a lot of work to be done in the domain still, there's a lot of empirical evidence that this approach does indeed decrease the data requirements (but is then no longer "pure" machine learning in some sense)
- joshgev 4y agoDamn, that was a really cool talk. Nice work and thanks for sharing.
- residual 4y agoI see two links. Which one did you think was cool?
- dekhn 4y agowhy would you build an earthquake-safe building using only 20-40 data points (sorry, I skimmed the video and I dont think it answered that question). That sounds irresponsible since you could collect far more data points, and unless you could demonstrate that collecting tiny amounts of data was "safe", you're always going to collect high quality data and lots of it. The advantage of physical systems is that it's straightforward to collect data from natural systems.
- ChrisRackauckas 4y ago> The advantage of physical systems is that it's straightforward to collect data from natural systems. For a lot of physical systems it might be "straightforward" but expensive. Testing new building designs for energy efficiency is just a matter of building a new test building, why not just build 1000? Testing new car designs to see if they reach the safety qualifications is just a matter of crashing a few of your prototypes, why not crash 1000? It's just a matter of getting more telescope time. It's just a matter of getting more robots to try more materials or proteins. Yes you can do theoretically do it, it's just cost prohibitive to get Google-level datasets for many scientific and engineering problems. Of course that's not true for all domains (high-throughput sequencing for bioinformatics is the prime example of something that became cheaply data-rich for standard machine learning), but there are many domains where getting another data point is always possible but just would cost another million.
- Kalanos 4y agoI think you'd be surprised at how few samples you need in order to get most deep learning algorithms learning at low loss. All you need are features with distributions that are roughly representative of the broader population. The real problem that I see is that people don't pick a task that is granular enough. github.com/aiqc/aiqc
- crabbygrabby 4y agoMachine learning on small data, oh you mean statistics? I'm trolling a bit here, but this is what domain experts and scientists have been doing for about a century. It's a much harder area to hang out in, and pays way less unless you are a world expert consultant. Stakes are often higher. I used to do this stuff, but there's so much hype around ML it's basically impossible to survive in industry even if you do very good work. For example if you create a model catered to a system and it takes two weeks but is as correct as it can be given what's known about the system, your manager will turn their chin too it while trying to find startups to buy(yes buy the company) that claim they can solve it with deep learning. The startups run away once they learn each sample costs 10k USD or deliver a boosted decision tree model that under performs, but it got so old I changed careers.
- yazanobeidi 4y agoWould you mind expanding on “each sample costs 10k USD”?
- jeffreyrogers 4y agoMost older machine learning techniques work with "small data". Most of the literature pre deep learning is on those techniques and what problems they work well for.