5 ms·
This article strongly resonates with me (thanks OP!). Models trained on huge datasets are truly very impressive, and it's easy to jump on the hype train of "BIG
by andersource 4y ago
This article strongly resonates with me (thanks OP!). Models trained on huge datasets are truly very impressive, and it's easy to jump on the hype train of "BIG MODEL GO VROOM" and overlook the cool / useful things one can do with a little data and solid domain knowledge.
Transfer learning can be immensely useful when applicable, but many times it's not (e.g. imagine a medical domain where you track a patient through some process, recording procedure choices, measurements and outcomes, etc., where it can be difficult to find relevant data elsewhere).
Some approaches I've found useful:
* Get to know the domain really well
* It's not a lot of data - that's a potential for rich interactive visualizations that allow you to get to know the data quite well, and grok how it relates to the domain knowledge
* Following the advice that ML models in production could/should start with simple heuristics, view your model more as an augmented heuristic than a powerful model to solve everything - that means also figuring out how to catch and handle cases where it's wrong (which is something one ought to do anyway)
* Invest in tailoring priors suitable to the problem, based on your domain knowledge and understanding of the data. This can range from writing your own loss function to training an ad-hoc type of model, not based on DL, e.g. using metaheuristics (genetic algorithms, simulated annealing etc.). The advantage of small data is that evaluation on it can be relatively fast and ad-hoc models using nonstandard techniques can be realistically optimized (sometimes, depending on context of course).
- nri 4y agoGlad you like the post. I strongly agree with all of your points, especially custom loss functions can be a great tool. If the problem you are trying to solve has some grounding in e.g. physics you can even go a step further and let the model itself mirror the physical equations. It's like you say, of course these big models are really cool, but I feel like most of the popular machine learning online courses are too narrowly focused on them and many people discard useful techniques if they are not popular in kaggle competitions.
- andersource 4y agoI fully agree with your perspective, and I think there's a lot of cargo-culting in that area that explains your observations. Sure, if you're a FAANG collecting massive amounts of data comes almost for free, and it makes sense to find a way to properly utilize that data. But for many startups collecting that amount of data doesn't make sense (either because they don't have a lot of users yet or because their domain isn't high-freq digital activity or both), and in those cases it can be more worthwhile figuring out what to do with the little data they have than shoe-horning their way into Big Data (TM). There's much more money to be had convincing startups to use big data infrastructure and train huge nets on GPU clusters etc. than convincing them to iterate on small ad-hoc models developed in-house. Not saying that big data or models are necessarily wrong for small startups, just that they don't have to be the default.
- Pamar 4y agoOne question (please take in account that I am no more than a dabbler in this field) - you conclude your piece with ... the nature of the problem slowly changes over time and prediction quality deteriorates. Isn't this a problem anyway, even with models based on much larger datasets?
- nri 4y agoYou are right, data/concept drift affects both approaches. What I meant to say was that retraining your model might not fix that if you baked strong assumptions into it. I edited the blog post to make it clearer.
- crabbygrabby 4y agoWith small datasets you need to know how the concept drifts well enough to model it or model the failure modes. That's half the game in my opinion.
- workingon 4y agoI’ve found constraints in the loss functions are key to finding the correct solution space. With small amount of training data and SGD you can get a lot of mathematically “correct” answers, but a well informed constraint based on your problem space can eliminate 99.9% of the mathematically correct but practically incorrect answers.
- andersource 4y agoMakes sense, and designing proper constraints is an art on its own. Out of curiosity, did you apply the constraint as a term in the loss or did you use constrained optimization?
- workingon 4y agoI’ve only ever played with constraints in the loss at this point. Would love more time to also deal with constraining the optimization. Do you have any examples that come to mind that worked well when doing this? Thank you for your suggestion!
- andersource 4y agoThanks for your reply! I've only seen (and used) this for linear / convex models and constraints, so that's actually linear / convex programming (highly recommend the cvxpy library [0]). I was curious if you've integrated that into DL models. Here's a fun example I did for a uni project: say you have a small recipe dataset, where for each recipe you have the ingredient list ("3 tbsp sugar", ".5 kg flour" etc.) and macronutrients (carbs, proteins, fat), and you want to learn to predict the nutritional content given a list of ingredients. With a bit of text manipulation you can split each ingredient text to "amount", "measurement unit" and "ingredient type". Then you can decompose the nutritional values into nutritional density of ingredients and the conversion of measurement unit -> actual mass for each ingredient type. Then you can introduce constraints in the form of known conversions between units, e.g. 1 cup = 16 x tbsp, known nutritional densities of some simple ingredients, and known unit conversions for specific ingredients. It worked better than simple regression (that didn't take into account unit conversions) and a simple MLP, though not sure how it would compare if you actually tried to finetune a language model to the task. The overall predictions weren't too accurate, but for common ingredients it actually gave very close nutritional density values (without being constrained), which was cool. Edit: another constraint I initially forgot about is that each gram of an ingredient cannot contain more than a gram of macronutrients (it can contain less because water), and of course nonnegative nutritional density. [0] https://www.cvxpy.org/ https://www.cvxpy.org/