5 ms·
Thanks for your reply! I've only seen (and used) this for linear / convex models and constraints, so that's actually linear / convex programming (highly recomm
by andersource 4y ago
Thanks for your reply!
I've only seen (and used) this for linear / convex models and constraints, so that's actually linear / convex programming (highly recommend the cvxpy library [0]). I was curious if you've integrated that into DL models.
Here's a fun example I did for a uni project: say you have a small recipe dataset, where for each recipe you have the ingredient list ("3 tbsp sugar", ".5 kg flour" etc.) and macronutrients (carbs, proteins, fat), and you want to learn to predict the nutritional content given a list of ingredients.
With a bit of text manipulation you can split each ingredient text to "amount", "measurement unit" and "ingredient type". Then you can decompose the nutritional values into nutritional density of ingredients and the conversion of measurement unit -> actual mass for each ingredient type. Then you can introduce constraints in the form of known conversions between units, e.g. 1 cup = 16 x tbsp, known nutritional densities of some simple ingredients, and known unit conversions for specific ingredients.
It worked better than simple regression (that didn't take into account unit conversions) and a simple MLP, though not sure how it would compare if you actually tried to finetune a language model to the task. The overall predictions weren't too accurate, but for common ingredients it actually gave very close nutritional density values (without being constrained), which was cool.
Edit: another constraint I initially forgot about is that each gram of an ingredient cannot contain more than a gram of macronutrients (it can contain less because water), and of course nonnegative nutritional density.
[0] https://www.cvxpy.org/ https://www.cvxpy.org/
- workingon 4y agoThank you too for this information and for the resource, I'm looking at it now it seems very interesting. The breakdown you give above seems to me to be more akin to something DL types tend to call 'feature engineering'. I also have a fun example, in this case it would be identifying land cover from satellite imagery. You can obviously just feed the raw reflectance values (RGB etc.) into a DL model to create a semantic segmentation of classes. However, it's been well established in literature at this point that that is not the most effective way to create a classification. This is similar to my previous comment, where there are lots of solutions that can be found through SGD based on these raw values. There's a lot of traditional satellite imagery analysis algorithms that are based on very simple 'band-ratios', i.e. NDVI (normalized difference vegetation index) is calculated by Near Infared - Red / Near Infared + Red. This index will visually highlight areas of vegetation that was extremely useful in human sight-based analysis to identify vegetative areas. Now, you'd expect a deep learning model that takes in all the bands to have this information already, it has the NIR band and the Red band. However, explicitly doing the NDVI calculation and using it as an input feature leads to increased accuracies for classification. The exact reasons for this are unknown, but I think you touched on some of this above. With machine learning, and DL even moreso, sometimes it's necessary to hand-hold the optimization to optimize for exactly what you want. It helps 'explainability' and oftentimes helps accuracy, at the cost of some preprocessing.
- andersource 4y agoThat's cool, I didn't know that, thanks!