3 ms·
Makes sense, and designing proper constraints is an art on its own. Out of curiosity, did you apply the constraint as a term in the loss or did you use constra
by andersource 4y ago
Makes sense, and designing proper constraints is an art on its own.
Out of curiosity, did you apply the constraint as a term in the loss or did you use constrained optimization?
- workingon 4y agoI’ve only ever played with constraints in the loss at this point. Would love more time to also deal with constraining the optimization. Do you have any examples that come to mind that worked well when doing this? Thank you for your suggestion!
- andersource 4y agoThanks for your reply! I've only seen (and used) this for linear / convex models and constraints, so that's actually linear / convex programming (highly recommend the cvxpy library [0]). I was curious if you've integrated that into DL models. Here's a fun example I did for a uni project: say you have a small recipe dataset, where for each recipe you have the ingredient list ("3 tbsp sugar", ".5 kg flour" etc.) and macronutrients (carbs, proteins, fat), and you want to learn to predict the nutritional content given a list of ingredients. With a bit of text manipulation you can split each ingredient text to "amount", "measurement unit" and "ingredient type". Then you can decompose the nutritional values into nutritional density of ingredients and the conversion of measurement unit -> actual mass for each ingredient type. Then you can introduce constraints in the form of known conversions between units, e.g. 1 cup = 16 x tbsp, known nutritional densities of some simple ingredients, and known unit conversions for specific ingredients. It worked better than simple regression (that didn't take into account unit conversions) and a simple MLP, though not sure how it would compare if you actually tried to finetune a language model to the task. The overall predictions weren't too accurate, but for common ingredients it actually gave very close nutritional density values (without being constrained), which was cool. Edit: another constraint I initially forgot about is that each gram of an ingredient cannot contain more than a gram of macronutrients (it can contain less because water), and of course nonnegative nutritional density. [0] https://www.cvxpy.org/ https://www.cvxpy.org/
- workingon 4y agoThank you too for this information and for the resource, I'm looking at it now it seems very interesting. The breakdown you give above seems to me to be more akin to something DL types tend to call 'feature engineering'. I also have a fun example, in this case it would be identifying land cover from satellite imagery. You can obviously just feed the raw reflectance values (RGB etc.) into a DL model to create a semantic segmentation of classes. However, it's been well established in literature at this point that that is not the most effective way to create a classification. This is similar to my previous comment, where there are lots of solutions that can be found through SGD based on these raw values. There's a lot of traditional satellite imagery analysis algorithms that are based on very simple 'band-ratios', i.e. NDVI (normalized difference vegetation index) is calculated by Near Infared - Red / Near Infared + Red. This index will visually highlight areas of vegetation that was extremely useful in human sight-based analysis to identify vegetative areas. Now, you'd expect a deep learning model that takes in all the bands to have this information already, it has the NIR band and the Red band. However, explicitly doing the NDVI calculation and using it as an input feature leads to increased accuracies for classification. The exact reasons for this are unknown, but I think you touched on some of this above. With machine learning, and DL even moreso, sometimes it's necessary to hand-hold the optimization to optimize for exactly what you want. It helps 'explainability' and oftentimes helps accuracy, at the cost of some preprocessing.
- andersource 4y agoThat's cool, I didn't know that, thanks!