4 ms·
This is a fantastic post. Would love to see an even more detailed walk through some code examples, and discussion of development to production of these models.
by lootsauce 7y ago
This is a fantastic post. Would love to see an even more detailed walk through some code examples, and discussion of development to production of these models. Do you have any other resources you would suggest?
- lootsauce 7y agoFor example, what ANN do you use in production, what are the reasons for using it. Another question is how much data does it take to make these models work, can they be primed effectively with little data if doing transfer learning from a more general model?
- eggie5 7y agoHi, I'm the author of the post. We actually just presented this work and more at PyData NYC. We share more implementation details. Here are the slides: https://www.slideshare.net/AlexEgg1/discover-yourlatentfoodgraphwiththis1weirdtrick-pydata-nyc-2019 https://www.slideshare.net/AlexEgg1/discover-yourlatentfoodg... But to answer you question here: For production ANN we have Annoy integrated into our backend as a service. Annoy was an easy choice bc it checked out box for JVM support. For training the models we have endless amounts of behavioral data, so we didn't even need to look at transfer learning. For this query2vec example, it was trained on 1 year of queries which takes 15min/epoch on an AWS p2 GPU. We do all our preprocessing (heavy normalization) in pyspark.