7 ms·
Hi, I'm the author of the post. We actually just presented this work and more at PyData NYC. We share more implementation details. Here are the slides: https://
by eggie5 7y ago
Hi, I'm the author of the post. We actually just presented this work and more at PyData NYC. We share more implementation details. Here are the slides: https://www.slideshare.net/AlexEgg1/discover-yourlatentfoodgraphwiththis1weirdtrick-pydata-nyc-2019 https://www.slideshare.net/AlexEgg1/discover-yourlatentfoodg...
But to answer you question here: For production ANN we have Annoy integrated into our backend as a service. Annoy was an easy choice bc it checked out box for JVM support.
For training the models we have endless amounts of behavioral data, so we didn't even need to look at transfer learning. For this query2vec example, it was trained on 1 year of queries which takes 15min/epoch on an AWS p2 GPU. We do all our preprocessing (heavy normalization) in pyspark.