Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
eggie5
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
10 ms
·
61.
▲
by
eggie5
7y ago
Although your post is orthogonal to what we present in this work, I think it still merits discussion. The high-level concept of a product is important for any commerce company w/ an unbounded and unstructured product catalog. In this c
62.
▲
by
eggie5
7y ago
thanks, I think we are in agreeance.
63.
▲
by
eggie5
7y ago
I agree: In the IR community you will see image2image search implemented by passing an image through a headless CNN and then do ANN search on the embedding space. Then that can lead to discussion about cross-model hashing which you suggeste
64.
▲
by
eggie5
7y ago
Author here. Thanks for the comment. This is a tough business to be in, but very exciting a the same time. I'm trying to effect change by making meaningful contributions to our search pipeline by applying novel techniques from informat
65.
▲
by
eggie5
7y ago
Correct, the model has a fixed vocabulary set at train-time. However, we can benefit from the fact that the distribution of queries in most search systems follows the Zipfian distribution. That means we can capture the vast majority of quer
66.
▲
by
eggie5
7y ago
Author here. Thanks, for the recommendation. I actually got inspiration for this idea from chatting w/ industry colleagues at SIGIR. Last SigIR in Paris, I spoke to some people from Apple, where they are using a technique like this for
67.
▲
by
eggie5
7y ago
Hi, I'm the author of the post. We actually just presented this work and more at PyData NYC. We share more implementation details. Here are the slides: https://www.slideshare.net/AlexEgg1/discover-yourlatentfoodg..
68.
▲
by
eggie5
7y ago
Author here: We hope to fix this w/ the techniques described in the post! The search engine will first collect a high-recall set of candidates which is then passed to a high-precision ranker. This should help you w/ your cuisine a
69.
▲
Query2vec: Search query expansion with query embeddings
(bytes.grubhub.com)
81 points
by
eggie5
7y ago
|
23 comments
70.
▲
by
eggie5
7y ago
CNNs are translation and scale invariant thanks mostly to the pooling operation. Good data augmentation (rotating images for example) would have build a model more robust to this effect.
71.
▲
by
eggie5
7y ago
There's loss of fpgas in space already. Space micro is one example
72.
▲
by
eggie5
7y ago
I have
73.
▲
by
eggie5
7y ago
no, I have never watched a course before and watched the first lecture of this last night w/ no issue.
74.
▲
by
eggie5
7y ago
instead of thinking about what it is in practice: skip-gram negative sampling, I think it's much more intuitive to think about what it is in theory: extreme multi-class classification. word2vec is a multi-class classification problem w
75.
▲
by
eggie5
7y ago
can you share the two numbers you are using for comparison? The Q4 numbers I see: Uber Eats: $165 million Grub Hub: $205 million* * https://investors.grubhub.com/investors/press-releases/press...
76.
▲
by
eggie5
8y ago
there's pipeline.ai, airbnb said they'd open source theirs this year and also, TFX suite is getting there. Platforms are becoming popular
77.
▲
by
eggie5
8y ago
he generated query-document features. Now he just needs to collect relevance labels for the documents, then he can learn a ranker a la LTR.
78.
▲
by
eggie5
8y ago
+1 partial of the loss w/ respect to the weight
79.
▲
by
eggie5
8y ago
Probably a good call. Slack trains recommender systems w/ all user data mixed together. Although w/ attempts to preserve privacy... See their paper from recsys '18 Vancouver.
80.
▲
by
eggie5
8y ago
cluster horse_js and Tom Dale's tweets in an embedding space and you can confirm your hypothesis.
81.
▲
by
eggie5
8y ago
This isn't hard to believe if you've worked w/ the Edgar system!
82.
▲
by
eggie5
8y ago
I've seen a rule of thumb that if the ratio of samples to words per sample is less than 1500 you prob don't have enough data for embeddings/cnn
83.
▲
by
eggie5
8y ago
Why not just use a 1d conv over the sequence of embeddings?
84.
▲
by
eggie5
8y ago
* factorization machines * word2vec * WARP loss for Matrix Factorization
85.
▲
by
eggie5
8y ago
Rooibos brother, This is my favorite blend I randomly found outside of the Schönefeld Airport in Berlin and have been importing the states ever since: https://shop.tgtea.com/Rooibush-Cream-Caramel-001312-100/561...
86.
▲
by
eggie5
8y ago
Plesently surprised w/ the choice of algorithms AWS provides. They skip classic MF and go straight to the good stuff: * Item-Item CF (workhorse amazon original from 2003 w/ modern enhancements) * Deep Pooling Models (à la Covingto
87.
▲
Amazon Personalize – Real-Time Personalization and Recommendation for Everyone
(aws.amazon.com)
4 points
by
eggie5
8y ago
|
1 comments
88.
▲
by
eggie5
8y ago
the pieces are: * CNN (plenty of pre-trained * Approximate Nearest Neighbors Database (Annoy, etc) * webserver to host CNN and serve UI It's popular to serve tensorflow models w/ tensorflow severing + kubernetes and that's w
89.
▲
by
eggie5
8y ago
The successive convolutional layers in popular CNN architectures learn representations from very general (edges) to very specific (dog breeds) from the input to the last conv. layer respectively. Depending on how different your new domain i
90.
▲
by
eggie5
8y ago
I gave a talk on the theory behind image to image search if anyone is interested. Image search is essentially what this backend well suited for and what the graphic on their home page uses: http://www.eggie5.com/126-semantic
More ›