4 ms·
When you said that you work in groups, the groups are discovered with a clustering technique and then you apply the learning to rank algos ?
by chudi 8y ago
When you said that you work in groups, the groups are discovered with a clustering technique and then you apply the learning to rank algos ?
- softwaredoug 8y agoIn most learning to rank data sets, the only notion of a "query" is a numerical identifier that groups a set of features and a grade. Your goal is to optimize the ordering in that group. That group can really be a user, a query pattern, a location, or a bazillion other groupings likely to share similar rankings. For example, a common training data storage format looks like grade(0-4) queryId list of feature values... Such as: 4 qid:1 1:0.5 2:0.24 # a doc with these features is 'good' for this query 0 qid:1 1:-0.5 2:0.24 # a bad doc for query id 1 is described by these features Features in an LTR context are some relationship between the query and doc (such as scores computed from Solr/Elasticsearch queries) Of course turning clicks/session info into grades isn't easy. One way to go about it is with click models (https://pdfs.semanticscholar.org/0b19/b37da5e438e6355418c726469f6a00473dc3.pdf https://pdfs.semanticscholar.org/0b19/b37da5e438e6355418c726...). Another way is to equate conversions in a session with one of the grades (a purchase is a 4). Needless to say, such usage data is sparse for long tail queries (queries that rarely happen - 'fuscha shoes'). Luckily it's not sparse for head queries (common queries - 'shoes'). So those head queries can provide some benefit, helping to generalize to the rare long tail queries. You'll see a lot of patterns that are modifiers on head queries. Like 'shoes' is a head query. But 'blue shoes' and 'purple shoes' and other 'color shoes' are tail queries. You can do a lot with that insight - Supervised grouping of queries into one 'query id' by watching how users refine queries. For example users type "shoes" then in the same session add one color adjective "red shoes". Using that to group '<color> shoes' queries - Unsupervised clustering of similar sessions or searches where users seem to expect ranking to perform similarly based on user behavior The downside/tradeoff is you lose a lot of nuance that you're getting with per-query training data. But it's a technique people use. It also substitutes one ML problem for another, which may or may not be harder than the original one
- kqr 8y agoIt's more of a hierarchy than groups, actually. If a user indicates that a pair of maroon walking shoes were relevant to their query for "red sneakers", then we have learned that "red sneakers" is associated, in order of decreasing strength, with e.g. - maroon walking shoes - walking shoes - casual shoes - footwear - clothes And obviously the same thing can be applied to generalize over the query. These hierarchies are constructed statically and dynamically with unsupervised learning, and then associations from query to groups happens dynamically.
- softwaredoug 8y agoVery cool, how exactly do you generate the hierarchies? Is there an existing site taxonomy or categorization you’re using? Or associating query strings with the docs clicked and using refinements to see the hierarchy? Or maybe LtR training data per site category?
- kqr 8y agoThey are constructed from a probabilistic similarity measure defined in terms of the metadata available for items, where their path tends to be weighed fairly heavily. Does that answer make sense?
- softwaredoug 8y agoCool stuff, appreciate the answer. Makes more sense. BTW feel free to join us on relevance slack community, a lot of us we’re curious what you were doing http://o19s.com/slack http://o19s.com/slack