4 ms·
They talk about using query logs to optimize their search results: >Queries performed by people, if associated to a web page, serve as even cleaner summaries t
by mgreg 7y ago
They talk about using query logs to optimize their search results:
>Queries performed by people, if associated to a web page, serve as even cleaner summaries than anchor text. This is because all the logic put in place by the search engine, who resolved the query with a list of web pages, and all human understanding and experience that led one to select the best page from the offered result list end up embedded in the association <query, url>.
This would seem to present a "rich get richer" problem where the oldest links that have the largest click-through tend to float to the top making it difficult for a new result that may be "better" to appear high in the search results. Anyone know how search engines tackle this problem?
- stanislavb 7y agoYes, that's an interesting problem to solve. Maybe some A/B testing and showing new results to a percentage of all searches.
- mc987 7y agoSurfacing new content in search engines is a very challenging problem. I am guessing they use a combination of social signals (twitter, facebook) popularity and domain popularity amongst other signals.
- alexis_fr 7y agoGPS tracking must also help Google a lot to determine mortar shop popularity.
- pheug 7y agoGoogle pretty much knows (or can accurately estimate) exactly when a new document appears on the web and how many people are visiting it, they don't even need to rely on second hand social signals for this. They control the web's dominant crawler (Googlebot), browser (Chrome, which sends everything you type in the address bar to them by default), ads (Adsense) and tracking (Google Analytics) platforms.
- netankit 7y ago[Disclaimer: I work at Cliqz] You are on point. Recency is a challenging problem in multiple ways for search engines. Not just limited to discovering new content, but also how does one index it? How does one balance out when you have for the same query "very new", "new", "slightly old" and "really old" results during ranking. This involves both news as well as new webpages surfacing on the web. On top of this, we have to remember that this is a fully autonomous real time system which requires solving some of the most difficult engineering challenges at scale and at the same time being mindful of the latency and quality constraints. At the end of the day, it's all about the final user experience that we ship. We are very much mindful of the same. We will be publishing more details about Cliqz search, on our blog https://0x65.dev/ https://0x65.dev/ in the coming days, so stay tuned.
- bryanrasmussen 7y agoIf I had a big enough user base I would give some users non optimized queries and check if the choices still matched up to previous choices, if they did I would increase rank on previous choices, if not I would start to down prioritize and increase the choices that were chosen. An ongoing test of is result X for search term y still best?
- solso 7y ago[Disclaimer, I work at Cliqz] Your point is spot on. Old pages tend to have more association to seen queries, which does not play in favor for new pages. That said, however, there are a couple of things to consider: 1) seen queries is not the only way to create queries, we are pretty good creating synthetic queries based on the content, descriptions, etc. This queries are more noisy that the seen queries of course, but good enough. And 2)novelty, freshness and popularity are very important features on the ranking. Feel free to try out any new topic you might think of on https://beta.cliqz.com https://beta.cliqz.com, you will see that is not only "stale" content.
- mgreg 7y agoThank you for the additional detail and this certainly appears to be a challenging problem. It is still a little fuzzy to me. What is a "synthetic query"? Is this basically generating queries that would match the content (i.e. essentially reversing the process)? Novelty, freshness are interesting but can lead back to the noise problem mentioned in the blog. If many pages are created that may match the query (e.g. "best new movies") many young pages will match this. Popularity would be useful but difficult to establish and then there's the clickbait and other gaming problems.
- kick 7y agoWhy does Cliqz use an analytics domain with a typo in it to get around user tracker-blockers? That's incredibly scummy, given how much Cliqz has been shouting about privacy. https://anolysis.privacy.cliqz.com/ https://anolysis.privacy.cliqz.com/
- JadoJodo 7y agoSeems like an excellent opportunity to apply Hanlon's Razor. "Never attribute to malice that which can be adequately explained by stupidity."
- kick 7y agoThey've been doing it for two years at minimum: https://news.ycombinator.com/item?id=15423936 https://news.ycombinator.com/item?id=15423936 Two years and three domains with the typo? It's malice.
- toast0 7y agoGoing strictly by volume of clicks (or volume of clicks for a keyword) is not going to be fair or particularly useful, you need to compare clicks received to clicks expected. If it's shown in position one, it should be expected to get more clicks than position ten. If something gets more than expected, move it up.
- MayeulC 7y agoWouldn't a multi-armed bandit help alleviate this issue? (Basically, randomly display a few other links, and use bayesian stats to figure out if the new links are more optimal). That, or any kind of exploration/optimization algorithm, to be honest.