3 ms·
Yes, we definitely have clustering. In fact, it's more extensive than simple word distances like cosine b/c we also take into account synonyms and word relation
by haidut 17y ago
Yes, we definitely have clustering. In fact, it's more extensive than simple word distances like cosine b/c we also take into account synonyms and word relationships (set membership, etc). For instance, in our system sentences like "Tiger chases antilope down the river" is very closely "related" to the sentence "Lion is pursuing a buffalo by the lake" b/c both sentences essentially say that a large cat is pursuing a prey of bovine origin near a water source.
In terms of the "most viewed" and how we control for that - like I said we cross-track news on multiple news sites and weight the cross posted one more often. We also cross-validated the most important tags for an article by using Google Trends data. Basically we tracked topics on multiple sites and then performed some statistical analysis to see how those topics did over time based on their presence on the web (topic momentum and longevity). We also run a partial search engine in house that crawls a subset of the web so we can ensure that the numbers we get from Google/Yahoo are legit. Finally, there is linguistic theory of topic popularity and how memes propagate over time. We use some of that theory to control for the crowd effect - i.e. sometimes people pick up and spread topics that are of no real importance to the world. Example: Paris Hilton's latest escapades may be widely discussed online and appear important news but the latest report on the recession estimates and projections is of much higher "impact" to society. So we try to account/estimate some of that "impact". Combining all factors gives an article a composite score. No two article really have the same score but a lot of articles cluster close to each other in terms of their "importance" cores. We fed the articles in a machine learning algorithm that is a combination of Support Vector Machine, Neural Network, and Naive Bayes and when a new article is fetched by our crawler the model "preditc" its various scores (controversial, engaging, popular) based on the data set that it has already learned. Deception detection is much trickier and is almost entirely analysis based - i.e. no machine learning there. There is quite a bit of research on deception detection published online. Just search google for "deception detection ext:pdf" and it will come back with a lot of results.