3 ms·
Also under-rated feature of 2010-era search was Matt Cutts, author of the article. He was an outlier at Google in that he did real community engagement as well
by choppaface 2y ago
Also under-rated feature of 2010-era search was Matt Cutts, author of the article. He was an outlier at Google in that he did real community engagement as well as anti-spam, which is a huge contrast to today’s Google and how the internet has reacted to present-day SEO.
While the Matt Cuts era search tech is interesting, it’s crucial to keep in mind that the dataset was very different then too as a result of Matt Cutts’ own attitude towards spam and SEO.
Back in 2010 LDA was big and Google had used probabilistic networks e.g. Rephil / large noisy-OR networks as models
https://uh.edu/nsm/computer-science/events/seminars/2016/1104-howes.php https://uh.edu/nsm/computer-science/events/seminars/2016/110...
Would the same things work today given how SEO spam and Google ads work? The same models are probably useful but it’s the noise and the long tail of the data that makes the problem hard.
- bpiche 2y agoI was a fan of LDA but would not agree that it is 'probably useful' today. It's an unsupervised clustering algorithm based on Gibbs sampling. Like k-means, it's gonna return a few buckets that will have to be reviewed by a human for data exploration. In this case instead of neatly labeled buckets, these are unlabeled distributions of distributions (lists of single word tokens). If you do some kind of multiword tokenization preprocessing, it'll return a few lists of words and multiword tokens for each document. How is this useful to an end user? Even internally, they're not useful embeddings/vectorizations. Would love to hear some contrary opinions
- choppaface 2y agoIn many applications like especially Google's display ad targeting market, the "accuracy" of the clusters isn't so import as the lift in key metrics (e.g. click rates or revenue) and the overall efficiency of the method. Indeed the clustering algo might get things "dead wrong" but somehow surface something that causes clicks and revenue to increase. LDA offered much improvement over e.g. TF-IDF models, just as t-SNE improved on LDA, and now LLM embeddings are on average better and potentially cheap to compute. LDA could be useful if your success metric is perplexity; k-means is useful if vector distance is very meaningful for your problem. Also well-studied algorithms are generally useful for initial studies in a new, unknown dataset. As always with ML, the dataset and setting are just as important as the model and algorithm.
- bpiche 2y agoThank you for the well considered response