3 ms·
Looking at the hn demo, I'm impressed. There are definitely relevant tags being generated. Unfortunately there also some noisy tags which clutter the results. T
by Goosey 12y ago
Looking at the hn demo, I'm impressed. There are definitely relevant tags being generated. Unfortunately there also some noisy tags which clutter the results. Taking one example, the post "DevOps? Join us in the fight against the Big Telcos" given the tags "phone tools sendhub we're news experience customers comfortable", I would say that "we're" is unarguably noise. Another example, "Questions for Donald Knuth" with tags "computer programming don i've knuth taocp algorithms i'm" I would call out "i've" and "i'm".
There are other words in both examples that I personally would not use as tags, but I can't really say they would be universally not-useful. I think a vast improvement could be made just by having a dictionary blacklist filled with things like these - from this tiny sampling contractions seem to be a big loser.
- doppenhe 12y agoAgreed. Actually we could turn up the number of times it runs inside LDA algorithm and those would fix up but it affects performance. This was just a quick and dirty example (with an expectation of high traffic). You can also seed LDA with a whitelist of words which we didn't do either - again all in the name of a quick and dirty solution to show. Glad you liked it!
- ppod 12y agoTry using tf-idf instead of raw word frequencies.
- doppenhe 12y agonot using raw word frequencies but http://en.wikipedia.org/wiki/Latent_Dirichlet_allocation http://en.wikipedia.org/wiki/Latent_Dirichlet_allocation. Didnt know about tf-idf thanks for the tip.
- hnriot 12y agoit's hard to imagine someone knowing an LDA without also knowing about TF.IDF (it's a dot product, not a hyphen)
- sprobertson 12y agoThat and the questionable use of stopwords makes it sound like they're just slapping some marketing on an out-of-the box LDA implementation (not that I blame them, it's a dense algorithm).
- GFK_of_xmaspast 12y agoThe OP is pushing "Algorithmica" whose manifesto is here: http://blog.algorithmia.com/post/75680476188/algorithm-development-is-broken http://blog.algorithmia.com/post/75680476188/algorithm-devel... and this doesn't really strike me as much of a victory for the idea that it's just the implementation of an algorithm being the sticking point in practice.
- doppenhe 12y agoWe are just showing the versatility of the platform through a real world use case. LDA is hard to implement/scale for the untrained same as many other machine learning, optimization, graph traversing,etc algorithms. What we are building is crowd-sourced and generalized API where all these algorithms can be combined and used together to really make any application smarter. The demo we show here is a version of how we used our platform to generate tags for all entries in our API by combining algorithms that existed already in Algorithmia. (modified for performance over quality due to the volume that HN would bring). Cheers.
- s0x 12y agoIt's only hard because most libriaries I've seen have so little documentation available. It's simple once you understand the library. We need people picking these libraries up, implementing them on weekends on fun projects, documenting their work and code, and publishing it for everyone to learn from.
- HNJohnC 12y agoHow about the plain old traditional 'stop words' list to solve this issue?
- r00fus 12y agoWhat about tag size (i.e., word length)? For the example of the knuth article, it'd be good if length was > 3.
- arg01 12y agoI think you'd need to whitelist some useful acronyms if you implemented that rule. USA NSA DOJ TCP POS etc.