3 ms·
The title should be "Clustering HTML pages". It's not a terribly interesting application. The only thing new I got from it was a de-noising technique.
by abadon 8y ago
The title should be "Clustering HTML pages". It's not a terribly interesting application. The only thing new I got from it was a de-noising technique.
- pagnol 8y agoWhat I'd really like to see is a presentation of an algorithm that automatically recognizes and hides the first dismissive HN comment that inevitably appears. Any takers?
- rimliu 8y agoWhy would you want to have your comment hidden? On the more serious note I would not mind a but more critical and less clickbaity attitude in the tech. Not every if statement or regexp is ML and AI, not eveything requires blockchain.
- jacquesm 8y agoRecursively?
- goostavos 8y agoRecognize, hide, and automatically post to /iamverysmart.
- abadon 8y agoYou can do it yourself, if you're so inclined. Out-of-the-box sentiment analysis is ~90% accurate. Feel free to provide training data for it.
- inputcoffee 8y agoNot sure why you're getting down-voted. You're right, the author didn't close the loop. Typically we would expect to see an insight after you apply the ML technique. So they extracted features, clustered the pages and found... what? I am sure they learned something but it might be proprietary.
- abadon 8y agoPeople just like to be cunts. I've been doing ML for 15 years and I'm not apologetic about calling shit shit. The article was 101-level stuff and there was no application to an interesting problem. "ML is hard. Let's go online shopping!" Given that clustering is unsupervised, it's one of the easiest ML methods. Did they use kNN to categorize new instances? Did they use PCA to make the model simpler? I can cluster shit all day with no effort, but my higher-ups wouldn't be pleased.