4 ms·
Teaching a Computer to Read: NLP Hacking in Python
- adpreese 13y agoHow do you deal with keeping your super rare words list sensible? For many forms of technical writing I could see things getting out of hand where you have lots of tiny dense clusters not really close to anything else if you didn't manage the list well.
- msalahi 13y agoAs with most successful applications of machine learning, it's about finessing your approach based on the problem at hand. In our case, we have classes divided on the level of "Medicine," "Real Estate," etc. So, we could throw away lots of words that only occurred once or twice in the massive corpus we crawled to build the language model and still have a pretty robust representation of the subject you're trying to represent.
- msalahi 13y agoIn fact, if your training corpus is sufficiently large, you'd be shocked how many words you can eliminate right away for a term frequency of one or two. I went from millions of words in the vocabulary to something like 60k just by ignoring words that happen once or twice in the corpus. Plus, you probably won't learn much about the relationships between words if they only occur a few times in the corpus.
- jlees 13y agoYeah, but consider that some rare words are much stronger indicators of topic than more common ones. Even more so if you look at n-grams. If you use something like wordnet you can get a lot of meaning out of low-frequency words and throw away the meaningless higher-frequency ones that occur in too many categories to be useful.
- adpreese 13y agoSure, there's value in rare words, but I don't think anything that occurs across the corpus fewer than 3 times is going to tell you anything useful. You need a certain amount just to have it be a real signal. What was the least frequent useful word in the data set, msalahi?
- samuizo 13y agoYou can often strike a balance between rare words that appear in only a couple of documents and very frequent words that occur all over the place by employing both a term frequency and a document frequency weighting scheme; 'tf-idf' in the nomenclature [1]. The basic idea is that you keep track of counts both within documents and among documents. For English, word like 'the' will be frequent in each document it occurs in. It will also occur in every document. The high document frequency counteracts the high term frequency. On the other hand, 'motherboard' might be infrequent overall (but not extremely so), but its low document frequency boosts its importance. The scheme is commonly employed and works quite well, sometimes obviating the need for careful vocabulary pruning. FWIW, scikit-learn implements it in their feature extraction library [2]. [1] http://en.wikipedia.org/wiki/Tf–idf http://en.wikipedia.org/wiki/Tf–idf [2] http://scikit-learn.org/stable/modules/generated/sklearn.feature_extraction.text.TfidfTransformer.html http://scikit-learn.org/stable/modules/generated/sklearn.fea...
- drakaal 13y agoThe CPU cost to do use this approach is terribly high. I don't think this approach is going to give better results than a few simple rules and NLTK would. This API will do a better job telling you what an article is about. https://www.mashape.com/stremor/stremor-noun-phrase-and-part-of-speech-tagging-alpha https://www.mashape.com/stremor/stremor-noun-phrase-and-part... That said, the approach we use for our TLDR software and search rankings doesn't rely on just frequency, the adjectives that amplify the content, the sentences with emotion attached to them, and the "charge" of words matters too much. Consider the following: That frakking loser Drakaal came over and hijacked my NLP thread. Just because he does NLP for a living, and thinks he knows everything doesn't mean a thing. My NLP is way cooler because it uses machine learning and that is the future of NLP, not the heuristics model he uses for his stuff. What is the "core" of that? Clearly it is about how Drakaal sucks, but we only mention him once. NLP is important, machine learning is important, but really it is about why Drakaal sucks.
- msalahi 13y agoi've actually found the performance of gensim (the topic modeling python module i use here) to be pretty great. we're not at a scale where CPU performance is make or break just yet, so i haven't done any comprehensive testing of performance. but i've definitely not run into any performance issues worth complaining about. however, gensim is 100% based on lazy evaluation where it can be, so it's relatively light on the CPU. i love NLTK as well, but it did lack in the dimensionality reduction/topic modeling department which gensim did so beautifully. LDA + SVM seemed like an interesting approach to go with, and it didn't disappoint.
- drakaal 13y agoThe issue with Genism is you have to know what you are trying to analyze before you analyze it. It doesn't do well if you use the wrong corpus or if like you mention start with a million word corpus. If you were analyzing emails in a single organization all day you could probably sort out topics really well. Doing all of the web it breaks down because it gets less accurate the larger the variety of content.
- taariqlewis 13y agoInteresting to see a content marketplace company using technology well ahead of its peers. I wonder how many other firms are pursuing this type of commercial research.
- rbucks 13y agoGreat question! Probably not very many.
- tdj 13y agoActually, I think you could save yourself some trouble and use scikit-learn's built-in text preprocessing utils: Word counter: http://scikit-learn.org/stable/modules/generated/sklearn.feature_extraction.text.CountVectorizer.html http://scikit-learn.org/stable/modules/generated/sklearn.fea... Hashing vectorizer if you want to trade off explainability for speed and scalability: http://scikit-learn.org/stable/modules/generated/sklearn.feature_extraction.text.HashingVectorizer.html http://scikit-learn.org/stable/modules/generated/sklearn.fea... TF-IDF weighing: http://scikit-learn.org/stable/modules/generated/sklearn.feature_extraction.text.TfidfVectorizer.html http://scikit-learn.org/stable/modules/generated/sklearn.fea... Also, if you transform bag-of-words vectors into a dense form, you're gonna have a bad time (insert appropriate meme picture here). In large corpora, dimensionality grows quite substantially - if you work with news corpora or Wikipedia, you're in the 100k-1M dimensional space pretty quickly. Great to see an approachable explanation for NLP. As they say sometimes, when you know how it's done, it stops being "Artificial Intelligence".
- aas48 13y agoGenius