8 ms·
A Guide to Natural Language Processing
- kinow 9y agoA lot to review, read, learn. Thanks a lot for sharing this. Any plans to extend it or have another one including even more, like Natural Language Generation (not limited to bots, we are using it in weather forecast), and co-reference?
- umilegenio 9y agoThanks. Well, there are interesting things that we had to cut because they were too advanced for an introductory article. We were thinking about making a new article for them in a few months. And Natural Language Generation would be another great topic to talk about. However, if you already have experience in the topic we would be happy if you would like to write a guest post for us.
- nl 9y agoHa, there's a whole section on clones of the summarizer from Classifier4J. I wrote that in 2003 (I think?) based on @pg's "A plan for spam" essay, and then "invented" the summarization approach (I'm sure others had done similar, but I thought it up myself anyway). Turns out it was rather well tuned. The 2003 implementation, presumably downloaded from sourceforge(!) still wins comparisons on datasets which didn't even exist when I wrote it[1]. I much prefer the Python implementation though[2], which I hadn't seen before. Also, Textacy on top of Spacy is awesome for any kind of text work. [1] https://dl.acm.org/citation.cfm?id=2797081 https://dl.acm.org/citation.cfm?id=2797081 [2] https://github.com/thavelick/summarize/blob/master/summarize.py https://github.com/thavelick/summarize/blob/master/summarize...
- visarga 9y agoFirst time I see reading time and readability score mentioned together with NLP.
- pencilcode 9y agoRegarding finding similar documents what is the state of the art nowadays, LDA, word2vec, something else? What do you normally use?
- nicklovescode 9y agoHave you heard of word mover’s distance? It works really well!
- zintinio5 9y agoLike everything else, depends on your use-case. I have personally used TF-IDF vectors and token sets with Cosine and Jaccard distances in practice. Some examples of use-cases: are you searching for "semantically similar", or "near duplicate"? You can compare documents under different metrics and different _representations_. Some representations are: LSA, PLSA, LDA, TF-IDF, and Set representations, along with metrics such as Jaccard Distance, Cosine Distance, Euclidean distance, etc. Doc2vec is the Word2vec analog for documents.
- nl 9y agoWord Mover Distance on Word2Vec vectors. There is an implementation in Textacy.
- betageek 9y agoYour 'send me a PDF' popup has the background fade div above the form so it's impossible to fill in the form (without opening dev tools).
- owlninja 9y agoHmm, worked fine for me.
- umilegenio 9y agoThanks for your comment! Now, we have fixed the issue.
- paultopia 9y agoFYI, still a glitch: email form for pdf doesn't work right on mobile Safari for me---the cursor shows up in strange places unrelated to the form fields, have to click in random places to go from editing the name field to the email field.
- umilegenio 9y agoThanks for your comment. We are going to look into it.
- raarts 9y agoThe 'send me the PDF' pop-up can not be closed on my iPhone. Had to close the page.
- amelius 9y agoThere are a few applications missing: - Answering a question by returning a search result from a large body of texts. E.g. "How do I change the background color of a page in Javascript?" - Improving the readability of a text. The article only mentions "understanding how difficult to read is a text". - Establishing relationships between entities in a body of text. E.g. we could build a fact-graph from sentences like "Burning coal increases CO2", and "CO2 increase induces global warming". Useful also in medical literature where there are millions of pathways. - Answering a question, using a large body of facts. Like search, but now it gives a precise answer. - Finding and correcting spelling/grammatical errors.
- 1maginary 9y agoAuthor profiling comes to mind as well
- fnl 9y ago- Text generation and dialogue systems
- umilegenio 9y ago> - Answering a question by returning a search result from a large body of texts. E.g. "How do I change the background color of a page in Javascript?" > - Answering a question, using a large body of facts. Like search, but now it gives a precise answer. That is essentially a Natural Language Interface. There are simple ways to implement one for bots that receives simple commands[1]. The problem is that it quickly become very hard if you are trying to do something more open ended that a bot. So, there was simply no room to include it. > - Improving the readability of a text. The article only mentions "understanding how difficult to read is a text". The issue is that the formulas to measure the readability of a text cannot really be used to suggest improvements. That's because the user ends up focusing on improving the score instead of improving the text. To suggest improvements you need a much more sophisticate system. > - Establishing relationships between entities in a body of text. E.g. we could build a fact-graph from sentences like "Burning coal increases CO2", and "CO2 increase induces global warming". Useful also in medical literature where there are millions of pathways. This is one of the things that were axed, because in some sense it is simple if you just want to link together concepts without any causality, i.e. stuff that happens together. To do that you could link named entity recogniton (to find entities) and a simple way to find a relationship between words (i.e., they happen in the same phrase therefore they have related). However a more sophisticated form of the process, like the one that results in the Knowledge Graph[2] would be quite hard to do. > - Finding and correcting spelling/grammatical errors. That's a great idea, we will add how to detect spelling errors. [1] https://medium.com/swlh/a-natural-language-user-interface-is-just-a-user-interface-4a6d898e9721 https://medium.com/swlh/a-natural-language-user-interface-is... [2] https://en.wikipedia.org/wiki/Knowledge_Graph https://en.wikipedia.org/wiki/Knowledge_Graph
- bhaak 9y ago> Essentially, when dealing with natural languages hacking a solution is the suggested way of doing things, since nobody can figure out how to do it properly. That's really the TL;DR I also got from the computational linguistic courses I attended. There's probably the Pareto principle at works. Having no solution is worse than having an 80% solution that works well enough when the 100% solution is much harder to achieve (and some of the problems not even humans would be able to solve properly).
- nlperguiy 9y agoYeah, when you look at some of the SemEval contest winners or top 3, many use fairly simple methods combined into a powerful solution (except when LSTM with attention grabs the throne).
- deleted 9y ago[deleted]
- cvs268 9y agoCouldn't agree more! Recently I wrote a web-extension for Firefox that displays funny "Deep thought" quotes. I wanted to analyse the quote text and fetch relevant images to animate in the background of the quote text. After reading several NLP tutorials, guess what I did as a first PoC - Pick the 3 longest words in a quote text and run an image search with those 3 words. 6 lines of plain javascript code that can be run anywhere almost instantly https://github.com/TheCodeArtist/deep-thought-tabs/blob/master/addon-src/deepThoughts.js#L87 https://github.com/TheCodeArtist/deep-thought-tabs/blob/mast... I get relevant images in the search results 99/100 times. The quirks of searching often result in the image adding to the funny-ness of the "Deep Thought" on display. Its so effective that i ended up publishing the "Deep Thought Tabs" web-extension with this approach itself: https://addons.mozilla.org/en-US/android/addon/deep-thought-tabs/ https://addons.mozilla.org/en-US/android/addon/deep-thought-... Later I tried using the nlp-compromise js library to identify "topics" of interest within a quote text - typically nouns, verbs, and adjectives. Comparing the results with my "3-longest-words" approach, I found that the longest words were anyways almost always the "topic" words that NLP identified for any given quote text.
- Boothroid 9y agoQuite an obnoxious website on my phone. Anyway I came here to point to GATE as a mature FLOSS option: https://gate.ac.uk/ https://gate.ac.uk/
- fnl 9y agoI'm always astonished how little mention gensim gets, considering that it can basically be used for all the listed tasks, including parsing, if you combine it with your favorite deep learning library (DyNet, anyone?).
- rpedela 9y agogensim is one of the best libraries for word vectors and summarization. For parsing and NER, Stanford CoreNLP works best in my experience.
- fnl 9y agoWell, a model you fine tune to your specific corpus/domain works even (in fact: much) better... And gensim there gives you the tools to build the best possible embeddings. But you do need a use case and an economic reward for the substantial increase in cost than a pre-trained, vanilla, off-the-shelf parser (model) can give you. Yet, if your domain is technical enough (pharma, finance, law, ... - essentially, all but parsing news, blogs, and tweets...) it might be the only way to get a NLP system that really works.
- bane 9y agoWas hoping for some discussion about word vectors like word2vec. I keep reading about them, but don't really understand what they're useful for.
- matt4077 9y agoLet me try: Take the famous example of [king] and [queen] being close neighbors in vector space after generating the word vectors ("embedding"). If you then use these vectors to represent the words in your text, a sentence about kings will also add information about the concept of queens, and vice versa. To a far lesser degree, such a sentence will also add to your knowledge of [ceo], and, further down, [mechanical engineer]. But it will not change the system's knowledge of [stereo].
- bane 9y agoThanks, yeah I get that, but I think I'm having a lack of imagination about what to do with that in terms of how to build something useful and user friendly out of it.
- umilegenio 9y agoThe interesting thing about word2vec is that is an unsupervised method that build vectors to represent each word in a way that makes easy to find relationship between them. There is a video by the creator of Gensim on word2vec and frieds: https://www.youtube.com/watch?v=wTp3P2UnTfQ https://www.youtube.com/watch?v=wTp3P2UnTfQ We didn't include it, simply because it relies on machine learning and we wanted to show simpler methods.
- laGrenouille 9y agoYes, I agree that the applications for word vectors are not made as clearly as it should be. One direct application is as the first layer of a neural network [1], which could be part of either a 1-dimensional convolution or a recurrent neural network. Using pre-trained word vectors is a form of transfer learning and allows for much more predictive models with smaller amounts of training data. [1] https://blog.keras.io/using-pre-trained-word-embeddings-in-a-keras-model.html https://blog.keras.io/using-pre-trained-word-embeddings-in-a...
- rpedela 9y agoUsing Chrome on both a Chromebook and Galaxy S5, the right sidebar is screwed up. On the phone, it completely blocks the content.
- arcanus 9y agoIs there an equivalent to MNIST for NLP? I've always wanted to play around in this space but I don't know a good, and simple, database to start with.
- nl 9y agoDepends on what you want to try. NLTK has built in datasets. 20 Newsgroups is useful for trying lots of things.
- gumby 9y agoWell there's word2vec, which while it isn't quite the same (its whole point is the vector classification it already embodies), I think is actually the kind of think you were asking for.
- jventura 9y agoI worked with NLP for my research, and I used to build my corpora from wikipedia documents. Here's a tool that I've built to do it: https://github.com/joaoventura/WikiCorpusExtractor https://github.com/joaoventura/WikiCorpusExtractor
- jimsmart 9y agoThere are a few different datasets that might be of use, depending on what you're playing with:- - bAbI https://research.fb.com/downloads/babi/ https://research.fb.com/downloads/babi/ and https://github.com/facebook/bAbI-tasks https://github.com/facebook/bAbI-tasks - SQuAD https://rajpurkar.github.io/SQuAD-explorer/ https://rajpurkar.github.io/SQuAD-explorer/ - WebQuestions https://github.com/brmson/dataset-factoid-webquestions https://github.com/brmson/dataset-factoid-webquestions Edit: there's also a great list of datasets on the ParlAI project page https://github.com/facebookresearch/ParlAI https://github.com/facebookresearch/ParlAI
- d23 9y agoMy experience with your site on mobile: https://m.imgur.com/5vLrEJH https://m.imgur.com/5vLrEJH Can't get it to go away, can't read the article.
- alexasmyths 9y agoRecommend Dan Jurafsky and Chris Manning @ Stanford online course: https://www.youtube.com/watch?v=nfoudtpBV68 https://www.youtube.com/watch?v=nfoudtpBV68