8 ms·
Deep Text Correcter
- camoby 10y agoIs the spelling of the name supposed to be ironic? ;)
- saycheese 10y ago(Comment removed.)
- delinka 10y agoIs the final word not supposed to be "corrector?" https://en.m.wikipedia.org/wiki/Corrector https://en.m.wikipedia.org/wiki/Corrector
- kurthr 10y agoWell, I think the dictionary disagrees. http://www.dictionary.com/browse/corrector http://www.dictionary.com/browse/corrector I'm going to assume that it was meant to be ironic (or errorful).
- hbornfree 10y agoThis is terrific! Very coincidentally, we were thinking of implementing a sentence de-noiser using sequence-to-sequence models only today evening. I work in the NLP domain writing Machine Translation systems. But NLP parsers are accurate for grammatically correct sentences only which necessitates the need for something like deep text correcter. Thank you for this. Will try this out this week and let you know how it goes.
- daveytea 10y agoThis is really cool. If you're looking for more datasets to train your model, here are a few relevant ones: - https://archive.org/details/stackexchange https://archive.org/details/stackexchange - http://trec.nist.gov/data/qamain.html http://trec.nist.gov/data/qamain.html - http://opus.lingfil.uu.se/OpenSubtitles2016.php http://opus.lingfil.uu.se/OpenSubtitles2016.php - http://corpus.byu.edu/full-text/wikipedia.asp http://corpus.byu.edu/full-text/wikipedia.asp OR https://en.wikipedia.org/wiki/Wikipedia:Database_download#English-language_Wikipedia https://en.wikipedia.org/wiki/Wikipedia:Database_download#En... - http://opus.lingfil.uu.se/ http://opus.lingfil.uu.se/ I'd love to see how good your model gets.
- atpaino 10y agoThanks for the links! I'll have to try out some of these. The data is definitely the limiting factor at this point.
- siscia 10y agoHave you consider project Gutenberg? Is there any reason why that huge corpus would not be useful? https://www.gutenberg.org/ https://www.gutenberg.org/
- gleenn 10y agoI think the issue is you need parallel corpuses of (sentences with grammatically errors) -> (same sentence but correct). Gutenberg has a lot of only the latter I think.
- stephanheijl 10y agoLooks like a cool project, I would love to see this as a browser plugin of some sort. As for the corpus, I suspect that using articles from Wikipedia would be appropriate. Especially large articles are routinely checked and cleaned up. It has the added benefit of being available in multiple languages. (https://en.wikipedia.org/wiki/Wikipedia:Database_download https://en.wikipedia.org/wiki/Wikipedia:Database_download) EDIT: I see this has already been suggested, along with a large amount of other source in another comment by daveytea.
- jmiserez 10y agoInteresting idea. I went ahead and tested: > Alex went to the kitchen to store the milk in the fridge. Corrected: > Alex went to the kitchen to the store the milk in the fridge. Gathering a large, high quality dataset from the internet is probably not so easy. A lot of the content on HN/Reddit/forums is of low quality grammatically and often written by non-native English speakers (such as myself). Movie dialogues don't necessarily consist of grammatically correct sentences like the ones you'd write in a letter. Perhaps there is some public domain contemporary literature available that could be used instead or alongside the dialogues? EDIT: Unrelated to this project, I have this general fear of language recommendation tools trained on just low-quality comments or emails. A simple thesaurus and a grammar-checker are often enough to find the right words when writing. But a tool that could understand my intent and then propose restructured or similar sentences and words that convey the same meaning could be a true killer application.
- babuskov 10y ago> A lot of the content on HN/Reddit/forums is of low quality grammatically and often written by non-native English speakers (such as myself). Corrected: > A lot of the content on the HN/Reddit/forums is of the quality grammatically and non-native written by written English speakers (such myself). as UNK Yeah. It's got a long way to go. No idea where "as UNK" came from.
- pilooch 10y agoUNK usually comes from the final sampling step: the distribution of words contains a special token UNK to represent all words with insufficient statistics in the training corpus.
- pizza 10y agoIt's like the equivalent of NaN; what you put when you don't know what should be there. coincidentally, the op corrected this as: It's like the equivalent of NaN
- tannerc 10y ago> A tool that could understand my intent and then propose structure or similar sentences and words I'm working on this exact thing right now fwiw, in my app called Prompts. I imagine machine assisted apps are going to percolate up over the coming years, or months, in a way we haven't seen yet. It's pretty exciting, if we can get it right. Written language seems like a decent place to start.
- topynate 10y ago> Unfortunately, I am not aware of any publicly available dataset of (mostly) grammatically correct English. How about books?
- wiz21c 10y agoGuess who's got that massive amount of data nowadays ? Google, facebook and co. So good luck if you want to work on AI stuff...
- mortehu 10y agoAnd Project Gutenberg: https://www.gutenberg.org/ https://www.gutenberg.org/ The language in many of those books may be a bit archaic, though. See also Distributed Proofreaders: http://www.pgdp.net/c/ http://www.pgdp.net/c/
- jcoffland 10y agoThere are a fair number of OCR mistakes in many of the Gutenberg texts. It would be interesting to try to correct them. They are however not the same types of errors which are addressed here.
- rspeer 10y agoWho's got that massive amount of data nowadays? Everybody who wants it. I recommend http://opus.lingfil.uu.se/ http://opus.lingfil.uu.se/ as a starting point to find lots of sources of quality, multilingual data. And if all you want is sheer quantity, at rather the expense of quality, then there's a several terabyte Web crawl at http://commoncrawl.org/ http://commoncrawl.org/.
- raverbashing 10y agoWikipedia, Books are usually ok Online forums are a disaster (HN is not too bad, but there are a lot of natives and non-natives fumbling some things)
- smoyer 10y ago'Kvothe went to the market' Off-topic but I've been waiting for the third and final book in the trilogy for a long time ... I've come to the conclusion that Rothfuss can't find a way to tie all the plot threads together. I'm also wondering if anyone else thinks Rothfuss looks like Longfellow in a lot of his publicity shots.
- saycheese 10y agoDeep Proofreading Tool Comparisons: http://www.deepgrammar.com/evaluation http://www.deepgrammar.com/evaluation https://blogs.nvidia.com/blog/2016/03/04/deep-learning-fix-grammar/ https://blogs.nvidia.com/blog/2016/03/04/deep-learning-fix-g...
- UhUhUhUh 10y agoRe. intent but with regards to spelling. I often wonder if there could be rules to correct errors due to key strokes in the immediate vicinity of the intended letter (e.g. "keu" vs. "key"). Would check combination of the surrounding letters, first oin (that's an unintended addition here) horizontal axis. That happens to me all the time, probably because I'm not a good typist but still.
- danso 10y agoTried out some of classic Garden path sentences [0], and of the 4 examples, it got all but one right: Original: The complex houses married and single soldiers and their families. Deep Text Corrector: The complex houses married and a single soldiers and their families. OT: does anyone know of a more substantial list of garden path sentences that people use in testing NLP software? [0] https://en.wikipedia.org/wiki/Garden_path_sentence https://en.wikipedia.org/wiki/Garden_path_sentence
- zitterbewegung 10y agoThis is a neat project . I think as a follow up step he should compare it to Word's grammar checker .
- brandonb 10y agoInteresting idea! I think this is analogous to the idea of a de-noising autoencoder in computer vision. Here, instead of introducing Gaussian noise at the pixel level and using a CNN, you're introducing grammatical "noise" at the world level and using an LSTM. I think that general framework applies to many different domains. For example, we trained a denoising sequence autoencoder on HealthKit data (sequences of step counts and heart rate measurements) in order to predict whether somebody is likely to have diabetes, high blood pressure, or a heart rhythm disorder based on wearable data. I've also seen similar ideas applied to EMR data (similar to word2vec). It's worth reading "Semi-Supervised Sequence Learning", where they use a non-denoising sequence autoencoder as a pretraining step, and compare a couple of different techniques: https://papers.nips.cc/paper/5949-semi-supervised-sequence-learning.pdf https://papers.nips.cc/paper/5949-semi-supervised-sequence-l... Toward the end, you start thinking about introducing different types of grammatical errors, like subject-verb disagreement. I think that's a good way to think about it. In the limit, you might even have a neural network generate increasingly harder types of grammatical corruptions, with the goal of "fooling" the corrector network. As the the corruptor network and corrector network compete with each other, you might end up with something like a generative adversarial network: https://arxiv.org/abs/1701.00160 https://arxiv.org/abs/1701.00160
- atpaino 10y agoI like the analogy to de-noising autoencoders; that's a good way of thinking about this. > In the limit, you might even have a neural network generate increasingly harder types of grammatical corruptions, with the goal of "fooling" the corrector network. Very interesting. I wonder how many constraints would need to be added to the corruptor model to ensure the corrupted sentence retains the same meaning as the original. Somewhat related to that, I've thought that a more basic curriculum learning setup could be deployed quite effectively here, and am hoping to try that out soon.
- unhammer 10y agoIt seems like it would be challenging to get the corruptor to generate examples that are of the same Kind that humans make, while still being "productive" (in the linguistic sense, ie. not just overfitting on examples from a corpus of low quality text). It's easy enough to just drop random words or run Levenshtein-edits on single words to create wrong-in-this-context makes (then/than), but grammar errors include much more than can be covered by that method, and the method will generate many errors that are of a kind never made by humans. And if you restrict your method to things already seen in a corpus, it's easy to overfit and miss out on a whole lot of good stuff.
- ashildr 10y agoAnd so it begins: http://www.goodreads.com/book/show/13184491-avogadro-corp http://www.goodreads.com/book/show/13184491-avogadro-corp
- OJFord 10y agoI can't make it work for anything other than the missing 'the' example. For example: > Do you know where I been 'corrects' to: > Do you know where I 's been
- sigmonsays 10y agobut will it correct "lets eat grandma"
- raverbashing 10y agoWhile there are limited errors it can officially correct, I tried a few phrases: Didn't fix misuse of its: "The tool worked on it's own power" "He should of gone yesterday" gets corrected to "the He should of gone yesterday" "To who does this belong?" doesn't get corrected "A Apple a day keeps the doctor away" doesn't change
- johnhenry 10y agoI tried these, which stem from the given example: "I'm going to store" gets corrected to "I'm going to the store" (looks good) "I'm going to store food" gets corrected to "I'm going to the store store" (I'll forgive this due to lack of a prepositional phrase) "I'm going to store food in the closet" gets corrected to "I'm going to the store food in the closet" (yeah, this is wrong) It seems like the AI is good at fixing sentences where adding a missing "the" is the solution. So when that does turn out to be the solution, it works fine. When that isn't the solution, well, when all you have is a hammer, every problem looks like a nail... Still really interesting, though. The AI may not [yet] be able to solve general grammatical errors, but it may good enough that it works in specific cases. It may even be the case that we'll design systems that depend on multiple agents offering suggestions and other aggregating suggestions into a more refined result.
- tyingq 10y agoIt does seem to pick up on the specific cases where adding "the" isn't right, like: Alex went to school, Alex went to bed, Alex went to prison
- deleted 10y ago[deleted]
- WhitneyLand 10y agoAlex nice work this is exciting. I've been wanting to work on something similar because the quality of common grammar checkers (like MS Word) has made such little progress. Have you considered combining your approach with rules based system? Some systems using only an elaborate set of rules for common mistakes have had pretty good performance. I wonder if these two approaches could be combined. Btw, what is the highest performing grammar checker you've found that's commonly available?
- guelo 10y agoWhat about using books such as from Project Gutenberg?
- macawfish 10y ago"I gotta take shit" -> "I gotta take the shit" Sorry to be airing out my personal business and everything but... everybody poops!
- BuuQu9hu 10y agohttps://www.youtube.com/watch?v=kQTW7Pd1vqc https://www.youtube.com/watch?v=kQTW7Pd1vqc
- koliber 10y agoWould the works archived in Project Gutenberg be a good training corpus?
- dado12500 10y agohttps://m.youtube.com/watch?v=Wo7DDBwPwzs https://m.youtube.com/watch?v=Wo7DDBwPwzs
- dado12500 10y agohttps://m.youtube.com/watch?v=Wo7DDBwPwzs https://m.youtube.com/watch?v=Wo7DDBwPwzs
- mikeflynn 10y agoInteresting project, and I love the example they used on the demo page. (Go Cardinals!)
- YeGoblynQueenne 10y ago>> Thus far, these perturbations have been limited to: + the subtraction of articles (a, an, the) + the subtraction of the second part of a verb contraction (e.g. “‘ve”, “‘ll”, “‘s”, “‘m”) + the replacement of a few common homophones with one of their counterparts (e.g. replacing “their” with “there”, “then” with “than”) Oooh, that's _very_ tricky what they're trying to do there. "Perturbations" that cause grammatical sentences to become ungrammatical are _very_ hard to create, for the absolutely practical reason that the only way to know whether a sentence is ungrammatical is to check that a grammar rejects it. And, for English (and generally natural languages) we have no (complete) such grammars. In fact, that's the whole point of language modelling- everyone's trying to "model" (i.e. approximate, i.e. guess at) the structure of English (etc)... because nobody has a complete grammar of it! Dropping a few bits off sentences may sound like a reasonable alternative (an approximation of an ungrammaticalising perturbation) but, unfortunately, it's really, really not that simple. For instance, take the removal of articles: consider the sentence: "Give him the flowers". Drop the "the". Now you have "Give him flowers". Which is perfectly correct and entirely plausible, conversational, everyday English. In fact, dropping words is de rigeur in language modelling, either to generate skip-grams for training, or to clean up a corpus by removing "stop words" (uninformative words like the the and and's) or generally, cruft. For this reason you'll notice that the NUCLE corpus used in the CoNLL-2014 error correction task mentioned in the OP is not auto-generated, and instead consists of student essays corrected by professors of English. tl;dr: You can't rely on generating ungrammaticality unless you can generate grammaticallity.
- Houshalter 10y agoSo what? Who says the training data has to be perfect? Your comment reminds me of an essay by Peter Norvig. Where he showed that a simple statistical model could tell the difference between a nonsense sentence that was grammatically correct and one that wasn't, but gave them both very low probability of occuring naturally. Similarly, the machine may learn that "Give him the flowers" has higher probability than "give him flowers". Or it may learn that both are possible sentences and not be able to correct it. Which is also OK, we can't expect any system to be perfect.
- 10y ago
- grizzles 10y agoFor training data, you could try ebook torrents, eg. books with a creative commons license.
- vacri 10y ago> "Kvothe went to market" This is not a grammatically incorrect sentence; it depends on context. Products are take to an abstract concept of 'market', for example.
- jrapdx3 10y agoThe "correcter" is a worthy effort, and it needs to start somewhere. It shows the magnitude of the task considering that missing articles are not the most crucial grammatical issues in on-line discourse. The meaning of a phrase is usually comprehensible with or without the article, and native speakers can easily overlook this kind of error made by non-native speakers. OTOH more troublesome to readers are common errors such as misuse of "its" vs. "it's", "to" vs. "too", and "their", "there" and "they're". These mistakes are quite prevalent among native-speaking writers so more ubiquitous than the missing article problem. The "correcter" didn't correct the latter class of errors. Understandably this would be a much harder goal to accomplish given the highly contextual nature of grammatically correct word choices. It prompts a question about how well the data-driven approach can handle the problem. Obviously that's what the research is trying to answer. It sure seems to point to something fairly easy for a human to do that's near or at the limit of what we can get a computer to do.
- ematvey 10y agoNice work! I was playing with exactly this idea for some time. Potentially it could be way bigger than simple grammatical corrections. My list of things to try, in addition to what you've already done: - replacing named entities with metadata-annotated tokens; - dropping random words, not just articles; - replacing random words with rarer synonyms; - annotate with POS tags from some external parser; - run syntax corrector before feeding sentences in grammatical model; I think this problem is easier that it appears on the surface. Generated deformation does not have to be a perfect replica of typical human errors. It just have to be sufficiently diverse. Also, I think seq2seq module is getting deprecated, as it doesn't do dynamic rollouts.
- burnbabyburn 10y agoisn't a ngram bayesian model sufficient here? also, isn't the test data linear dependent from the training set here so creating skewed performance measurement?