5 ms·
Nice! Curious to know what you're using to parse the essays into sentences. Isn't that quite difficult in the general case?
by jsomers 14y ago
Nice!
Curious to know what you're using to parse the essays into sentences. Isn't that quite difficult in the general case?
- eggbrain 14y agoLooks to be mostly just splitting by the period -- see article "Startup == Growth", page 17 near the bottom (splits on 1.7x at the period, although it shouldn't)
- flixic 14y agoBetter algorithm, just as simple, would be splitting at ". ". Making something smart enough to not split at "fig. 1" and the likes would be more difficult.
- mindcrime 14y agoSentence detection is a hard problem, and I don't know that anybody has solved it for the general case. But pragmatically speaking, it is solved for many - if not most - common cases, and there are Open Source NLP libraries out there that do a pretty good job. For Java, you have OpenNLP and Mallet, the C/C++ world has Ellogon and Freeling, Python has NLTK, etc.
- samsnelling 14y agoThought I'd jump in here. Yes, parsing sentences are actually quite difficult. Disclaimer: I am currently working on a text summarization startup, http://Summary.io http://Summary.io There are a few different approaches to parsing sentences, and only a few giant NLP libraries. What I have found the best in 95% of use cases to to write custom (RegEx) rules. Attempting a sentence such as: And then Mr. Bean (http://www.mrbean.com/?index http://www.mrbean.com/?index) said to Col. Sanders, "Holy moly sentence extraction is hard!" You have lots of little things like making sure you take the full quote and disregard periods in surnames, http links, etc. Parsing just by periods, question marks, and exclamtion points are going to lead to a lot of problems.