4 ms·
Nice landing page. If you Google "summarizer", you will find dozens of similar services for free. The mechanism behind it is very simple. A couple a years ago I
by blobster 7y ago
Nice landing page. If you Google "summarizer", you will find dozens of similar services for free. The mechanism behind it is very simple. A couple a years ago I built one from scratch in about 2 hours, then I accidentally deleted it and rewrote it in 15 minutes. Here's how most of them work:
1. Split the text into words
2. Rank each word based on how many times it appears in the text. For example, a word that appears 10 times gets 10 points, and so on.
3. Rank sentences based on the sum of the scores of each word inside them.
4. Return the top N sentences by score (N is up to the user), in the order in which they appear in the text.
For extra fancyness, exclude the most common articles and prepositions and give 2 points to proper nouns.
Works surprisingly well.
- bobbiechen 7y agoI first heard about this from the autotldr bot on Reddit [1] which uses a similar service SMMRY [2]. The SMMRY page points out some extra NLP-related grunt work in addition to the high-level steps you list, like: >Associate words with their grammatical counterparts. (e.g. "city" and "cities") >Detect which periods represent the end of a sentence. (e.g "Mr." does not). [1] https://www.reddit.com/r/autotldr/comments/31b9fm/faq_autotldr_bot/ https://www.reddit.com/r/autotldr/comments/31b9fm/faq_autotl... [2] https://smmry.com/about https://smmry.com/about
- radhakrsna 7y agoTrue, there are quite a few similar services but not many seem to work well. Our service provides better summarization (at least for the articles I tested), had additional features like extracting author name, publish data, important keywords etc and also comes with browsers extensions so you could summarize pages at the click of a button. The method you described is a part of our algorithms but more steps are needed to make it give meaningful results and make sure it works on different kinds of articles.
- bhl 7y agoYou can use tf-idf [1] to achieve step 2 and that extra fancy part of excluding commmon articles and prepositions: count the frequency of words in the article, but divide it by the sum of frequencies from past articles. Text summarization works as a good toy problem, because it leads to two harder problems: 1. text extraction (how to distinguish content from non-content like ads) 2. q&a (given text and a question about the text, how can you produce an answer). [1] https://en.wikipedia.org/wiki/Tf%E2%80%93idf https://en.wikipedia.org/wiki/Tf%E2%80%93idf
- shiredude95 7y agoI built a similar service a while back, with a small modification to the common algorithm. You can improve contextual summarization by splitting the x sentences into x/n buckets. Then based on the percent of article to be summarized (eg return 60% of the article), pick the sentences ranked in the top 60% of each bucket. Then do this for all the x sentences, ie top 60% across buckets, and combine them together. This prevents the bias rising from picking a sentence with a lot of critical words.
- mannykannot 7y agoI agree that the effectiveness is quite surprising, given the simplicity of the analysis. It can go very wrong, however, if a significant negation is overlooked, as in cautionary tales: ... So, don't do what the late Thag Simmons did... Maybe final paragraphs should be more highly weighted? That's often where the conclusion is.
- dstein64 7y agoI used a similar algorithm when developing a Chrome extension a few years ago: https://chrome.google.com/webstore/detail/auto-highlight/dnkdpcbijfnmekbkchfjapfneigjomhh https://chrome.google.com/webstore/detail/auto-highlight/dnk... Longer sentences had an inherent advantage, so I controlled for that by reducing sentence scores as a function of the sentence length.