6 ms·
Sorry, i've some experience in this field: - This is text-extraction, NOT text-generation - TextRank algorithm is so far fine, but it does not write a "summar
by hacker089 7y ago
Sorry,
i've some experience in this field:
- This is text-extraction, NOT text-generation
- TextRank algorithm is so far fine, but it does not write a "summary", instead it ranks the components of a text according to some "metrics" (simply spoken)
- Using this approach will still make you attackable by copyright claims from copyright owners
- Which stuff is summarized ("put to the final output") is not always clear to me in your implementation, i tried it on some newspaper & blog articles; on some it worked well, on others it didn't.
Funny thing is, i'm currently working on something similar with a slightly different twist - i will post it here if finished, than we can go into a battle :-)
- steve19 7y agoI'm interested in learning about true generation algorithms. Can you point me in the right direction?
- p1esk 7y agoGoogle “gpt-2”.
- steve19 7y agoThanks but that is not what the OP is claiming. That generates text from a seed, the OP is talking about an article that generates a summary of an article, but without using existing sentences.
- p1esk 7y agoDid you read the paper? https://d4mucfpksywv.cloudfront.net/better-language-models/language_models_are_unsupervised_multitask_learners.pdf https://d4mucfpksywv.cloudfront.net/better-language-models/l... You don’t need any seed, and can generate summaries (section 3.6). GPT-2 is the model to learn about if you’re interested in NLP.
- applecrazy 7y agoAre there any papers benchmarking a transformer NN architecture in comparison to something like a pointer-generator network? I'm doing a bit of work in this area (i.e. reimplementing papers), and I'm curious if GPT2-like models can derive greater semantic meaning.
- p1esk 7y agoBoth GPT-2 and pointer-generator network are open source, and pretrained models are available, so it should be straightforward to compare them.
- radhakrsna 7y agoHi, - Yes, you are totally right. TextRank is an extractive method. - Ah right. But we aren't storing any information on our servers, just showing selected sentences to the user. - We select the top 5 sentences which have the highest relevancy to the article. I am not an expert in this field so not too sure if that's the best way. Just started with NLP a few days back and wanted to test it out by developing a small application. Yes, it works on quite a few articles and but also there are some articles where it fails to give accurate results. Ah nice, I would like to hear more about what you are working on. Let me know if I could contribute to it in some way. Thank you again for your feedback.
- andreareina 7y ago> But we aren't storing any information on our servers, just showing selected sentences to the user. If this is in reference to the copyright comment, it doesn't matter -- you're still transmitting/redistributing the content, which is what matters. One way to get around this is to ship the code and have the code execute on the user's machine (i.e. what you're presumably doing with the extension).
- radhakrsna 7y agoAh right, thank you for the detailed explanation. Currently, we are processing the text using a Python backend. In order to process it on the user's side, I guess we'll have to use Javascript. I will try to fix that in the next version. Thank you very much.