3 ms·
I've got an IOException while trying to summarize http://matt-welsh.blogspot.com/2013/04/running-software-team-at-google.html http://matt-welsh.blogspot.com/201
by cpio 14y ago
I've got an IOException while trying to summarize http://matt-welsh.blogspot.com/2013/04/running-software-team-at-google.html http://matt-welsh.blogspot.com/2013/04/running-software-team...
And a different one for http://googleblog.blogspot.com http://googleblog.blogspot.com
I guess you should put more effort in your html parser. Try Apache Tika, perhaps.
- mohaps 14y agoah, it won't work directly for webpages (html). the url is expected to be that of a RSS/Atom feed. for the html web pages, copy pasting the text to the textarea works. Will try to add url content type detection in the next cut and summarizing non-feed url's next up
- mohaps 14y agotry this url: http://matt-welsh.blogspot.com/feeds/posts/default?alt=rss http://matt-welsh.blogspot.com/feeds/posts/default?alt=rss it works. same feed url pattern will work for google blog too http://googleblog.blogspot.com/feeds/posts/default?alt=rss http://googleblog.blogspot.com/feeds/posts/default?alt=rss
- mohaps 14y agookay, added a fix to try and extract article text from non-feed urls. try http://tldrzr.herokuapp.com/tldr/?feed_url=http://matt-welsh.blogspot.com/2013/04/running-software-team-at-google.html http://tldrzr.herokuapp.com/tldr/?feed_url=http://matt-welsh... :)