4 ms·
Are you talking about something that can go out to any web site, grab the HTML, and turn it into a feed that's consistent with any other feed and not a big hair
by HankStallone 1y ago
Are you talking about something that can go out to any web site, grab the HTML, and turn it into a feed that's consistent with any other feed and not a big hairy mess of stuff that wasn't actually part of the content?
That doesn't sound trivial at all to me, so maybe I'm misunderstanding.
- hombre_fatal 1y agoIt's a trivial task for an LLM to take the HTML at https://techcrunch.com/latest/ https://techcrunch.com/latest/ and extract a perfect list of only the article URLs. And if you can do this, then you can build the tool that I describe. From there you can extract each article content with an LLM or use something like readability.js or just download the whole page for later consumption. I have a prototype of it. You add a feed as { feedUrl, feedPrompt, articlePrompt }. `feedPrompt` lets you append rules for extracting the article url list from the `feedUrl` like "ignore video articles". `articlePrompt` lets you append rules for how to extract article content for a given website, like "translate to english".