3 ms·
The scoring approach seems interesting to extract the main content of web pages. I am aware of the large body of decades of research on that subject, with sophi
by la_fayette 2y ago
The scoring approach seems interesting to extract the main content of web pages. I am aware of the large body of decades of research on that subject, with sophisticated image or nlp based approaches. Since this extraction is critical to the quality of the LLM response, it would be good to know how well this performs. E.g., you could test it against a test dataset (https://github.com/scrapinghub/article-extraction-benchmark https://github.com/scrapinghub/article-extraction-benchmark). Also, you could provide the option to plugin another extraction algorithm, since there are other implementations available... just some ideas for improvement...
- leroman 2y agoThis totally makes sense, I will look into adding support for additional ways to detect the main content, super interesting!