4 ms·
Thanks for the links I had no idea those existed. For my article web scraper (wip) the current steps are: - Navigate with playwright + adblocker - Run mozill
by msp26 2y ago
Thanks for the links I had no idea those existed.
For my article web scraper (wip) the current steps are:
- Navigate with playwright + adblocker
- Run mozilla's readability on the page
- LLM checks readability output
If check failed
- Trim whole page HTML context
- Convert to markdown with pandoc
- LLM extracts from markdown
- privatenumber 2y agoMozilla has released Readability as a standalone package so you can avoid spinning up a browser entirely: https://github.com/mozilla/readability https://github.com/mozilla/readability
- asadalt 2y agoyou would still need to run. For js based websites.
- msp26 2y agoI still wanted the browser for UBlock Origin and handling sites with heavy JS. I was using the standalone Readability script already but today I ended up dropping it for Trafilatura. It works a lot better. The inefficiency of using a browser rather than just taking the html doesn't really matter because the limiting factor is the LLM here. And yes the LLM is essential for getting clean data. None of the existing methods are flexible enough for all cases even if people say "you don't need AI to do this".