4 ms·
> When I tried to do this I came to the conclusion that I needed to actually render the page to find out where on the page a particular piece of text was, what
by akiselev 3y ago
> When I tried to do this I came to the conclusion that I needed to actually render the page to find out where on the page a particular piece of text was, what font size it had, if it was even visible, etc. And then there's JavaScript of course.
Are there open source projects devoted to this functionality? It’s becoming more and more a sticking point for working with LLMs. Grabbing the text without navigation and other crap but while maintaining formatting and links, etc
- DeathArrow 3y agoThere are many software libraries that can output just the text from HTML or run JS. For C# there's HTML Agility Pack and PuppeteerSharp, for example. I did use them for web scrapping.
- fauigerzigerk 3y agoGood question (meaning I don't know :) For my specific purposes it has always been good enough to apply some simple heuristics. But that wouldn't have been possible without access to post rendering information, which only a real browser (https://pptr.dev https://pptr.dev) can reliably produce.