4 ms·
You can usually find a html parser for your language, that you can use xpath/xsl on. It will just make the same assumptions that the browser does, by adding mis
by qw 5y ago
You can usually find a html parser for your language, that you can use xpath/xsl on. It will just make the same assumptions that the browser does, by adding missing closing tags etc.
I made a tool that extracted parts of web pages 10-15 years ago, and it worked well. There are of course cases where the html is so unstructured that the results were unpredictable, but it worked well in general.