4 ms·
Nice work. Main content extraction based on the <main> tag won’t work with most of the web pages these days. Arc90 could help.
by gradientDissent 2y ago
Nice work. Main content extraction based on the <main> tag won’t work with most of the web pages these days. Arc90 could help.
- leroman 2y agoThank you! this is exactly why there's support for this specific use case- https://github.com/romansky/dom-to-semantic-markdown/blob/main/src/core/domUtils.ts https://github.com/romansky/dom-to-semantic-markdown/blob/ma... (see `findContentByScoring`) And if you pass an optional flag `extractMainContent` it will use some heuristics to find the main content container if there is no such tag..