3 ms·
The HTML parsing part appears to be "golang.org/x/net/html". Does anyone have experience parsing "real world" html with this? How does it do?
by gjem97 9y ago
The HTML parsing part appears to be "golang.org/x/net/html". Does anyone have experience parsing "real world" html with this? How does it do?
- jerf 9y agoIt's an HTML5-compliant parser. That means that, modulo bugs, it should produce the same results as any other modern HTML parser, which should also all be based on HTML5. For context, since I think a lot of people are still unaware of this, the HTML5 standard precisely specifies how HTML should be parsed: https://www.w3.org/TR/html5/syntax.html#parsing https://www.w3.org/TR/html5/syntax.html#parsing This is based on a survey of how the various browsers were handling it in reality, so it's not just one of those theoretical things that everybody ignores, it's an algorithm extracted from the brutal pragmatism of many separate code bases over many years. In theory now, all HTML parsing libraries should now be able to take the same input and produce the same DOM nodes. In practice I've not used a variety of such libraries, nor have I fed them very much pathological input, so I can't vouch for if this 100% true in practice, but in theory, there should no longer be any significant differences between HTML parsers in various languages, as they come on board with HTML5 compliance.
- anonacct37 9y agoI've used it to parse alot of broken HTML for a very large company. I have not used it to attempt to parse all the HTML on the web. My impression is that it's pretty good.