3 ms·
Maybe I am missing something here, but why would you run "AI" on that task when you go from formal language to formal language? I don't get the usage of "regex
by choeger 2y ago
Maybe I am missing something here, but why would you run "AI" on that task when you go from formal language to formal language?
I don't get the usage of "regex/heuristics" either. Why can that task not be completely handled by a classical algorithm?
Is it about the removal of non-content parts?
- nickpsecurity 2y agoIt’s informal language that has formal language mixed in. The informal parts determine how the final document should look. So, a simple formal-to-formal translation won’t meet their needs.
- baq 2y agoThere’s html and then there’s… html. A nicely formatted subset of html is very different from a dom tag soup that is more or less the default nowadays.
- JimDabell 2y agoTag soup hasn’t been a problem for years. The HTML 5 specification goes into a lot more detail than previous specifications when it comes to parsing malformed markup and browsers follow it. So no matter the quality of the markup, if you throw it at any HTML 5 implementation, you will get the same consistent, unambiguous DOM structure.
- mithametacs 2y agoyeah, you could just pull the parser out of any open source browser and voila a parser not only battle-tested, but probably the one the page was developed against
- faangguyindia 2y agoThat's why the best strategy is to feed the whole page into LLM. (After removing html tags) and just ask LLM to give you the date you need in the format you need. If there is lots of javascript dom manipulation happening after pageload. Then just render in webdriver and screenshot, ocr and feed the result into LLM and ask it the right questions.
- mithametacs 2y agoMy intuition is that you’d get better results emptying the tags or replacing them with some other delimiter. Keep the structural hint, remove the noise.