5 ms·
One precursor question would be whether an LLM can extract the data you want from raw html even when copy-pasted manually. In my limited experience we’re not q
by coderatlarge 3y ago
One precursor question would be whether an LLM can extract the data you want from raw html even when copy-pasted manually. In my limited experience we’re not quite there yet, but I’d be curious to hear of others have different experience - or better yet, actual measurements against a baseline scraper.
- FrenchDevRemote 3y agoOh it definitely can. I tried with GPT-4 and Anthropic(claude 2), it was relatively good, although I had to tweak the prompts to get everything I wanted(sometimes it forgot images or urls, because my prompt was too vague) You could also write a simple function that turn HTML into a json that only contained innerText, src and hrefs, and ask the LLM to only keep the relevant data