2 ms·
How do scrapers deal with randomized classes in web pages which is more and more common these days? Relying on the page structure only is not a robust alternat
by dmortin 4y ago
How do scrapers deal with randomized classes in web pages which is more and more common these days?
Relying on the page structure only is not a robust alternative.
- edmundsauto 4y agoI've had some success with running a meta-scraper that will search for known value on a page, then back out the page structure from there. It won't help with randomly generated class names, but 95% of tasks I've written aren't this complex. For sites that are hard to scrape (usually bigger sites that get scraped a lot), I pivot towards buying a data feed. Economies of scale incentivize these data companies towards putting someone on maintaining the feed full-time.