3 ms·
It is worth emphasizing that what the author is talking about is more akin to screen scraping than web crawling. Both tasks have their challenges, but screen sc
by kakaylor 16y ago
It is worth emphasizing that what the author is talking about is more akin to screen scraping than web crawling. Both tasks have their challenges, but screen scraping has several that are inherently difficult to overcome.
In particular, with screen scraping, you are trying to extract structured data from a markup language (in this case HTML) that simply doesn't guarantee the structure your looking for. With web crawling you only need the structural guarantees offered by the HTML markup (not even that, with the quality of libraries such as TagSoup or Neko).
Now, that isn't to say web crawling doesn't have its own challenges (URL canonicalization anyone?).