3 ms·
If you are scraping 1000 pages and your program crashes after the 800th page, you are going to be much happier if you have the data for the 800 pages saved some
by MaxDPS 5y ago
If you are scraping 1000 pages and your program crashes after the 800th page, you are going to be much happier if you have the data for the 800 pages saved somewhere vs having to start all over.
- cushychicken 5y agoGot it. Yeah, I can give this approach a shot. One of the things I've been doing with this scraping project is "health checking" job posting links. There's nothing more annoying than clicking an interesting looking link on a job site, only to find it's been filled. (This is one of the lousiest parts of Indeed and similar, IMO.) I wrote some pretty simple routines Caching solves the problem of potentially missing data while the scraper is running, but it doesn't really alleviate the network strain of requesting pages to see that they are actually still posted, valid job links.