3 ms·
I've been considering writing my own puppeteer docker image such that one could freeze the image at crawl time after a page has loaded. This would allow me to r
by princehonest 9y ago
I've been considering writing my own puppeteer docker image such that one could freeze the image at crawl time after a page has loaded. This would allow me to re-write the page-parsing logic after the page layout changes. Has anyone done this already or know of any other efforts to serialize the puppeteer page object to handle parsing bugs?
- radioo75555 9y agoIn large scale scraping they always separate loading pages from processing the data. Easiest thing would be to wait for the page to load and then save the html