2 ms·
1000% percent this. I write about Python web scraping a lot and the big one is that there's two parts. First is gathering the pages you need to scrape locally,
by jackschultz 9y ago
1000% percent this. I write about Python web scraping a lot and the big one is that there's two parts. First is gathering the pages you need to scrape locally, and the second is scraping the pages you've saved. You need need to separate those two to avoid hitting their servers over and over when you're tying to debug the scraping code. My way is to write the first in a file called gather.py, and then the other in scrape.py. Have that in mind before doing and heavy scraping.
Since my scrapes aren't always the biggest, feel free to just save the html in a local folder and then scrape from there. This part of the project depends on how many pages you need to scrape, the size of the files, whether you need to store the data, whether it's a one time scrape or croned, etc. Either way, save the files, and scrape from there.
Here are some of the posts I've done on the subject if people reading the comments want to see more about scraping.
- https://bigishdata.com/2017/05/11/general-tips-for-web-scraping-with-python/ https://bigishdata.com/2017/05/11/general-tips-for-web-scrap...
- https://bigishdata.com/2017/06/06/web-scraping-with-python-part-two-library-overview-of-requests-urllib2-beautifulsoup-lxml-scrapy-and-more/ https://bigishdata.com/2017/06/06/web-scraping-with-python-p...
- https://bigishdata.com/2017/05/11/general-tips-for-web-scraping-with-python/ https://bigishdata.com/2017/05/11/general-tips-for-web-scrap...