3 ms·
I really wouldn't recommend building a web scraper from scratch. You'll soon have to think about caching/rate-limiting/retries. Personally, I use Scrapy and it
by guskel 5y ago
I really wouldn't recommend building a web scraper from scratch. You'll soon have to think about caching/rate-limiting/retries.
Personally, I use Scrapy and it works fine. For best practice, I wouldn't use the Pipeline concept Scrapy provides -don't do data transformation inside scrapy. Simply save the responses and perform the validation and transformations outside of Scrapy. The Pipeline concept is flawed because you cannot create DAGs with it -only serially linked pipelines.