2 ms·
For folks who do this kind of disparate data-source scraping at scale, what does best practices look like? What kind of tools are used in industry? Maintaining
by celestialcheese 4y ago
For folks who do this kind of disparate data-source scraping at scale, what does best practices look like? What kind of tools are used in industry?
Maintaining scrapers for 18k county websites and PDs is no small task and looking through the docs for PDAP, it seems like this is still a very open question.
- electromech 4y agoOur World In Data is the largest open source data collection & analysis that I'm aware of. https://github.com/owid https://github.com/owid The 80000 Hours podcast has an interview with the (non-technical) creator of OWID. I seem to recall some interesting stories about them getting emailed PDFs with COVID data and such. I had the same question as you, and I was hoping to find ideas in the comments. It seems like the kind of thing that's both inherently messy and scrappy yet if you don't get at least somewhat organized it can't scale. Update: link to the podcast episode page with quotes, transcripts, etc. https://80000hours.org/podcast/episodes/max-roser-our-world-in-data/ https://80000hours.org/podcast/episodes/max-roser-our-world-...
- celestialcheese 4y agoThanks for sharing! It's interesting that even one of the largest still uses manual execution for almost all of their pipelines (at least in the covid data project[1]). This [2] seems like the bulk of their data importers (scrapers) but most are still operating as manual jobs. I guess with open-source data work, hours and minutes don't matter as much, and being a few days behind the latest data is acceptable. 1 - https://docs.owid.io/projects/covid/en/latest/data-pipeline.html https://docs.owid.io/projects/covid/en/latest/data-pipeline.... 2 - https://github.com/owid/importers https://github.com/owid/importers