3 ms·
Having done a lot of web scraping, I would say there are a few rules: 1) Follow any Terms of Service or Terms of Use. 2) Respect the site's robots.txt 3) Whe
by cblock811 11y ago
Having done a lot of web scraping, I would say there are a few rules:
1) Follow any Terms of Service or Terms of Use.
2) Respect the site's robots.txt
3) When in doubt...don't do it.
What is the purpose of your scraping? I found that if I was doing analysis, just asking people will prevent any issues. Half the time they just gave me the information I wanted.
- rebootthesystem 11y agoThe purpose is to enhance data available through API's. Every single bit of of it is available publicly without login. This is not to repost articles or data but rather to provide better actionable information. If you look at the article I linked to the common thread seems to be abusive use of scrapped data for direct financial gain. It seems TOS is irrelevant unless the data scrapped was behind a login and, even then, it depends on the TOS wording. Clearly it is a minefield u der certain circumstances. Definitely need to do more research.
- rebootthesystem 11y agoAlso, there seems to be a huge difference between scraping 100,000 pages per hour vs. one page every few seconds. The kind of scraping I am considering would be the latter and, even at that, it would only happen a few days per month.