3 ms·
It seems sort of questionable to use the list of things to not scrape as a starting point for scraping.... I mean, I get it's not actually enforced.
by calebegg 3y ago
It seems sort of questionable to use the list of things to not scrape as a starting point for scraping.... I mean, I get it's not actually enforced.
- das_keyboard 3y agoNot really sure why all the answers here are flagged, but you may be mistaken. The robots.txt does not exclusively list what not to scrape. It provides information on which parts are allowed and wich are not (disallowed). It also provides sitemaps for crawlers as a starting point with more information (eg. which sites are available and how often are they updated, etc.)
- xnx 3y agoSince ~2009 many crawlers recognize "Sitemap:" directives in robots.txt to link to sitemaps: https://en.wikipedia.org/wiki/Robots.txt#Sitemap https://en.wikipedia.org/wiki/Robots.txt#Sitemap
- fdsajfsldkj 3y ago[flagged]
- qefvss 3y ago[flagged]