4 ms·
If it changes the format significantly, the scraper will break, so for now you'll have to use the tool to rebuild. You will see on your API status page that it'
by pranade 13y ago
If it changes the format significantly, the scraper will break, so for now you'll have to use the tool to rebuild. You will see on your API status page that it's down. As for robots.txt, we do respect it... for now we're leaving that to the user, but we're trying to implement a proactive way of checking for disallows and stopping those scrapers from being built.
- thatthatis 13y agoPlease clarify: are you saying that right now you leave respecting robots.txt to the user?
- pranade 13y agoAt the moment, we rely on users to be responsible. We spell it out in the terms and FAQ. We've been in private beta, keeping usage very limited until today. We fully understand the seriousness of the issue as we scale. We're committed to becoming a responsible bot that respects robots.txt
- loceng 13y agoI would say how you're scraping differs from say how Google, a search engine, scrapes. I'm not sure there is a way in robots.txt to define for each use? Knowing the data in a structured way, but then allowing it to be displayed in full off-site is quite different than using the scraped data for linking into a website.
- thatthatis 13y agoBut robots.txt provides minimums: don't scrape this page, don't refresh more than once every x, these crawlers are allowed this access, etc.
- pranade 13y agoFeedback from webmasters is really helpful for us. We want to make sure we're making data available via API responsibly, so would love to hear your suggestions/ thoughts as we define a scalable solution.