4 ms·
I bet one of the leading factors is the number of sites blocking the AWS servers. As a scraping service, the fact that Amazon makes all their public IPs known (
by ecaron 13y ago
I bet one of the leading factors is the number of sites blocking the AWS servers. As a scraping service, the fact that Amazon makes all their public IPs known (https://forums.aws.amazon.com/ann.jspa?annID=1701 https://forums.aws.amazon.com/ann.jspa?annID=1701) is really inconvenient for anyone "crawling" the web. Rackspace also makes their list available, albeit incomplete (http://www.rackspace.com/knowledge_center/article/cloud-sites-outbound-ip-addresses http://www.rackspace.com/knowledge_center/article/cloud-site...).
As more websites setup default blocks for all accesses from AWS services, the Moz exit was inevitable.
- randfish 13y agoWe've actually never crawled out of Amazon (mostly because it was expensive to do so), so the crawl blocking stuff is unrelated, at least for us. As Sarah noted, it's really been costs, service, and support issues.
- briancurtin 13y ago> Rackspace also makes their list available, albeit incomplete (http://www.rackspace.com/knowledge_center/article/cloud-site... http://www.rackspace.com/knowledge_center/article/cloud-site...). FWIW, this is specific to Cloud Sites, which is like one-click Wordpress/Drupal deployments.
- hhw 13y agoYou can just use BGP to find out all that information easily enough for any network with an AS number (i.e. any who speaks BGP). Just take a look here: http://bgp.he.net http://bgp.he.net or a plethora of looking glasses like http://lg.level3.net http://lg.level3.net or publicly available route servers like telnet://route-views.routeviews.org