5 ms·
Yup this is true - I was just being lazy. But what surprised me was that Bing actually indexed them. (even though my robots.txt said not to)
by rsbadger 4y ago
Yup this is true - I was just being lazy. But what surprised me was that Bing actually indexed them. (even though my robots.txt said not to)
- bigtones 4y agoCan you prove that by linking to a Bing search where one of your pages show up ?
- rsbadger 4y agoyeh... https://www.bing.com/search?q=https%3A%2F%2Fshoprocket.io%2Femail-confirmation%2F&qs=n&form=QBRE&sp=-1&pq=shoprocket.io%2Femail-verification&sc=1-32&sk=&cvid=40B93AB2ADAA459CA2BB788274F18028&ghsh=0&ghacc=0 https://www.bing.com/search?q=https%3A%2F%2Fshoprocket.io%2F...
- scandinavian 4y agoThe URL in the search is this: https://shoprocket.io/email-confirmation/34b35b1 https://shoprocket.io/email-confirmation/34b35b1... I don't see that in the robots.txt https://shoprocket.io/robots.txt https://shoprocket.io/robots.txt User-agent: * Disallow: /cdn-cgi/l/email-protection Disallow: /login Disallow: /register Disallow: /404 Am I missing something?
- rsbadger 4y agoYou may be seeing a stale version, try this: https://shoprocket.io/robots.txt?bypass=1 https://shoprocket.io/robots.txt?bypass=1 (I made a lot of changes today when testing all, including "visit as Bingbot" from their webmaster tools with and without the URL blocked by robots.txt)
- logifail 4y ago> was that Bing actually indexed them. (even though my robots.txt said not to) Never mind indexing them (ie publishing them at Bing.com), if URLs are disallowed in robots.txt then Bing shouldn't even be retrieving them, even if only to scan the content for malware!
- snowwrestler 4y agoThis is a common misconception about robots.txt. It tells bots what they should do while directly crawling your site. But if a search engine gets to a URL some other way—for example if it follows a link from somewhere outside your site—it will still index that page. Robots.txt is not a reliable way to exclude pages from search engine indexes. That is not what it is for. It is for controlling crawler behavior. The only reliable way to exclude a URL from a search engine index is to serve “noindex” on that URL, either with a metatag or an HTTP header, or both.
- rsbadger 4y agoThis is very useful information. You’d really hope that private emails would be excluded by default…
- logifail 4y ago> It tells bots what they should do while directly crawling your site. But if a search engine gets to a URL some other way—for example if it follows a link from somewhere outside your site—it will still index that page. I must confess I've been sceptial of robots.txt for a very long time (if I want to stop bots I serve them HTTP 403 Forbidden using .htaccess or similar). Be that as it may, it appears I'm also confused about what robots.txt does and doesn't do. Assuming you're correct: let's say I run EvilBot which scrapes sites and want to scrape your site example.com, but your robots.txt only allows Googlebot and disallows everyone else. Am I really OK to: 1. scrape the SERPs from google.com which mention your site ("site:example.com") then 2. using that list of URIs, use my EvilBot to scrape your site, without needing to touch or respect your robots.txt, since I got the list of URIs on your site from Google, not by scraping example.com directly?
- snowwrestler 4y agoYour step 1 is enough for URLs to be indexed. Even a well-behaved search engine does not need to visit your site to index a URL, including whatever anchor text pointed at it. If the crawler does then visit your site, it will see your robots.txt and (if well-behaved) obey it and not crawl the contents of the page at that URL. But this does not mean it will remove the URL itself from its index. Again: robots.txt is intended to control crawler behavior, not search index visibility. Google's page is a pretty good overview of this distinction: https://developers.google.com/search/docs/advanced/robots/intro https://developers.google.com/search/docs/advanced/robots/in...