10 ms·
Google's robots.txt
- catmanjan 13y agoWhat a weird entry... https://www.google.com/maps?hq=http://maps.google.com/help/maps/directions/biking/mapleft.kml&ie=UTF8&ll=37.687624,-122.319717&spn=0.346132,0.727158&z=11&lci=bike&dirflg=b&f=d https://www.google.com/maps?hq=http://maps.google.com/help/m...
- hayksaakian 13y agois it a map focused on SF that highlights bikings paths?
- catmanjan 13y agoLooks like it? Very weird
- ryanpetrich 13y agoSee also: http://www.google.com/humans.txt http://www.google.com/humans.txt
- khyh 13y agoHow did you find this?
- Houshalter 13y agoThis doesn't work if you are using HTTPS everywhere.
- dfc 13y agoWeird, it works for me. Iceweasel Aurora / HTTPS Everywhere 4.0-dev
- daGrevis 13y agoI second this. Latest Chromium.
- ushi 13y ago% curl www.google.com/humans.txt Google is built by a large team of engineers, designers, researchers, robots, and others in many different sites across the globe. It is updated continuously, and built with more tools and technologies than we can shake a stick at. If you'd like to help us out, see google.com/jobs.
- jemfinch 13y agoFixed.
- deleted 13y ago[deleted]
- yRetsyM 13y agoI remember when this used to be a source of product leaks.
- zeckalpha 13y agoAnd now it is an archive of discontinued products.
- yRetsyM 13y agoI remember when this used to be a source of product leaks.
- blossoms 13y agoUnrelated but it looks like www.aol.com's robots.txt is served as text/html http://www.aol.com/robots.txt http://www.aol.com/robots.txt Is this a common mistake?
- mkonecny 13y agoIt just issues a "HTTP/1.1 302 Moved Temporarily" directed to their homepage. Requesting an invalid file such as "robots.txtsdfa32r523" has the same effect, so they probably don't have a robots file at all.
- judk 13y agoHuh? No, it is a regular robots.txt file
- tonyedgecombe 13y agoIt redirects requests from the UK.
- blossoms 13y agohttp://i.imgur.com/Jmjf3y2.png http://i.imgur.com/Jmjf3y2.png
- saltysugar 13y agoYes, from the server side's mishandling of TXT extension. Probably the server put the MIME type in the HTTP header as "HTML" instead of TXT, and the browser renders the page as such.
- pjscott 13y agoFrom the headers: Content-Type: text/html;charset=UTF-8 It also tries to set no less than four cookies.
- logotype 13y agoMine is more awesome: http://logotype.se/robots.txt http://logotype.se/robots.txt
- russellbeattie 13y agoHeh, if only Yandex and Baidu respected robots.txt.
- agwa 13y agoI've found this GitHub project to be an invaluable resource for blocking bad bots: https://github.com/bluedragonz/bad-bot-blocker https://github.com/bluedragonz/bad-bot-blocker
- bmelton 13y agoThat's amazing. Know of anything like it for Nginx?
- anon4 13y agoIt's blocking by user agent and source ip. You should be able to port the list easily for Nginx. I'd even say you can write a simple awk script in a few minutes to convert from Apache's format to Nginx's.
- coolj 13y agocurl -s https://raw.github.com/bluedragonz/bad-bot-blocker/master/.htaccess https://raw.github.com/bluedragonz/bad-bot-blocker/master/.h... | awk -F\" '/SetEnvIfNoCase/ {pattern = pattern $2 "|"} END {print "if ($http_user_agent ~ (" substr(pattern, 0, length(pattern)-1) ")) { return 403; }"}' Not tested. :)
- anonymfus 13y ago>Unless your website is written in Russian or Chinese, you probably don't get any traffic from them. They mostly just waste bandwidth and consume resources. THIS is evil. You could use this argument for banning any new search engine.
- oevi 13y agoWikipedia's robots.txt is quite verbose: http://en.wikipedia.org/robots.txt http://en.wikipedia.org/robots.txt
- iso8859-1 13y agoBecause it's part of the Wiki: https://en.wikipedia.org/w/index.php?title=MediaWiki:Robots.txt&action=history https://en.wikipedia.org/w/index.php?title=MediaWiki:Robots....
- MattHeard 13y agoMy favourite part: > # Folks get annoyed when XfD discussions end up the number 1 google hit for > # their name.
- 0_o 13y agotaobao.com has the shortest robots.txt http://www.taobao.com/robots.txt http://www.taobao.com/robots.txt
- cordite 13y agoA lot of these are 404's. It kinda makes me wonder what was behind things like /c/
- wupiass 13y agofacebook's... https://www.facebook.com/robots.txt https://www.facebook.com/robots.txt
- yeukhon 13y ago> Notice: Crawling Facebook is prohibited unless you have express written Wow, really? Who put up this sign?
- phinnaeus 13y agoSomeone with 80 character lines enforced in their editor? Here's the second line... > permission. See: http://www.facebook.com/apps/site_scraping_tos_terms.php http://www.facebook.com/apps/site_scraping_tos_terms.php
- Houshalter 13y agoThat's not even a text file.
- amaks 13y agoBing is allowed. But it's obvious why.
- aw3c2 13y agoDo you get a different file than me? No-one but ia_archiver is allowed, see https://pastee.org/zpjsa https://pastee.org/zpjsa
- monkeyspaw 13y agoI frequently wonder - is Facebook allowed to say, "bing, you can crawl us. NewCompetitor, you cannot." I feel like once a company allows public access by posting stuff on the web, they can specify terms, but not include/exclude groups specifically. (In a legal sense; I understand blocking systems that hammer servers but will respect robots.txt. IME bing is the worst offender -- they hammer my sites, send no traffic, but will stop if I specify in robots.txt.) Does anyone have an opinion about "once public, I can crawl"?
- d99kris 13y agoGlassdoor uses its for recruiting: http://www.glassdoor.com/robots.txt http://www.glassdoor.com/robots.txt
- te_chris 13y agoWhy would they disallow all their about content? (Bit of an SEO noob).
- JimmyM 13y agoOver-simply, because they do not want to indicate that their site is about those pages, for whatever reason. In practice, I'm not entirely sure but it looks like it's quite old since those pages don't seem to exist anymore, even as folders which don't exist as pages in themselves, and they aren't 301-redirected to the current relevant pages. In fact, they're all 404's. So perhaps they used to be pages, were deleted, and kept being crawled which made their site look crap (because of the 404's). Now they could use 301's, but I assume that the reason they didn't is because they might want to restructure the site in the future and re-use those pages. They don't use 302's because 302's are unreliable and freaky. Does that sound right to everyone else?
- mattmanser 13y agoThese guys are the best white-hat SEO growth hackers on the planet, just copy them and don't question it.
- JimmyM 13y agoOh, I wasn't questioning them - just asking if my assessment sounded right.
- coin 13y agoSadly they use Jobvite. Its usability is about as bad as Lotus Notes.
- avar 13y ago
- Theodores 13y agoThere are only 602 pages of Google.com indexed on Google.com, mostly 'plus' profiles. Quite a few show with this message: A description for this result is not available because of this site's robots.txt – learn more. Which is odd.
- eli 13y agoGoogle will still index pages blocked by robots.txt, it just won't crawl them (so it can't get a description/preview snippet). It indexes them based on the URL and how people link to them.
- techaddict009 13y agoCheckout youtube.com/robots.txt !
- dabit 13y agoFor the lazy: http://www.youtube.com/robots.txt http://www.youtube.com/robots.txt
- Mindless2112 13y agoAnd for those who don't know the reference: http://www.youtube.com/watch?v=WGoi1MSGu64 http://www.youtube.com/watch?v=WGoi1MSGu64
- plucas 13y agoYelp's is fun: https://www.yelp.com/robots.txt https://www.yelp.com/robots.txt
- ankurpatel 13y agoYelp has rules set for the robots: http://www.yelp.com/robots.txt http://www.yelp.com/robots.txt
- gergles 13y agohttp://www.google.com/baraza/en http://www.google.com/baraza/en What a weird little product. It's like Yahoo Answers, but somehow with even less sorting or categorization.
- gillis 13y agohttp://www.google.com/baraza/en/help?file=inboundsms http://www.google.com/baraza/en/help?file=inboundsms They have an SMS companion service that only works in Ghana..
- lesiki 13y ago'baraza' is Swahili for forum/meeting place. It was very much Yahoo Answers, targeted at the African market - here in Kenya, most people only have internet access via mobile connections, often using feature phones, hence the minimalistic stylesheet. Baraza never really took off. More about it here: http://whiteafrican.com/2010/10/05/google-baraza-qa-for-africa/ http://whiteafrican.com/2010/10/05/google-baraza-qa-for-afri...
- MarkTee 13y agoAmazingly, the questions seem to be even dumber than the ones found on YA: - why is computer an idiot machine - i casted a love spell on my ex,should i tell her now that she is back - What is the colour of the black box which is using in planes ?
- codr 13y agoApple allows robots access to everything! http://www.apple.com/robots.txt http://www.apple.com/robots.txt
- oneeyedpigeon 13y agoIf you're big enough, doesn't this just make sense? Why waste time maintaining a robots.txt policy when it must represent a tiny fraction of your traffic, which your servers can surely handle? And the really 'bad' guys are going to ignore it anyway. And if you really care, you'll have some much more sophisticated bandwidth throttling in place. For the smaller guys, sure it makes sense to have some kind of simple robots.txt policy.
- keule 13y agoThats not quite true. Apple has multiple subdomains, each for some part of their site. www.apple.com hosts most of the marketing stuff, but have a look at other pages, e.g.: http://store.apple.com/robots.txt http://store.apple.com/robots.txt
- datawander 13y agoActually, robot traffic accounts for a much larger share of traffic for any of the big websites than you think.
- 0x0 13y agoCurious as to why someone sat down and added this line to that file: Allow: /maps?hq=http://maps.google.com/help/maps/directions/biking/mapleft.kml&ie=UTF8&ll=37.687624,-122.319717&spn=0.346132,0.727158&z=11&lci=bike&dirflg=b&f=d
- joliv 13y agoHere's where that link goes to: http://i.imgur.com/CmdBeQy.png http://i.imgur.com/CmdBeQy.png
- ZirconCode 13y agohttps://www.google.com/search?num=100&site=&source=hp&q=https%3A%2F%2Fwww.google.com%2Fmaps%3Fhq%3Dhttp%3A%2F%2Fmaps.google.com%2Fhelp%2Fmaps%2Fdirections%2Fbiking%2Fmapleft.kml%26ie%3DUTF8%26ll%3D37.687624%2C-122.319717%26spn%3D0.346132%2C0.727158%26z%3D11%26lci%3Dbike%26dirflg%3Db%26f%3Dd&oq=https%3A%2F%2Fwww.google.com%2Fmaps%3Fhq%3Dhttp%3A%2F%2Fmaps.google.com%2Fhelp%2Fmaps%2Fdirections%2Fbiking%2Fmapleft.kml%26ie%3DUTF8%26ll%3D37.687624%2C-122.319717%26spn%3D0.346132%2C0.727158%26z%3D11%26lci%3Dbike%26dirflg%3Db%26f%3Dd&gs_l=hp.3...6256.6256.0.6770.1.1.0.0.0.0.0.0..0.0....0...1c.1.35.hp..1.0.0.rCg1jO0B9q4 https://www.google.com/search?num=100&site=&source=hp&q=http... Turns up three results in google. Very weird indeed.
- Istof 13y ago404 for http://www.gstatic.com/trends/websites/sitemaps/sitemapindex.xml http://www.gstatic.com/trends/websites/sitemaps/sitemapindex...