13 ms·
What one may find in robots.txt
- look_lookatme 11y agoI like: [...] User-agent: nsa Disallow: / From slack.com/robots.txt
- psykovsky 11y agoI like to put a Disallow rule to a randomly named directory with an index.php file that blocks any IP that accesses it. Then, for a bit of added fun, I put another Disallow rule to a directory named "spamtrap" which does an .htaccess redirection to the block script on the randomly named directory.
- kjjw 11y agoWonderful. So your users may occasionally be blocked due to you having a bit of fun. IP blocking is a terribly over broad way of stopping intruders.
- psykovsky 11y agoIf my users want to go where I ask them not to go when they have no reason at all to go there, their problem. There are no links anywhere to those dirs, except for the robots.txt. Also, the blocks lift after a couple days. Honestly, nobody ever complained, and I'm sure it stops/hampers some attacks dead on.
- kjjw 11y agoUsers share IPs.
- psykovsky 11y agoThe websites I manage aren't Facebook or Google sized. Or even HN sized. I don't see that as a real problem at all.
- castell 11y agoYou see it not as problem, because user don't see you (your sites). As IPv4 are getting rare, many share the same IP. So it's really a bad practice to ban IP for a longer period.
- malka 11y agowith bots ? If an Ip does not respect the host rules, it deserves to be blocked.
- germanier 11y agoFor a start, the entire nation of Qatar shares 82.148.97.69.
- kjjw 11y agoWhat do you mean 'with bots'? We're talking about anyone at all hitting a link blocking the IP they have. Some mobile networks have every person on the network originating traffic from the same IP. Some large institutions, universities, government departments, large companies have all their traffic coming from one IP. This person has effectively created a feature that will perform a denial of service attack on their own website.
- AgentME 11y agorobots.txt tells robots where not to go.
- comeonnow 11y agoI don't understand why a user would access a randomly created directory that is only mentioned in the robots.txt file. Can you explain as to why you'd think they'd stumble across this?
- gpvos 11y agoCuriosity?
- psykovsky 11y agoCuriosity killed the cat, or so they say.
- beaumartinez 11y ago"But satisfaction brought him back."
- gpvos 11y agoAs long as the block lasts not more than a few days, as he says in a nearby comment, I don't think it's much of a problem.
- psykovsky 11y agoOh, believe me they only last a few days at most. I don't use an infinite IP blocklist. It has like 100 IP's or so. When one goes in, one must come out. And lets just say there are enough badly behaved bots around that it doesn't take much time for 100 IP's to rotate.
- protomyth 11y agoThis is silly. If you have a user of your site going through your bots file and specifically going to directories listed as Disallow then you deal with that user. Blocking based on the robots.txt for a directory that doesn't exist anywhere but that file is fine. I did a two bad directory ban, it seemed to work fine.
- 11y ago
- asddubs 11y agoYou understand that all you are doing with that is giving people possibly looking for attack vectors an attack vector, right? All an evildoer has to do is embed the blacklist directory as an image somewhere, send it to someone they want to lock out of the service, etc.
- psykovsky 11y agoWhile you have a very valid point, in my use case I just ignore it. I don't see why anyone would want to block anyone else from accessing a small independent label website. Or the personal blog of a friend of mine. Or the portfolio site of another friend who is a designer. (edit)Or what real harm could come from it if it happened.
- njharman 11y agoI agree with you (on risk is acceptable) but I also believe you underestimate the pettiness of people. I'm sure there are people who believe the ind label or your friends are enemies (for whatever real or imagined reason). But the set of those people who also have technical ability or access to technical ability to enact this "revenge" is almost certainly empty.
- laumars 11y agoIt wouldn't be too hard to prevent that. eg the honeypot directory name could be a hash of the originating IP. You'd need to then have a dynamic robots.txt but that's easily done. The destination directory doesn't even need to exist. Worst case scenario you could handly the hash via your 404 handler or via .htaccess file if all of your hashes are prefixed. Those are only examples though - there's a multidude of ways you could handle the incoming request.
- egeozcan 11y agoAnd then our hypothetical attacker can figure out how you generate the honeypot URL and embed an image with that URL for their visitors. Of course you can easily make it impossible to guess but you need to make sure that they can't obtain robots.txt via GET requests from a visitors browser (No Access-Control-Allow headers). Also don't forget a single visit from a network, sharing an IP would still ban all the network. And there is the rule that a GET request should not change any state. Banning an IP address is a change of state. In short, it's just too much pain for little gain. Maybe a login form can be served from that URL, and any attempts to login would then get the visitor banned via a session cookie / browser fingerprint combo (Easy to get around but at least then you're not blocking IP addresses).
- liviu 11y agoWhat if I put a link to this URL and google bot will follow? You will block google bot?
- Ded7xSEoPKYNsDd 11y agoGooglebot won't follow the link because it's listed in robots.txt.
- ceequof 11y agoGooglebot caches robots.txt for a very, very long time. If you disallow a directory it may take months for the entire googlebot fleet to start ignoring it. Google's official stance is that you should manage disallow directives through webmaster tools.
- ryanlol 11y agoYes, but it will index it.
- lucb1e 11y agoDitto. My favorite name is /porn and then see who visits it. Mostly bots, though.
- Gigablah 11y ago> "At worse, only the internet etiquette has been breached." Proceeds to announce the name of a stalking victim. Classy.
- camillomiller 11y agoIt'a name that's in a plain text document that anybody can personally look up on its browser. What's the point of hiding it, if it's a very good point for his article?
- werid 11y agoNow it's indexed by search engines?
- icebraining 11y agohttps://www.google.com/search?q=inurl%3Arobots.txt https://www.google.com/search?q=inurl%3Arobots.txt Not that I agree it should be further divulged, mind you.
- MatthewWilkes 11y agoHe could have made exactly the same point with a fake name and institution. The information became public due to incompetence, he's made it much more visible… I don't care to speculate on what personal failing might be the reason.
- rlidwka 11y agoIt won't be exactly the same point. With a fake name there is no proof that the information in question ever existed. I would've probably masked a name, but institution url should stay there, so anyone could check that the point is valid.
- MatthewWilkes 11y agoOkay, that's fair, but it would be the same point for practical purposes, in my opinion. I don't personally believe that it's necessary for people to get hand-holding to independently verify his individual claims. Certainly other people could copy his methods to reproduce the results, but I'm not sure what the benefit is of links into the individual leaks of sensitive information.
- noobie 11y agoWhat's with the domain name? it says thiébaud.fr but when copied/pasted it becomes xn--thibaud-dya.fr Edit: Thanks for the links! :)
- talideon 11y agoIt's an IDN: http://en.wikipedia.org/wiki/Internationalized_domain_name http://en.wikipedia.org/wiki/Internationalized_domain_name
- learnstats2 11y agoThat's how internationalized domains work. Wikipedia: https://en.wikipedia.org/wiki/Uniform_Resource_Locator#Internationalized_URL https://en.wikipedia.org/wiki/Uniform_Resource_Locator#Inter...
- mrwizrd 11y agoThe domain name is being converted to PunyCode (https://en.wikipedia.org/wiki/Punycode https://en.wikipedia.org/wiki/Punycode). This is a defence developed to prevent spoofing of domain names using indistinguishable characters from other character sets than ASCII. Before this was fixed by browser developers, it was possible to execute an IDN homograph attack (https://en.wikipedia.org/wiki/IDN_homograph_attack https://en.wikipedia.org/wiki/IDN_homograph_attack) and spoof sites using characters that looked like e.g., google.com/paypal.com when they rendered in your browser address bar, but were actually a completely different domain as far as automated systems were concerned. Edit: typo.
- Flimm 11y agoInterestingly, Hacker News doesn't support this standard, and the link on the front page shows the unfriendly version.
- ozh 11y agoIt's because IDN domains became a standard after tables stopped to be used in page layouts </sarcasm>
- _lce0 11y agoI wonder how could you gather a huge _domain name_ list. I guess using DNS, or by query some engine, like google, or archive.org is there a service somewhere?
- anc84 11y ago1 million domains ranked by Alexa: http://s3.amazonaws.com/alexa-static/top-1m.csv.zip http://s3.amazonaws.com/alexa-static/top-1m.csv.zip For more try https://commoncrawl.org/ https://commoncrawl.org/
- zuzun 11y agoYou can apply for Zone File access at some Internet registries.
- dredmorbius 11y agoOr ... you could use the methods listed in the article.
- briandh 11y agoMaybe the DNS ANY or reverse DNS datasets on scans.io? (the former, I presume, ultimately covering more domain names but containing more extraneous information)
- xyzzy123 11y agoSee also: https://www.premiumdrops.com/zones.html https://www.premiumdrops.com/zones.html They have some pretty good zone files for major TLDs.
- roel_v 11y agoI needed some real-world MS Word & Excel documents years ago to test some parsing code against. So I started crawling Google results for 'filetype:.doc .xls' style queries. Left it running for a weekend, then the whole Monday was wasted as I was sucked into looking through the results - some stuff in there was certainly not meant for public disclosure...
- Ntrails 11y agoGreat. There goes any chance of me actually doing work this week. ;p
- elorant 11y agoPardon me for asking, but how did you crawl Google for a whole weekend. From what I know, Google blocks you if you request too many search queries in a short period of time. Did you use proxies?
- roel_v 11y agoMaybe now, I don't remember I had do anything sneaky back then (8-10 years ago)
- wanderingstan 11y agoIn 2004 I attempted to do some automatic crawling of Google for my masters thesis and was astonished to get an unfriendly server response saying it was disallowed and "don't even bother asking for an exception for research, it won't be granted." So at least 11 years ago it was blocked. (I didn't know about spoofing a user agent back then, so it might not have been as easy as that to get around it.)
- frik 11y agoDid you really crawl Google? That has to be a long time ago. But speaking about searching on Google as a user: Google's Advanced search used to be a great tool, until around 2007/08. For some reason it never received an upgrade and several things are broken or don't work any more or were removed (e.g. '+' which is now a keyword for Google+, the '"' does mean the same; e.g some filetypes are blocked, some show only a few results).
- ar-jan 11y agoHm, the article several times says "required not to be indexed", while robots.txt is more like "request not to be crawled". An important distinction, because a page may well be indexed without being crawled, typically when there is a link to it. Better to use the noindex meta tag (at least if you are only concerned with search indexes, not access control).
- polaco 11y agoYes. This need to be emphasized: Disallowing URLs in robots.txt will not necessarily exclude them from search results. Search engines will still find those pages if they are linked to or mentioned somewhere else. The search result will consist only of the URL, and the snippet will say "A description for this result is not available because of this site's robots.txt" Use the noindex tag, folks. Also, Google Webmaster Tools allows you to remove URLs from Google's index. Ref.: https://yoast.com/prevent-site-being-indexed/ https://yoast.com/prevent-site-being-indexed/
- nickhalfasleep 11y agoUser-agent: * Allow: / # A robot may not injure a human being or through inaction allow a human being to come to harm. # A robot must obey the orders given it by human beings, except where such orders would conflict with the First Law # A robot must protect its own existence, as long as such protection does not conflict with the First or Second Laws.
- grandfish456 11y agoRegarding the Knesset website, it is actually just a boring recordings of the parliament discussions. Nothing to see here, move along... :-)
- wtbob 11y ago> xn--thibaud-dya.fr Looks like HN needs to learn how to decode Punycode…
- neil_s 11y agoWow, all the US Department of State files have just gone missing from archive.org. The servers hosting those files are conveniently down.
- benyami 11y agoCheck out the Internet Archive FAQ on how to remove a document from their archives. https://archive.org/about/exclude.php https://archive.org/about/exclude.php It looks like they used robots.txt to do that.
- neil_s 11y agoHuh, so the wild-card user-agent will block not just searchbots, but also archivebots. Wonder how OP managed to get screenshots of archive.org having archives available for those documents.
- maxmcd 11y agoI have been able to view multiple pdfs and view the page screenshotted by the author.
- kjell 11y agoThey're there, at least the two I looked at. https://web.archive.org/web/20130413152316/http://www.state.gov/documents/organization/193775.pdf https://web.archive.org/web/20130413152316/http://www.state.... Each line is missing `/documents` in the snippet of the `robots.txt`
- Scoundreller 11y agoNow if someone could do an analysis of humans.txt, that would be cool.
- jenscow 11y agoYou'd have to get a bot to do that, since us human's can't access the content it points to.
- philip1209 11y agoI find this submission of this article interesting because it underscores inconsistent handling of I18N/punycode domains. The domain is "thiébaud.fr". Should submission sites (like HN) show the sites in ASCII? Is there a fraud risk? Should the web browser show the domain in ASCII? For me - at no point was I shown the domain decoded to ASCII (either on HN or in the browser). I recognized the pattern and decoded it manually. For users who are not technical - this is a failed experience because the domain looks suspicious and at no point was it decoded. I wonder when punycode decoding will begin to get attention from developers. Last year's Google IO had a great talk about how Google realized the inconsistency of their domain handling with regard to I18N: https://www.google.com/events/io/schedule/session/22ce27dc-7cbf-e311-b297-00155d5066d7 https://www.google.com/events/io/schedule/session/22ce27dc-7...
- aidenn0 11y agoI'm on firefox, and it showed me the decoded domain.
- tlb 11y agoIt's a vector for putting in disruptive utf-8 characters, such as a huge stack of accents, or spoofing a reputable domain. It's not clear yet that the benefits to HN outweigh the risks. But if we start seeing a lot of quality content from domains that look better with punycode decoded, it'll be considered.
- ErikDub 11y agoIt's actually quite different how all the browsers handle the Punycode domains in different places: http://blog.dubbelboer.com/2015/05/10/unicode-domain-support.html http://blog.dubbelboer.com/2015/05/10/unicode-domain-support...