11 ms·
I like to put a Disallow rule to a randomly named directory with an index.php file that blocks any IP that accesses it. Then, for a bit of added fun, I put anot
by psykovsky 11y ago
I like to put a Disallow rule to a randomly named directory with an index.php file that blocks any IP that accesses it.
Then, for a bit of added fun, I put another Disallow rule to a directory named "spamtrap" which does an .htaccess redirection to the block script on the randomly named directory.
- kjjw 11y agoWonderful. So your users may occasionally be blocked due to you having a bit of fun. IP blocking is a terribly over broad way of stopping intruders.
- psykovsky 11y agoIf my users want to go where I ask them not to go when they have no reason at all to go there, their problem. There are no links anywhere to those dirs, except for the robots.txt. Also, the blocks lift after a couple days. Honestly, nobody ever complained, and I'm sure it stops/hampers some attacks dead on.
- kjjw 11y agoUsers share IPs.
- psykovsky 11y agoThe websites I manage aren't Facebook or Google sized. Or even HN sized. I don't see that as a real problem at all.
- castell 11y agoYou see it not as problem, because user don't see you (your sites). As IPv4 are getting rare, many share the same IP. So it's really a bad practice to ban IP for a longer period.
- malka 11y agowith bots ? If an Ip does not respect the host rules, it deserves to be blocked.
- germanier 11y agoFor a start, the entire nation of Qatar shares 82.148.97.69.
- kjjw 11y agoWhat do you mean 'with bots'? We're talking about anyone at all hitting a link blocking the IP they have. Some mobile networks have every person on the network originating traffic from the same IP. Some large institutions, universities, government departments, large companies have all their traffic coming from one IP. This person has effectively created a feature that will perform a denial of service attack on their own website.
- AgentME 11y agorobots.txt tells robots where not to go.
- comeonnow 11y agoI don't understand why a user would access a randomly created directory that is only mentioned in the robots.txt file. Can you explain as to why you'd think they'd stumble across this?
- gpvos 11y agoCuriosity?
- psykovsky 11y agoCuriosity killed the cat, or so they say.
- beaumartinez 11y ago"But satisfaction brought him back."
- gpvos 11y agoAs long as the block lasts not more than a few days, as he says in a nearby comment, I don't think it's much of a problem.
- psykovsky 11y agoOh, believe me they only last a few days at most. I don't use an infinite IP blocklist. It has like 100 IP's or so. When one goes in, one must come out. And lets just say there are enough badly behaved bots around that it doesn't take much time for 100 IP's to rotate.
- protomyth 11y agoThis is silly. If you have a user of your site going through your bots file and specifically going to directories listed as Disallow then you deal with that user. Blocking based on the robots.txt for a directory that doesn't exist anywhere but that file is fine. I did a two bad directory ban, it seemed to work fine.
- 11y ago
- asddubs 11y agoYou understand that all you are doing with that is giving people possibly looking for attack vectors an attack vector, right? All an evildoer has to do is embed the blacklist directory as an image somewhere, send it to someone they want to lock out of the service, etc.
- psykovsky 11y agoWhile you have a very valid point, in my use case I just ignore it. I don't see why anyone would want to block anyone else from accessing a small independent label website. Or the personal blog of a friend of mine. Or the portfolio site of another friend who is a designer. (edit)Or what real harm could come from it if it happened.
- njharman 11y agoI agree with you (on risk is acceptable) but I also believe you underestimate the pettiness of people. I'm sure there are people who believe the ind label or your friends are enemies (for whatever real or imagined reason). But the set of those people who also have technical ability or access to technical ability to enact this "revenge" is almost certainly empty.
- laumars 11y agoIt wouldn't be too hard to prevent that. eg the honeypot directory name could be a hash of the originating IP. You'd need to then have a dynamic robots.txt but that's easily done. The destination directory doesn't even need to exist. Worst case scenario you could handly the hash via your 404 handler or via .htaccess file if all of your hashes are prefixed. Those are only examples though - there's a multidude of ways you could handle the incoming request.
- egeozcan 11y agoAnd then our hypothetical attacker can figure out how you generate the honeypot URL and embed an image with that URL for their visitors. Of course you can easily make it impossible to guess but you need to make sure that they can't obtain robots.txt via GET requests from a visitors browser (No Access-Control-Allow headers). Also don't forget a single visit from a network, sharing an IP would still ban all the network. And there is the rule that a GET request should not change any state. Banning an IP address is a change of state. In short, it's just too much pain for little gain. Maybe a login form can be served from that URL, and any attempts to login would then get the visitor banned via a session cookie / browser fingerprint combo (Easy to get around but at least then you're not blocking IP addresses).
- liviu 11y agoWhat if I put a link to this URL and google bot will follow? You will block google bot?
- Ded7xSEoPKYNsDd 11y agoGooglebot won't follow the link because it's listed in robots.txt.
- ceequof 11y agoGooglebot caches robots.txt for a very, very long time. If you disallow a directory it may take months for the entire googlebot fleet to start ignoring it. Google's official stance is that you should manage disallow directives through webmaster tools.
- ryanlol 11y agoYes, but it will index it.
- lucb1e 11y agoDitto. My favorite name is /porn and then see who visits it. Mostly bots, though.