4 ms·
These are some cute "fuck off"s but its unlikely that these sites actually respect the robots.txt, right? Correct me if I'm wrong: After the recent web scrapin
by evv 4y ago
These are some cute "fuck off"s but its unlikely that these sites actually respect the robots.txt, right?
Correct me if I'm wrong: After the recent web scraping ruling[1] it seems that it's perfectly legal to ignore the robots.txt.
[1] https://news.ycombinator.com/item?id=31075396 https://news.ycombinator.com/item?id=31075396
- dave5104 4y agoDepends on the bot owner on whether they want to be respectful. Following the link to the TurnItIn bot... https://www.turnitin.com/robot/crawlerinfo.html https://www.turnitin.com/robot/crawlerinfo.html > Q: How can I completely exclude TurnitinBot from my site? > To exclude TurnitinBot from all or portions of your site all you have to to do is create a file called robots.txt and put it in the top most directory of your web site.
- bartread 4y agoWell, it's possible to also return a 403 (forbidden) to any request based off the user agent. Of course, this can be relatively easily circumvented, but then it's also possible to block IP ranges and suchlike. You can return a 403 off of any detectable aspect of the client that you don't like if you so wish. I don't know how well this would work with a CDN, but presumably if you pay for the right tier of Cloudflare (or whatever) you can perform similar operations to prevent content being hoovered from their by clients you'd prefer not to serve.
- superkuh 4y agoYep. I 403 turnitin and similar companies via nginx configuration, if ($http_referer ~* (TurnitinBot|PaperLiBot|idmarch|FairShare|Lightspeedsystems|ZmEu|BPImageWalker|semrushBot|ias_crawler|360spider|copyrightinfringementportal|PetalBot|Adsbot|SlySearch|NPBot)) { return 403; } But my favorite robots.txt is, User-agent: Zombies Disallow: /brains
- rcarmo 4y agoShouldn’t that be… User-agent: Zombies Disallow: /braaains ?
- amazing_stories 4y ago
- easrng 4y agoWhy are you blocking PetalBot? It's an actual search engine.
- superkuh 4y agoLegit Huawei IP ranges identifying as Huawei PetalBot were being abusive, definitely not obeying robots.txt, and searching for subsets of content that indicated they were looking to identify political dissidents with no worries about actually indexing the full site. I don't consider it a real search engine. But yeah, maybe not a good fit for this list of educational and copyright parasites.
- easrng 4y agoOh yikes, that sounds bad. It is a real search engine though, https://petalsearch.com/ https://petalsearch.com/
- jrochkind1 4y agoSo... not necessarily. 1. So that case was about the CFAA (Computer Fraud and Abuse Act). So at most it would say that ignoring the robots.txt does not violate the CFAA -- a law that makes some things felonies as "hacking", basically. I agree that ignoring the robots.txt (say if you are Archive Team? [1]) should not be considered a criminal "hacking" felony. But there can still be other reasons ignoring the robots.txt is against a law -- or cause for a civil tort action. (Most copyright violation is a civil tort action for instance, the CFAA is, again, a law that establishes some felonies with many years of jail time, intended to punish "hackers"). The decision in that case said nothing about anything except the CFAA. For instance, taking copyrighted content from the public web and re-selling is probably still going to put you in various kinds of legal trouble -- just not a CFAA violation. It's possible ignoring a robots.txt could put you in other kinds of criminal or civil trouble, depending on the particular circumstances -- just not a CFAA violation. It would be interesting to research what other possible liability there might be. If for instance you caused harm to the site by ignoring the robots.txt (say, an accidental or intentional DOS), I bet there'd at least be cause for civil tort. 2. Even so, even under that case, if that specific case didn't involve a robots.txt (did it?), it's always possible the presence of a robots.txt would result in a differnet outcome. My sense is probably not though, that Supreme Court decision referenced by the ninth circuit on remand -- probably does mean ignoring a robots.txt is not a violation of the CFAA. (And again, I say, PHEW, that would have been terrible if it were -- if say someone trying to archive MySpace before it went away could be put in prison for a couple decades for disrespecting the robots.txt). [1] https://wiki.archiveteam.org/ https://wiki.archiveteam.org/
- henryfjordan 4y agoThat case isn't even decided yet, the court only ruled on a preliminary injunction so there's still quite a bit of case left to go before any final decisions are made. For now it's only "likely" that HiQ will prevail (though that means it's pretty likely). In this case Linkedin sent HiQ a cease and desist letter before they sued and claimed that letter revoked access for the purpose of the CFAA, so not quite the same as a robots.txt but legally it's probably close enough. If anything it's stronger because HiQ can't claim they didn't see it.
- erk__ 4y agoWould it not have to be tried in a French court since that is where VideoLan is located?