6 ms·
Videolan.org robots.txt
- kmeisthax 4y agoSo, I can understand the hate towards copyright enforcement bots, but... did TurnItIn hammer the shit out of VLC's website? Or do VLC's developers just hate the idea of automated enforcement in general? (I doubt they're pro-plagiarism - not even copyright abolitionists go that far.)
- KennyBlanken 4y agoI imagine the objection is that it is consuming bandwidth, electricity, and computing capacity of an open source project as part of a profit-making service, with an extra fuck-you of a)crawling the site making no sense whatsoever for the service and b)the service being of no possible use to VLC or its users
- jrochkind1 4y ago> consuming bandwidth, electricity, and computing capacity of an open source project as part of a profit-making service, And yet they don't disallow Googlebot! For obvious reasons.
- ozfive 4y agoI don't get it...
- woliveirajr 4y agoProbably because of the ># --> fuck off. comment that is added after 3 specific robots.
- wormer 4y agofuck off.
- deleted 4y ago[deleted]
- ozfive 4y agoHow do you report people like this on HN?
- teaearlgraycold 4y agoSeems like you don't understand the joke. Not something I'd take action on if I was moderating HN.
- oehpr 4y agoThough, you gotta figure it's tough to be a moderator when you have a massive report queue, full of toxic behavior. You come across this post with just a plain "Fuck off", not much more or less toxic than anything else you have to deal with. But~ instead of just getting your backlog done, you instead click into the post, look at the wider context of not just the person they're responding to, but also the post itself. Toss a coin to your admin, oh valley of plenty~~~
- kome 4y agoi'm going to copy them. brilliant.
- evv 4y agoThese are some cute "fuck off"s but its unlikely that these sites actually respect the robots.txt, right? Correct me if I'm wrong: After the recent web scraping ruling[1] it seems that it's perfectly legal to ignore the robots.txt. [1] https://news.ycombinator.com/item?id=31075396 https://news.ycombinator.com/item?id=31075396
- dave5104 4y agoDepends on the bot owner on whether they want to be respectful. Following the link to the TurnItIn bot... https://www.turnitin.com/robot/crawlerinfo.html https://www.turnitin.com/robot/crawlerinfo.html > Q: How can I completely exclude TurnitinBot from my site? > To exclude TurnitinBot from all or portions of your site all you have to to do is create a file called robots.txt and put it in the top most directory of your web site.
- bartread 4y agoWell, it's possible to also return a 403 (forbidden) to any request based off the user agent. Of course, this can be relatively easily circumvented, but then it's also possible to block IP ranges and suchlike. You can return a 403 off of any detectable aspect of the client that you don't like if you so wish. I don't know how well this would work with a CDN, but presumably if you pay for the right tier of Cloudflare (or whatever) you can perform similar operations to prevent content being hoovered from their by clients you'd prefer not to serve.
- superkuh 4y agoYep. I 403 turnitin and similar companies via nginx configuration, if ($http_referer ~* (TurnitinBot|PaperLiBot|idmarch|FairShare|Lightspeedsystems|ZmEu|BPImageWalker|semrushBot|ias_crawler|360spider|copyrightinfringementportal|PetalBot|Adsbot|SlySearch|NPBot)) { return 403; } But my favorite robots.txt is, User-agent: Zombies Disallow: /brains
- rcarmo 4y agoShouldn’t that be… User-agent: Zombies Disallow: /braaains ?
- frereubu 4y agoI get some of the feeling behind this. But in terms of Turnitin my wife taught in art college and a close friend taught a Masters in Economics, and the amount of plagiarism was ridiculous. Sure, in theory there should be smaller class sizes, teachers should have more time per student, etc. etc., but Turnitin was an extremely helpful tool that meant they could offload the cognitive effort of detecting mechanical reproduction and get into reading the work. Unless there's something about Turnitin that I'm not aware of which tarnishes what they do? (Beyond making money out of already cash-strapped universities, I suppose...)
- woofcat 4y agoWhy does Turnitin get to keep a copy of all of my work for free? Do I not own the copyright of my papers?
- hackmiester 4y agoThis was my issue and was why I refused to use it.
- frereubu 4y agoOK, this objection sort of makes sense to me. Do they have something in their Ts and Cs which says "by submitting your work you consent to us storing your work..."? Presumably people who submit their work to them also benefit to some extent though, because then plagiarisers of your work will be caught?
- woofcat 4y agoOften you don't really have a choice, if you're enrolled in a school. Tough luck you get to give up your rights.
- google234123 4y agoYou don’t have a choice if your grader/professor keeps a copy either.
- paxys 4y agoWhat content is videolan.org hosting that would be relevant for these bots?
- ahmetkun 4y agoYou don't need to be relevant, just being accessible is enough reason for all sorts of bots, legit or sketchy, to shove hundreds of thousands of requests down your throat.
- jokoon 4y agoI once started logging user agents and ips for my small old php website, which is a bit hard to find. I was quite surprised to see all the weird bots that were crawling it.
- jamespwilliams 4y ago> iThenticate® Unexpected place to see latin1 -> utf8 mojibake
- a3w 4y agoWhat would a webcrawler be called which reads only the disallowed robots.txt routes? Still just an unfriendly webscraper? Shodan? Shodan on steroids?
- nixcraft 4y agoI used the ultimate Nginx bad bot blocker on a couple of my side projects, and it is a pretty good project https://github.com/mitchellkrogza/nginx-ultimate-bad-bot-blocker https://github.com/mitchellkrogza/nginx-ultimate-bad-bot-blo... . Apart from the Cloudflare offers UA blocking and AI driven bot management too. Most of these bots are for content scrapping and then creating search spam results. I am a one-person show, and it hurts both financially and resources wise on my tiny severs. So I block them.
- rickstanley 4y agoWhat is that "# $Id$" at the top? Just a comment or serves a purpose?
- tedunangst 4y agoIf the file lived in cvs, it would be replaced with the revision.
- JNRowe 4y agoIf you want a little background on Ted's answer, the keywords and their use are described in the RCS docs¹. TIL, RCS still gets releases² ;) ¹ https://www.gnu.org/software/rcs/manual/html_node/Concepts.html#Fundamental-operations https://www.gnu.org/software/rcs/manual/html_node/Concepts.h... ² https://lists.gnu.org/archive/html/info-gnu/2022-02/msg00001.html https://lists.gnu.org/archive/html/info-gnu/2022-02/msg00001...
- notorandit 4y agoFuck off: I won't crawl it through!
- 1vuio0pswjnm7 4y agoWhy crawl when a sitemap is provided. Honest question. IME, using a sitemap is much more efficient. For example, HTTP/1.1 pipelining can be used to reduce the number of TCP connections needed. Is resource exhaustion what draws a public website^1 operator's attention to "bots". If it is not resource exhaustion then what is it. 1. For this question, assume "public website" means a website serving public information where there are no legitimate intellectual property rights in the information that can be asserted by the site operator.
- spamtarget 4y agoYeah, NPBot and SlySearch can just fuck the fuck off, but what is wrong with fighting plagiarism? (honest question)
- Vladimof 4y agoMy homeworks aren't written to become public (they probably could because of their business)... I guess the schools are probably more to blame though.
- spamtarget 4y agoit's not your homework that is public, but you may sourced text from the public
- Vladimof 4y agoThat company might leak and/or sell the data that they get from the schools?
- bastawhiz 4y agoThe bot is clearly crawling enough to be noticed, and consider the site: how much are you plagiarizing from videolan.org? If they're wasting even a small amount of resources, they're worth blocking.
- spamtarget 4y agoreasonable
- Kaze404 4y agoThere are arguments on whether plagiarism is a bad thing in an academic context. I'm not nearly qualified enough to make them, but they exist if you want to go looking.