6 ms·
On the linked page, I see comments about ignoring the webmasters' wishes et al. All I can say is f*ck that. It's a free and open internet. If you put content u
by libeclipse 9y ago
On the linked page, I see comments about ignoring the webmasters' wishes et al.
All I can say is f*ck that. It's a free and open internet. If you put content up on a public site, anyone has the right to go and look at it. Stop complaining when someone saves it.
And sure some people complain that scrapers slow down their site and that's why they use robots.txt, but really? Really? It's 2017 and your site is affected by that. I think you have bigger things to worry about.
- chii 9y ago> scrapers slow down their site and that's why they use robots.txt a poorly written scraper may really slow down your site, especially if it wasn't intended to be scrapped repeatedly. There should be something to be said about frequency which scrapers should follow (specified by the website owner via a robots.txt like spec). But website owners cannot demand unreasonable frequencies (such as once a year!), and what constitutes unreasonable is up for debate.
- davb 9y agoAdditionally, if excessive scraping became an issue for my site I'd consider rate limiting client.
- timClicks 9y agoThe Crawl-delay directive is the de facto standard for this.
- eknkc 9y agoI don't think a poorly written scraper would follow robots.txt rules according to spec. So, in any case the site should have other measures (rate limiting?) anyway.
- dingo_bat 9y ago> (specified by the website owner via a robots.txt like spec). Nope, if a website wants such a restriction, it must enforce it. Robots.txt is a request. It's worthless.
- bkor 9y agoIf a robot misbehaves, it'll either be blocked or it'll go to the networks abuse section and that bot will be taken down. That a site possibly could have some kind of technical solution to this doesn't matter.
- problems 9y agoPrecisely - the solution here needs to be that the server blocks the robot - if it can differentiate it from other traffic that is. That's all well and good and that's the solution which should be used here. If you don't want to be archived, block the IP.
- blowski 9y agoWhat's your opinion on Google scraping your content and putting minimally attributed snippets?
- hoschicz 9y ago"minimally attributed"? On Google? Really?
- fenwick67 9y agoSee this article for examples of how snippetsa don't always attribute correctly: https://theoutline.com/post/1399/how-google-ate-celebritynetworth-com https://theoutline.com/post/1399/how-google-ate-celebritynet... The author runs CelebrityNetWorth.com, which BusinessInsider cites, but the snippets cite BusinessInsider. So the user doesn't see the proper attribution.
- LoSboccacc 9y agothey only ever needed to honor the robots.txt at the date of archival. archive.org fucked up by making robots retroactive, if they used the archived robots.txt as a filter for a site at the relevant date, they'd have had the best of both worlds - respecting how sites work without losing how sites appeared at a date.
- deleted 9y ago[deleted]
- pbhjpbhj 9y agoThey'd be the flouting copyright laws, like Google do, but nonetheless tortuously. They're making a copy, which is already an infringement, distributing it seemingly against the owners express wishes is treated as a crime in some jurisdictions.
- Retric 9y agoRobots.TXT is explicit permission to make a copy, otherwise crawling is meaningless and that is not nessisarily reversible. Like putting up a yard sale sign, then trying to get the poeple that show up yesterday arrested for trespassing. What the archive can do after that point is a different issue, but they clearly can keep a copy. Further, someone else is using the domain they don't nessisarily have anything to do with the archived data.
- reitanqild 9y agoAnyone knows why this was downvoted?
- pbhjpbhj 9y agoRobots.txt is usually explicitly permission for a robot not to crawl. But a robot crawling your site and an archive, cache, or duplicate page are all different propositions. Google and others have enhanced robots.txt to enable permission for crawling (allow, sitemap), meta tags can deny archiving and various means allow permission to be explicitly denied for caching. To use your analogy of raising a sign: if you don't put up a 'no trespassing' sign then it doesn't make trespassing legal. FWIW I disprove of this state of affairs and consider copyright to be hugely defective in these respects. >but they clearly can keep a copy // It's nuanced but permission to access a page =/= permission to keep a copy. Just as you have explicit permission to access a video on YouTube but in most jurisdictions will not have permission to download it for later (commercial) use.
- bkor 9y ago> Really? It's 2017 and your site is affected by that. That someone wants to use a robot to completely scrape an entire dynamic website is their goal. A site is not responsible to make that possible. One bot causes _way_ more traffic and CPU usage than just a normal visitor or 1000s of visitors. Saying '2017' or anything else: meh. Various network operators are pretty helpful. Sending abuse complaints regarding misbehaving bots has resulted in actions before. I've seen action being taken from universities, ISPs, etc. Though normally the bots are auto-blocked (on IP address or ranges; quite easy to script). robots.txt is an established / de facto standard. Ignore it, be prepared to explain why. IMO pretty much any country have computer hacking laws which are vague enough that to consciously ignore such a standard can be seen as "invading". A "not my problem" approach: I think you should really think a little bit more.
- madshiva 9y agoI totally ignore it and my bot never get caught. If they catch me I will say that the script wasn't working correctly, but what you are saying is wrong, there's is NO LAW stating that /robots.txt must be obeyed. Therefore it's not my problem, I just don't follow your rule, but I have the choice too and you have the choice to block my IP too which I think is more harmful. Also thanks for spreading bad information.
- adventured 9y ago> there's is NO LAW stating that /robots.txt must be obeyed. Therefore it's not my problem You're not wrong about robots.txt, you're wrong in a much more broad way. There is in fact an extremely dangerous law that could easily ensnare what you're talking about: https://en.wikipedia.org/wiki/Computer_Fraud_and_Abuse_Act https://en.wikipedia.org/wiki/Computer_Fraud_and_Abuse_Act
- madshiva 9y agoI don't know if the CFAA apply to my country, I know moreover that we don't need to comply with DCMA. I don't thing that browsing a web page and saving it's content it's the same than scamming people by doing fake online site. This is growing in our country and the local police don't have any rights. If it's a global problem we need to have global rules, we can't have Chinese not respecting Authors' rights and in the other hand only blame local people it's stupid. Specially when it's non-tech people that do the rules, they don't know tech therefore should not say anything about it. EDIT: You can be mad at me and down vote, but what I say is true and relevant. There's not only US in the world, specially when there's other way than protecting your site behind a robots.txt
- tomjen3 9y agoThey also make the content available to the public, which is directly competing with the site owners. Thats inviting lawsuits they can't win and expecting people to pay the bandwidth for it too.
- un-devmox 9y ago> If you put content up on a public site, anyone has the right to go and look at it. Fine. > Stop complaining when someone saves it. Fine. What you don't say is that it is fine to recreate and publish that content against the owner's wishes, especially when said content is copyrighted in one way or another. You're failing to see the whole picture from the content owner's point of view.
- Houshalter 9y agoConsider this very website, hacker news. Every single comment has it's own url. Which creates a huge number of redundant urls. There's no reason for archive sites to be scraping a separate copy of every single comment after it's indexed the thread. And it makes it harder to use google to search HN.
- bigbugbag 9y agoAw man you should have been there 10 years ago, this website hacker news had a url to access every single comment. And quite often some of the comment were more informative than the link poster, I've bookmarked quite a few of those myself. But with hacker news gone they're gone too, because archive.org failed: though they do have the data they broke the links and it is now inaccessible to me.
- Houshalter 9y agoI see your point. Perhaps a better solution would be to get rid of permalinks to individual comments entirely and point to them with fragment identifiers.