6 ms·
I agree. It always surprises me when people are indignant about scrapers ignoring robots.txt and throw around words like "theft" and "abuse." robots.txt is a p
by Calavar 11mo ago
I agree. It always surprises me when people are indignant about scrapers ignoring robots.txt and throw around words like "theft" and "abuse."
robots.txt is a polite request to please not scrape these pages because it's probably not going to be productive. It was never meant to be a binding agreement, otherwise there would be a stricter protocol around it.
It's kind of like leaving a note for the deliveryman saying please don't leave packages on the porch. It's fine for low stakes situations, but if package security is of utmost importance to you, you should arrange to get it certified or to pick it up at the delivery center. Likewise if enforcing a rule of no scraping is of utmost importance you need to require an API token or some other form of authentication before you serve the pages.
- hsbauauvhabzb 11mo agoHow else do you tell the bot you do not wish to be scraped? Your analogy is lacking - you didn’t order a package, you never wanted a package, and the postman is taking something, not leaving it, and you’ve explicitly left a sign saying ‘you are not welcome here’.
- bakql 11mo agoStop your http server if you do not wish to receive http requests.
- vkou 11mo agoTurn off your phone if you don't want to receive robo-dialed calls and unsolicited texts 300 times a day. Fence off your yard if you don't want people coming by and dumping a mountain of garbage on it every day. You can certainly choose to live in a society that thinks these are acceptable solutions. I think it's bullshit, and we'd all be better off if anyone doing these things would be breaking rocks with their teeth in a re-education camp, until they learn how to be a decent human being.
- bigbuppo 11mo agoAh yes, and unplug the mail server to stop all spam. Great idea!
- Calavar 11mo agoIf you are serving web pages, you are soliciting GET requests, kind of like ordering a package is soliciting a delivery. "Taking" versus "giving" is neither here nor there for this discussion. The question is are you expressing a preference on etiquette versus a hard rule that must be followed. I personally believe robots.txt is the former, and I say that as someone who serves more pages than they scrape
- yuliyp 11mo agoHaving a front door physically allows anyone on the street to come to knock on it. Having a "no soliciting" sign is an instruction clarifying that not everybody is welcome. Having a web site should operate in a similar fashion. The robots.txt is the equivalent of such a sign.
- halJordan 11mo agoNo soliciting signs are polite requests that no one has to follow, and door to door salesman regularly walk right past them. No one is calling for the criminalization of door-to-door sales and no one is worried about how much door-to-door sales increases water consumption.
- oytis 11mo ago> door to door salesman regularly walk right past them. Oh, now I understand why Americans can't see a problem here.
- ahtihn 11mo agoIf a company was sending hundreds of salesmen to knock at a door one after the other, I'm pretty sure they could successfully get sued for harassment.
- hsbauauvhabzb 11mo agoCan’t Americans literally shoot each other for trespassing?
- davsti4 11mo agoIts simple, and I'll quote myself - "robots.txt isn't the law".
- ColinWright 11mo agoQuoting Cervisia : > robots.txt. This is not the law In Germany, it is the law. § 44b UrhG says (translated): (1) Text and data mining is the automated analysis of one or more digital or digitized works to obtain information, in particular about patterns, trends, and correlations. (2) Reproductions of lawfully accessible works for text and data mining are permitted. These reproductions must be deleted when they are no longer needed for text and data mining. (3) Uses pursuant to paragraph 2, sentence 1, are only permitted if the rights holder has not reserved these rights. A reservation of rights for works accessible online is only effective if it is in machine-readable form. -- https://news.ycombinator.com/item?id=45776825 https://news.ycombinator.com/item?id=45776825
- davsti4 11mo agoIts the law pertaining to copyright. https://www.gesetze-im-internet.de/englisch_urhg/englisch_urhg.html https://www.gesetze-im-internet.de/englisch_urhg/englisch_ur... ... and that's only in Germany. If what you're protecting with robots.txt isn't copyright-able, then you'll need to find another legal means.
- bigbuppo 11mo agoViolating norms makes you an abusive jerk at best.
- nkrisc 11mo agoPut your content behind authentication if you don’t want it to be requested by just anyone.
- kelnos 11mo agoBut I do want my content accessible to "just anyone", as long as they are humans. I don't want it accessible to bots. You are free to say "well, there is no mechanism to do that", and I would agree with you. That's the problem!
- 1gn15 11mo agoWhat the hell? That is incredibly discriminatory. Fuck off. I support those that counter those discriminatory mechanisms.
- deleted 11mo ago[deleted]
- 9rx 11mo ago> as long as they are humans. I don't want it accessible to bots. A curious position. There isn't a secondary species using the internet. There is only humans. Unless you foresee some kind of alien invasion or earthworm uprising, nothing other than humans will ever access your content. Rejecting the tools humans use to bridge their biological gaps is rather nonsensical. > You are free to say "well, there is no mechanism to do that", and I would agree with you. That's the problem! I suppose it would be pretty neat if humans were born with some kind of internet-like telepathy ability, but lacking that mechanism isn't any kind of real problem. Humans are well adept at using tools and have successfully used tools for millennia. The internet itself is a tool! Which, like before, makes rejecting the human use of tools nonsensical.
- stray 11mo agoYou require something the bot won't have that a human would. Anybody may watch the demo screen of an arcade game for free, but you have to insert a quarter to play — and you can have even greater access with a key. > and you’ve explicitly left a sign saying ‘you are not welcome here’ And the sign said "Long-haired freaky people Need not apply" So I tucked my hair up under my hat And I went in to ask him why He said, "You look like a fine upstandin' young man I think you'll do" So I took off my hat and said, "Imagine that Huh, me workin' for you"
- michaelt 11mo ago> You require something the bot won't have that a human would. Is this why the “open web” is showing me a captcha or two, along with their cookie banner and newsletter pop up these days?
- bigbuppo 11mo agoUp until people started making a big stink about CAPTCHAs being used for unpaid labor at scale, uh, well they had two purposes.
- whimsicalism 11mo agoThere's an evolving morality around the internet that is very, very different from the pseudo-libertarian rule of the jungle I was raised with. Interesting to see things change.
- sethhochberg 11mo agoThe evolutionary force is really just "everyone else showed up at the party". The Internet has gone from a capital-I thing that was hard to access, to a little-i internet that was easier to access and well known but still largely distinct from the real world, to now... just the real world in virtual form. Internet morality mirrors real world morality. For the most part, everybody is participating now, and that brings all of the challenges of any other space with everyone's competing interests colliding - but fewer established systems of governance.
- hdgvhicv 11mo agoBased on the comments here the polite world of the internet where people obeyed unwritten best practices is certainly over in favour of “grab what you can might makes right”
- whimsicalism 11mo agothat was never the internet. the old internet was “information wants to be free, good luck if you want to restrict my access or resharing”
- bigbuppo 11mo agoYou're very much wrong. Two of the key tennets of libertarianism is that your rights end where my nose begins and the respect of property rights . Your AI bot is causing problems for me, then you should be compensating me for the damage or other expense you caused. But the AI bros think they should be able to take anything they want whenever they want without compensation, and they'll use every single shady behavior they can to make that happen. In other words, they're robber barrons.
- bigbuppo 11mo agoSeriously. Did you see what that web server was wearing? I mean, sure it said "don't touch me" and started screaming for help and blocked 99.9% of our IP space, but we got more and they didn't block that so clearly they weren't serious. They were asking for it. It's their fault. They're not really victims.
- jMyles 11mo agoSexual consent is sacred. This metaphor is in truly bad taste. When you return a response with a 200-series status code, you've granted consent. If you don't want to grant consent, change the logic of the server.
- jraph 11mo ago> When you return a response with a 200-series status code, you've granted consent. If you don't want to grant consent, change the logic of the server. "If you don't consent to me entering your house, change its logic so that picking the door's lock doesn't let me open the door" Yeah, well… As if the LLM scrappers didn't try anything under the sun like using millions of different residential IP to prevent admins from "changing the logic of the server" so it doesn't "return a response with a 200-series status code" when they don't agree to this scrapping. As if there weren't broken assumptions that make "When you return a response with a 200-series status code, you've granted consent" very false. As if technical details were good carriers of human intents.
- ryandrake 11mo agoThe locked door is a ridiculous analogy when it comes to the open web. Pretty much all "door" analogies are flawed, but sure let's imagine your web server has a door. If you want to actually lock the door, you're more than welcome to put an authentication gate around your content. A web server that accepts a GET request and replies 2xx is distinctly NOT "locked" in any way.
- jraph 11mo ago
- mxkopy 11mo agoThe metaphor doesn’t work. It’s not the security of the package that’s in question, but something like whether the delivery person is getting paid enough or whether you’re supporting them getting replaced by a robot. The issue is in the context, not the protocol.
- kelnos 11mo ago> robots.txt is a polite request to please not scrape these pages People who ignore polite requests are assholes, and we are well within our rights to complain about them. I agree that "theft" is too strong (though I think you might be presenting a straw man there), but "abuse" can be perfectly apt: a crawler hammering a server, requesting the same pages over and over, absolutely is abuse. > Likewise if enforcing a rule of no scraping is of utmost importance you need to require an API token or some other form of authentication before you serve the pages. That's a shitty world that we shouldn't have to live in.
- wslh 11mo ago> People who ignore polite requests are assholes, and we are well within our rights to complain about them. If you are building a new search engine and the robots.txt only include Google, are you an asshole indexing the information?
- kijin 11mo agoYes, because the site owner has clearly and explicitly requested that you don't scrape their site, fully accepting the consequence that their site will not appear in any search engine other than Google. Whatever impact your new search engine or LLM might have in the world is irrelevant to their wishes.
- DoctorOetker 11mo agoWhenever one forms a sentence, it is worthwhile to try to form a sentence that you believe to be generally true. If someone politely requests you to suck their genitalia, and you ignore that request, does that make you an asshole?
- watwut 11mo agoIf you ignore polite request, then it is perfectly ok to give you as much false data as possible. You have shown yourself not interested in good faith cooperation, that means other people can and should treat you as a jerk.
- grayhatter 11mo ago> I agree. It always surprises me when people are indignant about scrapers ignoring robots.txt and throw around words like "theft" and "abuse." This feels like the kind of argument some would make as to why they aren't required to return their shopping cart to the bay. > robots.txt is a polite request to please not scrape these pages because it's probably not going to be productive. It was never meant to be a binding agreement, otherwise there would be a stricter protocol around it. Well, no. That's an overly simplistic description which fits your argument, but doesn't accurately represent reality. yes, robots.txt is created as a hint for robots, a hint that was never expected to be non-binding, but the important detail, the one that is important to understanding why it's called robots.txt is because the web server exists to serve the requests of humans. Robots are welcome too, but please follow these rules. You can tell your description is completely inaccurate and non-representative of the expectations of the web as a whole. because every popular llm scraper goes out of their way to both follow and announce that they follow robots.txt. > It's kind of like leaving a note for the deliveryman saying please don't leave packages on the porch. It's nothing like that, it's more like a note that says no soliciting, or please knock quietly because the baby is sleeping. > It's fine for low stakes situations, but if package security is of utmost importance to you, you should arrange to get it certified or to pick it up at the delivery center. Or, people could not be assholes? Yes, I get it, the reality we live in there are assholes. But the problem as I see it, is not just the assholes, but the people who act as apologists for this clearly deviant behavior. > Likewise if enforcing a rule of no scraping is of utmost importance you need to require an API token or some other form of authentication before you serve the pages. Because it's your fault if you don't, right? That's victim blaming. I want to be able to host free, easy to access content for humans, but someone with more money, and more compute resources than I have, gets to overwhelm my server because they don't care... And that's my fault, right? I guess that's a take... There's a huge difference between suggesting mitigations for dealing with someone abusing resources, and excusing the abuse of resources, or implying that I should expect my server to be abused, instead of frustrated about the abuse.
- smsm42 11mo ago"Theft" may be wrong, but "abuse" certainly is not. Human interactions in general, and the web in particular, are built on certain set of conventions and common behaviors. One of them is that most sites are for consuming information at human paces and volumes, not downloading their content wholesale. There are specialized sites that are fine with that, but they say it upfront. Average, especially hobbyist site, is not that. People who do not abide by it are certainly abusing it. > Likewise if enforcing a rule of no scraping is of utmost importance you need to require an API token or some other form of authentication before you serve the pages. Yes, and if the rule of not dumping a ton of manure on your driveway is so important to you, you should live in a gated community and hire round-the-clock security. Some people do, but living in a society where the only way to not wake up with a ton of manure in your driveway is to spend excessive resources on security is not the world that I would prefer to live in. And I don't see why people would spend time to prove this is the only possible and normal world - it's certainly not the case, we can do better.
- o11c 11mo agoTheft is correct but for a different reason. The #1 reason for all AI scrapers is to replace the content they are scraping. This means no "fair use" defense to the copyright infringement they inevitably commit.
- bigiain 11mo ago> robots.txt is a polite request to please not scrape these pages At the same time, an http GET request is a polite request to respond with the expects content. There is no binding agreement that my webserver sends you the webpage you asked for. I am at liberty to enforce my no-scraping rules however I see fit. I get to choose whether I'm prepared to accept the consequences of a "real user" tripping my web scraping detection thresholds and getting firewalled or served nonsense or zipbombed (or whatever countermeasure I choose). Perhaps that'll drive away a reader (or customer) who opens 50 tabs to my site all at once, perhaps Google will send a badly behaved bot and miss indexing some of my pages or even deindexing my site. For my personal site I'm 100% OK with those consequences. For work's website I still use countermeasures but set the thresholds significantly more conservatively. For production webapps I use different but still strict thresholds and different countermeasures. Anybody who doesn't consider typical AI company's webscraping behaviour over the last few years to qualify as "abuse" has probably never been responsible for a website with any volume of vaguely interesting text or any reasonable number of backlinks from popular/respected sites.
- overfeed 11mo agoIt may be naivete, but I love the standards-based open web as a software platform and a s a fabric that connects people. O It makes my blood boil that some solipsistic, predatory bastards are eager to turn the internet into a dark forest