10 ms·
>These were scrapers, and they were most likely trying to non-consensually collect content for training LLMs. "Non-consensually", as if you had to ask for perm
by bakql 11mo ago
>These were scrapers, and they were most likely trying to non-consensually collect content for training LLMs.
"Non-consensually", as if you had to ask for permission to perform a GET request to an open HTTP server.
Yes, I know about weev. That was a travesty.
- XenophileJKO 11mo agoWhat about people using an LLM as their web client? Are you now saying the website owner should be able to dictate what client I use and how it must behave?
- aDyslecticCrow 11mo ago> Are you now saying the website owner should be able to dictate what client I use and how it must behave? Already pretty well established with Ad-block actually. It's a pretty similar case even. AI's don't click ads, so why should we accept their traffic? If it's un-proportionally loading the server without contributing to the funding of the site, get blocked. The server can set whatever rules it wants. If the maintainer hates google and wants to block all chrome users, it can do so.
- XenophileJKO 11mo agoThat was kind of what I was really hinting at, as the HN community tends to embrace things like ad blockers and archive links on stories, but god forbid someone read a site using an LLM.
- 1gn15 11mo agoHumans are usually hypocritical. They support whatever they personally use while opposing whatever inconveniences them, even though they're basically the same thing. This whole thing has made me hate humans, so so much. Robots are much better.
- aDyslecticCrow 11mo agoI use adblock myself, and don't feel bad for using it (it's a security and privacy tool). But i don't blame websites that kick me out for it; hosting costs money. Server owners should have all the right to set the terms of their server access. Better tools to control LLMs and scrapers are all good in my book. I really wish ad platforms were better at managing malware, trackers and fraud through. It is rather difficult to fully argue for website owner authority with how bad ads actually are for the user.
- grayhatter 11mo agoYes? I'd suggest that you understand that's not an unreasonable expectation either. Your browser has a bug, if you leave my webpage open in a tab, because of that bug, it's going to close the connection, reconnect, new tls handshake and everything and re-request that page without any cache tag, every second, everyday, for as long as you have the tab open. That feels kinda problematic, right? Web servers block well formed clients all the time, and I agree with you, that's dumb. But servers should be allowed to serve only the traffic they wish. If you want to use some LLM client, but the way that client behaves puts undue strain on wy server, what should I do, just accept that your client, and by proxy you, are an asshole and just accept that? You shouldn't put your rules on my webserver, exactly as much I my webserver shouldn't put my rules on yours. But i believe that ethically, we should both attempt to respect and follow the rules of the other. Blocking traffic when it starts to behave abusively. It's not complex, just try to be nice and help the other as much as you reasonably can.
- Calavar 11mo agoI agree. It always surprises me when people are indignant about scrapers ignoring robots.txt and throw around words like "theft" and "abuse." robots.txt is a polite request to please not scrape these pages because it's probably not going to be productive. It was never meant to be a binding agreement, otherwise there would be a stricter protocol around it. It's kind of like leaving a note for the deliveryman saying please don't leave packages on the porch. It's fine for low stakes situations, but if package security is of utmost importance to you, you should arrange to get it certified or to pick it up at the delivery center. Likewise if enforcing a rule of no scraping is of utmost importance you need to require an API token or some other form of authentication before you serve the pages.
- hsbauauvhabzb 11mo agoHow else do you tell the bot you do not wish to be scraped? Your analogy is lacking - you didn’t order a package, you never wanted a package, and the postman is taking something, not leaving it, and you’ve explicitly left a sign saying ‘you are not welcome here’.
- bakql 11mo agoStop your http server if you do not wish to receive http requests.
- vkou 11mo agoTurn off your phone if you don't want to receive robo-dialed calls and unsolicited texts 300 times a day. Fence off your yard if you don't want people coming by and dumping a mountain of garbage on it every day. You can certainly choose to live in a society that thinks these are acceptable solutions. I think it's bullshit, and we'd all be better off if anyone doing these things would be breaking rocks with their teeth in a re-education camp, until they learn how to be a decent human being.
- bigbuppo 11mo agoAh yes, and unplug the mail server to stop all spam. Great idea!
- isodev 11mo agoAh yes, the “it’s ok because I can” school of thought. As if that was ever true.
- munk-a 11mo agoI think there's a massive shift in what the letter of the law needs to be to match the intent. The letter hasn't changed and this is all still quite legal - but there is a significant different between what webscraping was doing to impact creative lives five years ago and today. It was always possible for artists to have their content stolen and for creative works to be reposted - but there was enough IP laws around image sharing (which AI disingenuously steps around) and other creative work wasn't monetarily efficient to scrape. I think there is a really different intent to an action to read something someone created (which is often a form of marketing) and to reproduce but modify someone's creative output (which competes against and starves the creative of income). The world changed really quickly and our legal systems haven't kept up. It is hurting real people who used to have small side businesses.
- Lionga 11mo agoSo if a house is not not locked I can take whatever I want?
- Ylpertnodi 11mo agoYes, but you may get caught, and there suffer 'consequences'. I can drive well over 220kmh+ on the autobahn (Germany, Europe), and also in France (also in Europe). One is acceptable, the other will get me Royale-e fucked. If the can catch me.
- arccy 11mo agoyeah all open HTTP servers are fair game for DDoS because well it's open right?
- sdenton4 11mo agoThe problem is that serving content costs money. Llm scraping is essentially ddos'ing content meant for human consumption. Ddos'ing sucks.
- 2OEH8eoCRo0 11mo agoScraping is legal. DDoSing isn't. We should start suing these bad actors. Why do techies forget that the legal system exists?
- ColinWright 11mo agoThere is no way that you can sue the people responsible for DDoSing your system. Even if you can find them ... and you won't ... they're likely as not either not in your jurisdiction (they might be in Russia, or China, or Bolivia, or anywhere) and they will have a lot more money than you. People here on HN are laughing at the UKs Online Safety Act for trying to impose restrictions on people in other countries, and yet now you're implying that similar restrictions can be placed on people in other countries and over whom you have neither power nor control.
- anon10484810573 11mo ago> Why do techies forget that the legal system exists? Simple, the gamble is it often makes business sense to "forget". Initially it could be they are unaware of the specific law but after a certain period of time it can really only be assumed the motivation is it is more convenient to ignore it. Not all techies are like this. Until the regulatory system corrects this calculus it will keep happening. Reputational costs or social costs sure wont correct it in this day and age.
- herbst 11mo agoFacebook and Bing sometimes are 80% of my daily hits and don't respect my IP bans and other bot filterings at all. You think I can just sue them and have any change to win before being broke?
- dylan604 11mo agorunning the scraping bots cost money too.
- jraph 11mo agoWhen I open an HTTP server to the public web, I expect and welcome GET requests in general. However, (1) there's a difference between (a) a regular user browsing my websites and (b) robots DDoSing them. It was never okay to hammer a webserver. This is not new, and it's for this reason that curl has had options to throttle repeated requests to servers forever. In real life, there are many instances of things being offered for free, it's usually not okay to take it all. Yes, this would be abuse. And no, the correct answer to such a situation would not be "but it was free, don't offer it for free if you don't want it to be taken for free". Same thing here. (2) there's a difference between (a) a regular user reading my website or even copying and redistributing my content as long as the license of this work / the fair use or related laws are respected, and (b) a robot counterfeiting it (yeah, I agree with another commenter, theft is not the right word, let's call a spade a spade) (3) well-behaved robots are expected to respect robots.txt. This is not the law, this is about being respectful. It is only fair bad-behaved robots get called out. Well behaved robots do not usually use millions of residential IPs through shady apps to "Perform a get request to an open HTTP server".
- Cervisia 11mo ago> robots.txt. This is not the law In Germany, it is the law. § 44b UrhG says (translated): (1) Text and data mining is the automated analysis of one or more digital or digitized works to obtain information, in particular about patterns, trends, and correlations. (2) Reproductions of lawfully accessible works for text and data mining are permitted. These reproductions must be deleted when they are no longer needed for text and data mining. (3) Uses pursuant to paragraph 2, sentence 1, are only permitted if the rights holder has not reserved these rights. A reservation of rights for works accessible online is only effective if it is in machine-readable form.
- klntsky 11mo ago> A reservation of rights for works accessible online is only effective if it is in machine-readable form. What if MY machine can't read it though?
- 11mo ago
- deleted 11mo ago[deleted]
- codyb 11mo agoThe sign on the door said "no scrapers", which as far as I know is not a protected class.
- anon10484810573 11mo agoThis mindset really baffles me. Just because it is not illegal doesn't mean one should do it. And for anything truly innovative there are bound to be gaps in the current law. It's pretty obvious that there is an asymmetry in benefit between those creating the models and those creating the content. If that doesn't bother you consider the fact that this currently undermines the economic and social model for open content creation on the internet. What happens when the content significantly decreases? Should those who create content not have some say in how their content is used?
- davesque 11mo agoI mean, it costs money to host content. If you are hosting content for bots fine, but if the money you're paying to host it is meant to benefit human users (the reason for robots.txt) then yeah, you ought to ask permission. Content might also be copyrighted. Honestly, I don't even know why I'm bothering to mention these things because it just feels obvious. LLM scrapers obviously want as much data as they can get, whether or not they act like assholes (ignoring robots.txt) or criminals (ignoring copyright) to get it.
- j2kun 11mo agoYou should not have to ask for permission, but you should have to honestly set your user-agent. (In my opinion, this should be the law and it should be enforced)
- grayhatter 11mo agoIf you're lying in the requests you send, to trick my server into returning the content you want, instead of what I would want to return to webscrapers, that's non-consensual. You don't need my permission to send a GET request, I completely agree. In fact, by having a publicly accessible webserver, there's implied consent that I'm willing to accept reasonable, and valid GET requests. But I have configured my server to spend server resources the way I want, you don't like how my server works, so your configure your bot to lie. If you get what you want only because you're willing to lie, where's the implied consent?
- batch12 11mo agoBrowser user agents have a history of being lies from the earliest days of usage. Official browsers lied about what they were- and still do.
- jraph 11mo agoLies in user agent strings where for bypassing bugs, poor workarounds and assumptions that became wrong, they are nothing like what we are talking about.
- gkbrk 11mo agoA server returning HTML for Chrome but not cURL seems like a bug, no? This is why there are so many libraries to make requests that look like they came from browser, to work around buggy servers or server operators with wrong assumptions.
- grayhatter 11mo ago> A server returning HTML for Chrome but not cURL seems like a bug, no? tell me you've never heard of https://wttr.in/ https://wttr.in/ without telling me. :P It would absolutely be a bug iff this site returned html to curl. > This is why there are so many libraries to make requests that look like they came from browser, to work around buggy servers or server operators with wrong assumptions. This is a shallow take, the best counter example is how googlebot has no problem identifying it itself both in and out of thue user agent. Do note user agent packing, is distinctly different from a fake user agent selected randomly from the list of most common. The existence of many libraries with the intent to help conceal the truth about a request doesn't feel like proof that's what everyone should be doing. It feels more like proof that most people only want to serve traffic to browsers and real users. And it's the bots and scripts that are the fuckups.
- malfist 11mo agoIf I set out a bowl of candy for ticker treaters, I wouldn't expect to be okay with the first adult strolling by and taking everything.
- dylan604 11mo agoand if they do, you have no recourse just like with scrapers. with the candy example, you spend you time sitting near the candy bowl supervising. for servers, we have various anti-bot supervisors. however, some asshat with no scruples can still just walk right up to your bowl and empty the contents into their bag and then just walk away even with you sitting right there. Unless you're willing to commit violence, there's nothing stopping them. now you're the assailant and the asshat is the victim. you still loose.
- righthand 11mo agoThen cutting up the candy and taping candy together in the most statistically pleasing way and finally selling all of the stolen frankenstein’s monster candy as innovative new candy and the future of humanity.
- smsm42 11mo agoYou are still trying to pretend that accessing HTTP server once and burying it under an avalanche of never-stopping bot crawlers is the same thing? And spam is the same as "sending an email" and should be treated the same? I thought in this day and age we're past that.