6 ms·
The worrying thing here is that so many people accepted outsourcing the decision on who can see their website to a large company (Cloudflare). If the company de
by jwr 2mo ago
The worrying thing here is that so many people accepted outsourcing the decision on who can see their website to a large company (Cloudflare). If the company decides that a certain user should not see the website, the user will not see the website, and no one will know about it, and the user will have no recourse. That is not the open web that I would like to see.
A second side effect of a knee-jerk reaction to bots crawling websites is that if you try to fight all bots, you also end up hurting real users that use "bots". If I run a local script or an LLM that needs to fetch a web page from your site and you block that script or LLM as being a bot, you hurt me, the user. I wanted the information and I have no way of getting at it without investing my own time and effort in going to your website. That might or might not be what you expected, but it's worth taking into account.
And finally, something worth noting is that there are so many websites whose owners complain about bots, but the real problem is that the website is poorly built and should be improved anyway. Bot traffic is not necessarily bad.
- eigencoder 2mo ago> I wanted the information and I have no way of getting at it without investing my own time and effort in going to your website. I think you've hit the nail on the head here. I think that's a big reason why people want to ban bots.
- aomix 2mo agoThe social contract of accepting scraping for visibility was always tenuous and is now fully dead. But it’s shocking the number of people who are essentially victim blaming here. You should spend YOUR time to optimize your free website so MY use case is not impacted. How is that not incredibly selfish on its face?
- iwontberude 2mo ago[dead]
- autoexec 2mo agoThe social contract for putting a website on the public internet involves following standards and making your website accessible. You SHOULD make sure that your website is readable to people no matter what user agent they are using as long as it's standards compliant. For example, making a website that only works in Edge and refuses to load for anything else would be a shitty practice. You are free to do it. There's no one stopping you, but you should expect people to bitch about it and they might think you're kind of a jerk or maybe just bad at making websites. It's the same thing when a site actively rejects traffic from curl, wget, or any other common utility that many people find useful.
- bakugo 2mo agoThe social contract for putting a website on the public internet was built on the assumption that browsers would be used to access those websites, and that most browsers would generally behave in a similar way (click on page, display page, pretty simple). LLMs are not browsers, nor are they people, so I don't see why they should be part of this equation at all. They do not display the page for a real person to read, they merely ingest its contents and then spit out something completely different.
- autoexec 2mo ago> The social contract for putting a website on the public internet was built on the assumption that browsers would be used to access those websites, and that most browsers would generally behave in a similar way (click on page, display page, pretty simple). No, not browsers, user agents (some of which would be browsers) and there was never an assumption of how a website would behave on the user's end, that's the job the user agent. All that HTML and CSS are only suggestions, but the power was always intended to be left to the user to decide if/how they wanted that data presented to them and it was always intended that the user be able to choose whatever tools they wanted to collect, process, and display content pulled down from the internet. That same principle is how we have ad-blockers. You are free to infest your website with ads, but as the person requesting the website I'm under zero obligation to display any part of that site I don't want. > LLMs are not browsers, nor are they people, so I don't see why they should be part of this equation at all. LLMs are just another tool used by people to collect and process the information available on websites. Maybe there is a distinction to be made between people using LLMs to get web content and corporations scraping websites to take training data, but even scraping has always been a common and expected practice. It's the current scale that is making things different.
- paul7986 2mo agoExactly and Cloudflare (there should be others out there too - open source options) has proposed that AI bots need to pay for access to our websites. If they and or others could pull off blocking AI access until it pays creators then AI is forced to pay as it should and always should've! Overall systems and different business models need to be created that forces AI to pay it's fair share! Elon says there will be an abundance thanks to AI and we wont have to work -- ok how's that gonna work without different systems in place paying us? I feel strongly about this topic and proposed some systems in this Substack post https://ryanspahn.substack.com/p/ai-to-pay-for-all-americans-content https://ryanspahn.substack.com/p/ai-to-pay-for-all-americans...
- andai 2mo ago...because they want to make it harder for people to get the information they need? I think I must be missing something here.
- moralestapia 2mo ago>That is not the open web that I would like to see. Cloudflare is opt-in so I don't see that being an issue (yet).
- skinfaxi 2mo agoBeing opt-in doesn't negate the fact that it is closing the web.
- moralestapia 2mo agoI agree but it's a bit nuanced. If you can still buy a domain, publish a site, and other people read it as usual, then the web is still open imo. But I can see a lot of negative network effects if/when Cloudflare gets to control 60%+ of web traffic.
- IsTom 2mo agoInternet walled gardens are all opt-in and still they've made the web a worse place.
- aeddZX_0 2mo agoHow do you monetise bot traffic?
- gmerc 2mo agohttps://claw-guard.org/adnet https://claw-guard.org/adnet
- whynotmaybe 2mo ago> Revenue settles in $CLAW tokens Yes, but no.
- gmerc 2mo agopa-ro-dy
- paytonjjones 2mo agohttps://blog.cloudflare.com/introducing-pay-per-crawl/ https://blog.cloudflare.com/introducing-pay-per-crawl/
- autoexec 2mo agoYou take your site off the public internet and paywall it off. Websites on the open internet should be accessible and ideally the goal should be to share something cool with the world, not just to make yourself rich.
- testing22321 2mo ago> You take your site off the public internet and paywall it off I did that. I even printed and bound it in a bunch of dead trees and put it up for sale. The LLMs still stole it.
- autoexec 2mo agoPaywalling off the site solves the problem of monetizing bot traffic, any bot crawling your pages paid you to be there, but it can't fix the plagiarism/copyright infringement problem
- hk__2 2mo ago> If I run a local script or an LLM that needs to fetch a web page from your site and you block that script or LLM as being a bot, you hurt me, the user. No; in this case you are not a user, you are a bot user.
- ako 2mo agoThe best way to read the information on the internet today is via a LLM. Just like you don't process raw data in a data warehouse or datalake by looking at tables, you use a SQL or BI tool, to process all the information on the internet you need to tool to digest it for you. For many, today that tool is a LLM chat interface or agent.
- deleted 2mo ago[deleted]
- ballooney 2mo agoThis is such a grim thing to read.
- doc_ick 2mo ago100%
- aprdm 2mo agoWhy do you feel that ?
- sdellis 2mo agoBecause AI cannot be trusted as a reliable source of information. Information is more trustworthy when it comes from from primary sources on the open web.
- aprdm 2mo agoWhy not ? Can google be trusted ? Or facebook ? How do people have been accessing the internet for the last decade you reckon and how's this any different ?
- bob1029 2mo ago> the real problem is that the website is poorly built This is almost always the problem. This is currently the problem that GitHub is having too. If they had remained with their crusty old rails architecture, very little of this mess [0] would be occurring right now. They would have been able to focus all their engineering talent on scaling the product rather than inventing elaborate client side state synchronization mechanisms. [0]: https://www.githubstatus.com/incidents/qcvjkzcs7j74 https://www.githubstatus.com/incidents/qcvjkzcs7j74
- throw93003838 2mo agoWe tried to deploy private cloud github enterprise node back in 2018. It was pure garbage without CI integration. I am happy microsoft took harder long term decision!
- tonyhart7 2mo agowhy you blaming cloudflare that try to solve botting issue and not the Botters ??? you literally can turn off cloudflare and use your own solution
- skinfaxi 2mo agoSometimes the cure is worse than the disease.
- compiler-guy 2mo agoTrue, but for whom? The cure here isn't perfect, but much better than the disease of paying hundreds of dollars a month for scrapers which will never be beneficial. Worse for the scrapers really isn't anyone's problem but the scrapers'.
- skinfaxi 2mo agoSome of us are labelled as bots, much like dolphins getting caught in fishing nets. I guess it's not material since it's not life or death (yet? if access to essential services is gated by bot detection we are all screwed).
- nullsanity 2mo ago[dead]
- dspillett 2mo ago> If I run a local script or an LLM that needs to fetch a web page from your site and you block that script or LLM as being a bot, you hurt me, the user You are mistaking yourself, well your bot, as his target audience. You might as well say “If I want to send you my commercial email, and you block it, you hurt me, the email user.”. While your point of being concerned about cloudflair becoming a global arbiter of who gets in and who does not (which may at times not just mean blocking bots, intentionally or through technical issues), the need to block the deluge of bot traffic is very real for many sites and that is one of the easy options for them to deal with that. There are other methods like directives in robots.txt and nofollow attributes on links, but so many bots simply ignore those that they are not really useful. > Bot traffic is not necessarily bad. Nor is it necessarily wanted. In fact, it often isn't. Unfortunately practically all bot runners seem to either assume that their traffic is the special good kind or not care either way. > I wanted the information and I have no way of getting at it without investing my own time and effort in going to your website. I WANT! I WANT!! I WANT!!! Well, that site runner wants you to access the site as a human, if at all, not via bots. Sorry to be the one to break it to you, but what you want isn't always the most important factor for the rest of us. > but the real problem is that the website is poorly built and should be improved anyway Firstly: just no. Secondly: if you and your bot don't like our badly built sites, feel free to go get your information from those that you consider to be better built. Problem solved.
- BLKNSLVR 2mo agoMy understanding would be that, if an API is available, then bots are essentially welcome. If no API is available and the website has to be 'visited' then it seems that it's intended for human consumption 'the old fashioned way'.
- autoexec 2mo ago> If I run a local script or an LLM that needs to fetch a web page from your site and you block that script or LLM as being a bot, you hurt me, the user. People trying to block bots end up keeping out all kinds of users. I get blocked frequently for using a regular browser, just with JS disabled. 99% of the time, I just close the browser tab and move on with my life.
- lostmsu 2mo agoThe situation is getting worse in some OSS communities too. I go to a bugtracker just to read the discussion on the issue I am facing, often to understand what's the roadmap to fix it if any, and some of them immediately demand me to enable JS and solve a captcha. Even GitHub doesn't do it!
- hluska 2mo agoThis is a whole lot of things you want and feel entitled to. Nobody has to cater to you - we can block whatever we want to block. And if our poorly designed sites bother you, that’s too bad for you. But your wants are not my problem. If you want someone to cater to you, pay them. You aren’t entitled to anything.
- BLKNSLVR 2mo agoThe attitude you're describing seems startlingly common, as if it's not the actual site owner that has made the active choice to block something that's causing them trouble. They all sound as if they've been logic-twisted by some product idea they think is going to make them rich, and these blocks on bots are costing them access to the raw materials for their magnificently worthwhile project.
- buzer 2mo ago> If the company decides that a certain user should not see the website, the user will not see the website, and no one will know about it, and the user will have no recourse. If the the processing is subject to GDPR (e.g. if controller is in EU) then you do have recourse. You can complain to DPA or sue the company. The company is ultimately responsible for the decision to block you, at least in cases where you personally tried to access the site.
- wolrah 2mo ago> If I run a local script or an LLM that needs to fetch a web page from your site and you block that script or LLM as being a bot, you hurt me, the user. I wanted the information and I have no way of getting at it without investing my own time and effort in going to your website. The article you're replying to describes in explicit detail how the bots and their operators have directly caused and continue to knowingly cause real harm to the author and others in similar positions, both financial costs and administrative/maintenance burdens that would not have otherwise been required. You then respond "but if you block the bots then I won't be able to use the bots, and that harms me because I might have to read your web site myself..." Are you serious? > That might or might not be what you expected, but it's worth taking into account. I would wager that for almost everyone who is blocking bots after getting functionally DDoSed by them this is absolutely an expected and desired outcome. > And finally, something worth noting is that there are so many websites whose owners complain about bots, but the real problem is that the website is poorly built and should be improved anyway. Both can be true. If you operate a git repository with a public-facing web interface for example there are going to be a lot of possible operations that are inherently expensive but also incredibly rarely used by normal users so it doesn't really matter, but the bots now ignore your robots.txt and are programmed to go after every link they can find, so they trigger every single possible expensive operation more times in a night than your actual users ever have in the history of the site while dividing requests across so many different IP addresses that rate limiting becomes impossible at the individual scale. These days they're even feeding the discovered URLs back in to their models to have them invent new possible URLs and trying those in hopes of finding content never publicly linked. They will send you thousands of requests for URLs that they literally made up. Sometimes the site is in fact badly coded and operations that should be simple have higher costs due to bad design but you don't have to look very far to find situations where legitimately high-cost resources are exposed to the public because they're expected to be used in a non-abusive way. We should always be standing up against abuse of public resources, unless we want to lose them altogether. > Bot traffic is not necessarily bad. You are right, but whether it's good or bad more or less comes down to a cost/benefit analysis. As we've already covered infinite times, these bots being used to train LLMs cause significant real costs to the operators of these sites. What benefits do they offer in return? We know the clickthrough rates are terrible, so what other reasons would site operators have to make those real costs worth it? So someone can get a response back from an algorithm that confidently misinterprets or even entirely misreports what the data actually meant? Even the most die-hard "information wants to be free" types who absolutely want their datasets trained on would probably prefer that the bots accessed the data directly via an API or downloaded a database dump rather than spidering and scraping a web interface intended for humans.
- testing22321 2mo ago> I wanted the information and I have no way of getting at it without investing my own time and effort in going to your website. Substitute the word “website” for book, or training course, or documentary or published paper, or patent or one of hundreds of examples. Now you see the problem.
- ethin 2mo agoAnd if you believe that blocking bots is bad, then by all means provide a better alternative that doesn't increase costs for sysadmins. If you believe that the website should just be improved, then by all means feel free to provide instructions on what should be improved and exactly how so that we don't ever have to block bots anymore. I'm sure all the sysadmins having to deal with issues like this one will thank you
- lostmsu 2mo agoHuh? I regularly ask codex to not just summarize or extract a crux of published papers, but sometimes to do a topic directed search and give a comparison table. If you aren't doing that, do you see the problem?
- andersmurphy 2mo agoAgreed! Cloudflare absolutely destroys user experience and honestly doesn't seem that effective in practice. What's worked for me is I block any client that don't support brotli compression and http2. Seems to work well enough for stopping scrapers.
- arcrevenant 2mo agoJokes on them, the second I see that “Verifying you are human…” redirect I leave and never return.
- m463 2mo agoI am denied by cloudflare CONSTANTLY on one system. I have an old os (macos 10.11), running the highest firefox esr I can run, and I get denied by cloudflare. But not always immediately - I get to enable javascript/cookies sometimes just to be denied. they are not the folks we want gatekeeping the internet, they are opportunists increasing their OKRs
- qingcharles 2mo agoAnd here's the rub -- the bots are "running" the latest "MacOS" and have no problem accessing the site.
- deleted 2mo ago[deleted]
- binaryturtle 2mo agoI'm getting a "browser not supported" by the Cloudflare check. So I guess the "job is well done", and the user is lost.
- BLKNSLVR 2mo ago> I wanted the information and I have no way of getting at it without investing my own time and effort in going to your website. And this is a complaint? I find nothing in this statement to sympathize with whatsoever. Isn't the contract of, essentially everything, that effort is required to obtain something worthwhile? Please let me know where I'm getting this long-standing, fairly fundamental understanding of the world, wrong.
- lostmsu 2mo agoBecause the effort you are talking about is completely unnecessary. Imagine every grocery store would require you do a little dance when you buy a carton of milk. Your vision needs some narrowing (which I believe will justify the parent point), because as-is it has obvious counterexamples. I believe your problem is that "effort" is unspecified. "some effort" would make the statement correct, but some effort does not justify arbitrary effort, therefore you have no point here.
- account42 2mo agoGrocery stores do in fact make you walk through a maze to get to the items you want and absolutely would block any technology that would let you get only what you had in mind without seeing what else is on offer.
- TheRealPomax 2mo agoNo one was hurt, you were at most inconvenienced. And that's perfectly fine. Also, using jargon that doesn't apply: a knee-jerk reaction is one that does way too much to address a small problem. Objectively, this is the literal opposite: targeted reactions to different aspects of a huge problem.
- bluefirebrand 2mo ago> I wanted the information and I have no way of getting at it without investing my own time and effort in going to your website Yes, that's the social contract. Bots are not a part of it
- unsungNovelty 2mo ago> If I run a local script or an LLM that needs to fetch a web page from your site and you block that script or LLM as being a bot, you hurt me, the user. I wanted the information and I have no way of getting at it without investing my own time and effort in going to your website. You remind me of a friend who asks TLDR for every single thing! If u aren't bothered to consume the content as envisioned by the OP/Author, go away. You have plenty of other sources. We are living in the age of abundant information. Otherwise? Play by the rules. Even when you are doing this on a service u paid for. If it's not on their terms and conditions, u are simply not supposed to use it that way.