14 ms·
Everyone loves the dream of a free for all and open web. But the reality is how can someone small protect their blog or content from AI training bots? E.g.: Th
by TIPSIO 1y ago
Everyone loves the dream of a free for all and open web.
But the reality is how can someone small protect their blog or content from AI training bots? E.g.: They just blindly trust someone is sending Agent vs Training bots and super duper respecting robots.txt? Get real...
Or, fine what if they do respect robots.txt, but they buy the data that may or may not have been shielded through liability layers via "licensed data"?
Unless you're reddit, X, Google, or Meta with scary unlimited budget legal teams, you have no power.
Great video: https://www.youtube.com/shorts/M0QyOp7zqcY https://www.youtube.com/shorts/M0QyOp7zqcY
- anigbrowl 1y agoMaybe this is a naive question, but why not just cut an IP off temporarily if it sends too many requests or sends them too fast?
- ceejayoz 1y agoThey use many IPs, often not identifiable as the same bot.
- clvx 1y agoYou can lock it up with a user account and payment system. The fact the site is up on the internet doesn’t mean you can or cannot profit from it. It’s up to you. What I would like it’s a way to notify my isp and say, block this traffic to my site.
- inetknght 1y ago> What I would like it’s a way to notify my isp and say, block this traffic to my site. I would love that, and make it automated. A single message from your IP to your router: block this traffic. That router sends it upstream, and it also blocks it. Repeat ad nauseum until source changes ASN or (if the originator is on the same ASN) reaches the router from the originator, routing table space notwithstanding. Maybe it expires after some auto-expiry -- a day or month or however long your IP lease exists. Plus, of course, a way to query what blocks I've requested and a way to unblock.
- deleted 1y ago[deleted]
- Gud 1y agoBy developing Free Software combating these hostile softwares. Corporations develop hostile AI agents, Capable hackers develop anti-AI-agents. This defeatist atittude "we have no power".
- banku_brougham 1y agoThis is the attitude I like to see. As they say, actually I hate this because of past connotations but "freedom isn't free"
- TIPSIO 1y agoYes, I obviously agree with you. My comment's point is missed a little I think by you. CF is making these tools and giving access to it to millions of people.
- supriyo-biswas 1y agoWell there's open source stuff like https://github.com/TecharoHQ/anubis https://github.com/TecharoHQ/anubis; one doesn't need a top-down mandated solution coming from a corporation. In general Cloudflare has been pushing DRMization of the web for quite some time, and while I understand why they want to do it, I wish they didn't always show off as taking the moral high ground.
- Klonoar 1y agoAnubis doesn’t necessarily stop the most well funded actors. If anything we’ve seen the rise in complaints about it just annoying average users.
- supriyo-biswas 1y agoThe actual response to which Anubis was created is seemingly a strange kind of DDOS attack that has been misattributed to LLMs, but is some kind of attacker that makes partial GET requests that are aborted soon after sending the request headers, mostly coming from residential proxies. (Yes, it doesn’t help that the author of Anubis also isn’t fully aware of the mechanics of the attack. In fact, there is no proper write up of the mechanism of the attack which I hope to write about someday). Having said that, the solution is effective enough, having a lightweight proxy component that issues proof of work tokens to such bogus requests works well enough, as various users on HN seem to point out.
- eldenring 1y agoI recently found out my website has been blocked by AI agents, when I had never asked for it. It seems to be opt-out by default, but in an obscure way. Very frustrating. I think some of these companies (one in particular) are risking burning a lot of goodwill, although I think they have been on that path for a while now.
- DanOpcode 1y agoAre you talking about Cloudflare? The default seems indeed to be to block AI crawlers when you set up a new site with them.
- deleted 1y ago[deleted]
- gausswho 1y agoWhat we need is some legal teeth behind robots.txt. It won't stop everyone, but Big Corp would be a tasty target for lawsuits.
- notatoad 1y agoIt wouldn’t stop anyone. The bots you want to block already operate out of places where those laws wouldn’t be enforced.
- qbane 1y agoThen that is a good reason to deny the requests from those IPs
- literalAardvark 1y agoI've run a few hundred small domains for various online stores with an older backend that didn't scale very well for crawlers and at some point we started blocking by continent. It's getting really, really ugly out there.
- deleted 1y ago[deleted]
- account42 1y agoIf that was the case then I am I getting buttflare-blocked here in the EU.
- stronglikedan 1y agoIt should have the same protections as an EULA, where the crawler is the end user, and crawlers should be required to read it and apply it.
- immibis 1y agoSo none at all? EULAs are mostly just meant to intimidate you so you won't exercise your inalienable rights.
- jMyles 1y ago> Everyone loves the dream of a free for all and open web. > protect their blog or content from AI training bots It strikes me that one needs to chose one of these as their visionary future. Specifically: a free and open web is one where read access is unfettered to humans and AI training bots alike. So much of the friction and malfunction of the web stems from efforts to exert control over the flow (and reuse) of information. But this is in conflict with the strengths of a free and open web, chief of which is the stone cold reality that bytes can trivially be copied and distributed permissionlessly for all time.
- deleted 1y ago[deleted]
- pessimizer 1y agoIt's the new "ban cassette tapes to prevent people from listening to unauthorized music," but wrapped in an anti-corporate skin delivered by a massive, powerful corporation that could sell themselves to Microsoft tomorrow. The AI crawlers are going to get smarter at crawling, and they'll have crawled and cached everything anyway; they'll just be reading your new stuff. They should literally just buy the Internet Archive jointly, and only read everything once a week or so. But people (to protect their precious ideas) will then just try to figure out how to block the IA. One thing I wish people would stop doing is conflating their precious ideas and their bandwidth. The bandwidth is one very serious issue, because it's a denial of service attack. But it can be easily solved. Your precious ideas? Those have to be protected by a court. And I don't actually care iff the copyright violation can go both ways; wealthy people seem to be free to steal from the poor at will, even rewarded, "normal" (upper-middle class) people can't even afford to challenge obviously fraudulent copyright claims, and the penalties are comically absurd and the direct result of corruption. Maybe having pay-to-play justice systems that punish the accused before conviction with no compensation was a bad idea? Even if it helped you to feel safe from black people? Maybe copyright is dumb now that there aren't any printers anymore, just rent-seekers hiding bitfields?
- andy99 1y ago> But the reality is how can someone small protect their blog or content from AI training bots? A paywall. In reality, what some want is to get all the benefits of having their content on the open internet while still controlling who gets to access it. That is the root cause here.
- notatoad 1y agoWhich is really all that cloudflare is building here that people are mad about. It’s a way to give bots access to paywalled content.
- positiveblue 1y agoWhere everyone needs a cloudflare account to be able to pay*
- notatoad 1y ago“Everyone” in this context being bot operators who want to access websites who have decided to use cloudflare to block unauthenticated bot traffic. Which is not everyone.
- littlecranky67 1y agoThis. We need to get rid of the ad-supported free internet economy. If you want your content to be free, you release it and have no issues with AI. If you want to make money of your content, add a paywall. We need micropayments going forward, Lightning (Bitcoin backend) could be the solution.
- rustc 1y ago> If you want your content to be free, you release it and have no issues with AI. If you want to make money of your content, add a paywall. What about licenses like CC-BY-NC (Creative Commons - Non Commercial)?
- 1y ago
- PaulRobinson 1y agoYou might have this the wrong way around. It's not the publishers who need to do the hard work, it's the multi-billion dollar investments into training these systems that need to do the hard work. We are moving to a position whereby if you or I want to download something without compensating the publisher, that's jail time, but if it's Zuck, Bezos or Musk, they get a free pass. That's the system that needs to change. I should not have to defend my blog from these businesses. They should be figuring out how to pay me for the value my content adds to their business model. And if they don't want to do that, then they shouldn't get to operate that model, in the same way I don't get to build a whole set of technologies on papers published by Springer Nature without paying them. This power imbalance is going to be temporary. These trillion-dollar market cap companies think if they just speed run it, they'll become too big, too essential, the law will bend to their fiefdom. But in the long term, it won't - history tells us that concentration of power into monarchies descends over time, and the results aren't pretty. I'm not sure I'll see the guillotine scaffolds going up in Silicon Valley or Seattle in my lifetime, but they'll go up one day unless these companies get a clue from history as to what they need to do.
- deleted 1y ago[deleted]
- FlyingSnake 1y agoIt is a service available to Cloudflare customers and is opt-in. I fail to see how they’re being gatekeepers when site owners have option not to use it.
- sneak 1y ago> But the reality is how can someone small protect their blog or content from AI training bots? First off, there's no harm from well-behaved bots. Badly behaved bots that cause problems for the server are easily detected (by the problems they cause), classified, and blocked or heavily throttled. Of course, if you mean "protect" in the sense of "keep AI companies from getting a copy" (which you may have, given that you mentioned training) - you simply can't, unless you consider "don't put it on the web" a solution. It's impossible to make something "public, but not like that". Either you publish or you don't. If anything, it's a legal issue (copyright/fair use), not a technical one. Technical solutions won't work. I'm not sure why people are so confused by this. The Mastodon/AP userbase put their public content on a publicly federated protocol then lost their shit and sent me death threats when I spidered and indexed it for network-wide search. There are upsides and downsides to publishing things you create. One of the downsides is that it will be public and accessible to everyone.
- buyucu 1y agoEveryone loves a free for all and open web because it works really well. Basic tools like Anubis and fail2ban are very effective at keeping most of this evil at bay.
- deadbabe 1y agoI care more about the dream of a wide open free web than a small time blogger’s fears of their content being trained on by an AI that might only ever emit text inspired by their content a handful of times in their life.
- avazhi 1y agoNobody cares about robots.txt, nor should they. If this is your primary argument against being scraped (viz that your robots.txt said not to) then you’re naive and you’re doing it wrong. If the internet is open, then data on it is going to be scraped lol. You can’t have it both ways.
- verdverm 1y agoIt seems the Open Internet is idealistic. If others respected robots.txt, we would not need solutions like what Cloudflare is presenting here. Since abuse is rampant, people are looking for mitigations and this CF offering is an interesting one to consider.
- mannanj 1y agohow about we discuss and design and implement a system that charges them for their actions? we could put some dark patterns in our sites that specifically have this cost through some sort of problem solving thing in the site that harvests their energetic scraping/LLM tools into directing their energy onto causes that give us profit on our site, in exchange for revealing some content in return that achieves their mission of scraping too. Looks like these exist to degrees.
- m463 1y agononsense. I'm routinely denied access to websites now. enable javascript and unblock cookies to continue
- account42 1y agoJavascript and cookies are far from enough, your browser also needs to look like a recent mainstream one without niche privacy extensions.
- tonetegeatinst 1y agoOnion sites have bots and scrapers. They don't use cloudlfare AFAIK. They normally use a puzzle that the website generates, or the use a proof of work based capcha. I've found proof of work good enough out of these two, and it also means that the site owner can run it themselves instead of being reliant on cloudflare and third parties.
- wvenable 1y ago> Everyone loves the dream of a free for all and open web... But the reality is how can someone small protect their blog or content from AI training bots? Aren't these statements entirely in conflict? You either have a free for all open web or you don't. Blocking AI training bots is not free and open for all.
- BrenBarn 1y agoI think that was the point. Everyone loves the dream, but the reality is different.
- wilson090 1y agoHow so? If you don't want AI bots reading information on the web, you don't actually want a free and open web. The reality of an open web is that such information is free and available for anyone.
- gradstudent 1y agoHow is it available for everyone if the AI bots bring down your server?
- sebasvisser 1y agoBuild better
- mikestorrent 1y agoUltimately, you have to realize that this is a losing battle, unless we have completely draconian control over every piece of silicon. Captchas are being defeated; at this point they're basically just mechanisms to prove you Really Want to Make That Request to the extent that you'll spend some compute time on it, which is starting to become a bit of a waste of electricity and carbon. Talented people that want to scrape or bot things are going to find ways to make that look human. If that comes in the form of tricking a physical iPhone by automatically driving the screen physically, so be it; many such cases already! The techniques you need for preventing DDoS don't need to really differentiate that much between bots and people unless you're being distinctly targeted; Fail2Ban-style IP bans are still quite effective, and basic WAF functionality does a lot.
- tiahura 1y agoWhy should your blog be protected? Information wants to be free.
- sumeno 1y agoIt's amazing how this catchphrase has reversed meanings for some people. It was previously used against walled gardens and paywalls, but these corporate LLMs are the ultimate walled garden for information because in most cases you can't even find out who created the information in the first place. "Information wants to be free! That's why I support hiding it behind a chatbot paywall that makes a few people billionaires"
- infecto 1y agoI personally love the idea of a free and open internet and also have no issues with bots scraping or training off of my data. I would much rather have it open for all, including companies, than the coming dystopian landscape of paywall gates. I don’t care about respecting robots.txt or any other types of rules. If it’s on the internet it’s for all to consume. The moment you start carving out certain parties is the moment it becomes a slippery slope.
- TIPSIO 1y agoFor what it’s worth, I think CF will lose this battle and fundamentally feeding the bots will just become normal and wanted
- deleted 1y ago[deleted]
- rsync 1y ago“But the reality is how can someone small protect their blog or content from AI training bots?” Why would you need to? If your inability to assemble basic HTML forces you to adopt enormous, bloated frameworks that require two full cores of a cpu to render your post… … or if you think your online missives are a step in the road to content creator riches … … then I suppose I see the problem. Otherwise there’s no problem.
- lovich 1y agoSo by a free and open for all web you mean only for the tech priests competent enough to build the skills and maintain them in light of changes to the spec(hope these people didn’t run across xml/xslt dependent techniques building their site), or have a rich enough family that you can casually learn a skill while not worry about putting food on the table? There’s going to be bad actors taking advantage of people who cannot fight back without regulations and gatekeepers, suggesting otherwise is about as reasonable as ancaps idea of government
- nc0 1y agoIt's not a question of languages or frameworks, but hardware. I cannot finance servers large enough to keep up with AI bots constantly scrapping my host, bypassing cache indications, or changing IP to avoid bans.
- jeroenhd 1y agoI have had to disable at least one service because AI bots kept hitting it and it started impacting other stuff I was running that I am more interested in. Part of it was the CPU load on the database rendering dozens of 404s per second (which still required a database call), part of it was that the thumbnail images were being queried over and over again with seemingly different parameters for no reason. I'm sure there are AI bots that are good and respect the websites they operate on. Most of them don't seem to, and I don't care enough about the AI bubble to support them. When AI companies stop people from using them as cheap scrapers, I'll rethink my position. So far, there's no way to distinguish any good AI bot from a bad one.
- 1y ago
- mikae1 1y ago> Great video: https://www.youtube.com/shorts/M0QyOp7zqcY https://www.youtube.com/shorts/M0QyOp7zqcY Here's an even greater video: https://www.youtube.com/watch?v=mAUpxN-EIgU&t=4m24s https://www.youtube.com/watch?v=mAUpxN-EIgU&t=4m24s
- ctoth 1y ago"I want an open web!" "Okay, that means AI companies can train on your content." "Well, actually, we need some protections..." "So you want a closed web with access controls?" "No no no, I support openness! Can't we just have, like, ethical openness? Where everyone respects boundaries but there's no enforcement mechanism? Why are you making this so black and white?"
- 1gn15 1y ago> “When we started the “free speech movement,” we had a bold new vision. No longer would dissenters’ views be silenced. With the government out of the business of policing the content of speech, robust debate and the marketplace of ideas would lead us toward truth and enlightenment. But it turned out that freedom of the press meant freedom for those who owned one. The wealthy and powerful dominated the channels of speech. The privileged had a megaphone and used free speech protections to immunize their own complacent or even hateful speech. Clearly, the time has come to denounce the naïve idealism of the past and offer a new movement, Speech 2.0, which will pay more attention to the political economy of media and aim at “free-ish” speech — the good stuff without the bad.” https://openfuture.eu/paradox-of-open-responses/misunderestimating-openness/ https://openfuture.eu/paradox-of-open-responses/misunderesti...
- account42 1y ago"I want a free an open society!" "But criminals are people too." See how stupid that sounds?
- ForHackernews 1y agoYou could run https://zadzmo.org/code/nepenthes/ https://zadzmo.org/code/nepenthes/ to punish the AI scrapers.
- BoredPositron 1y agoWe have thousands of engineers of these companies right here on hackernews and they cry and scream about privacy and data governance on every topic but their own work. If you guys need a mirror to do some self reflection I am offering to buy.
- csomar 1y agoI'll contribute for the mirror. The hypocrisy is so loud, aliens in outer space can hear it (and sound doesn't even travel in vacuum).
- tucnak 1y agoIn the recent days, the biggest delu-lulz was delivered by that guy who'd bravely decided to boycott Grok out of... environmental concerns, apparently. It's curious how everybody is so anxious these days, about AI among other things in our little corner of the web. I swear, every other day it's some new big fight against something... bad. Surely it couldn't ALL be attributed to policy in the US!
- davepeck 1y ago> Everyone loves the dream of a free for all and open web. But the reality is how can someone small protect their blog or content from AI training bots? I'm old enough to remember when people asked the same questions of Hotbot, Lycos, Altavista, Ask Jeeves, and -- eventually -- Google. Then, as now, it never felt like the right way to frame the question. If you want your content freely available, make it freely available... including to the bots. If you want your content restricted, make it restricted... including to the humans. It's also not clear to me that AI materially changes the equation, since Google has for many years tried to cut out links to the small sites anyway in favor of instant answers. (FWIW, the big companies typically do honor robots.txt. It's everyone else that does what they please.)
- BobaFloutist 1y agoWhat if I want my content freely available to humans, and not to bots? Why is that such an insane, unworkable ask? All I want is a copyleft protection that specifically allows humans to access my work to their heart's content, but disallows AI use of it in any form. Is that truly so unreasonable?
- dragonwriter 1y ago> What if I want my content freely available to humans, and not to bots? Why is that such an insane, unworkable ask? Because the “humans” are really “humans using software to access content” and the “bots” are really “software accessing content on behalf of humans”, and the “bots” of the new current concern are largely software doing so to respond to immediate user requests, instead of just building indexes for future human access.
- davepeck 1y agoIt's not unreasonable to ask but I think it probably is unreasonable to expect a strictly technical solution. It feels like we're in the realm of politics, policy, and law.
- BobaFloutist 1y ago
- arjie 1y agoThe dream is real, man. If you want open content on the Internet, it's never been a better time. My blog is open to all - machine or man. And it's hosted on my home server next to me. I don't see why anyone would bother trying to distinguish humans from AI. A human hitting your website too much is no different from an AI hitting your website too much. I have a robots.txt that tries to help bots not get stuck in loops, but if they want to, they're welcome to. Let the web be open. Slurp up my stuff if you want to. Amazonbot seems to love visiting my site, and it is always welcome.
- 0x3f 1y agoIt's traditional to include a link when claiming to be invulnerable. :)
- arjie 1y agoHaha, sounds a bit self-promotional to do that but link in profile. Not claiming that the site is technologically invulnerable. Just that it's not a big deal if LLMs scrape it (which bizarrely they do).
- danudey 1y ago> I don't see why anyone would bother trying to distinguish humans from AI. Because a hundred thousand people reading a blog post is more beneficial to the world than an AI scraper bot fetching my (unchanged) blog post a hundred thousand times just in case it's changed in the last hour. If AI bots were well-behaved, maintained a consistent user agent, used consistent IP subnets, and respected robots.txt, I wouldn't have a problem with them. You could manage your content filtering however you want (or not at all) and that would be that. Unfortunately at the moment, AI bots do everything they can to bypass any restrictions or blocks or rate limits you put on them; they behave as though they're completely entitled to overload your servers in their quest to train their AI bots so they can make billions of dollars on the new AI craze while giving nothing back to the people whose content they're misappropriating.
- immibis 1y ago
- msgodel 1y agoDon't publish things if you don't want them published. Get real yourself.
- deleted 1y ago[deleted]
- paool 1y agoYou can't trust everyone will be polite or follow "standards". However, you can incentivize good behavior. Let's say there's a scraping agent, you could make a x402 compatible endpoint and offer them a discount or something. Kinda like piracy; if you offer a good, simple, cheap service people will pay for it versus go through the hassle of pirating.
- deknos 1y ago> But the reality is how can someone small protect their blog or content from AI training bots? E.g.: They just blindly trust someone is sending Agent vs Training bots and super duper respecting robots.txt? Get real... baking in hashcash into http 1.0/1.1/1.2/2/3, smtp, imap, pop3, tls and ssh. then this will all to expensive for spammers and training bots. but IETF is infiltrated by government and corporate interests..
- Taek 1y agoSpammers will buy ASICs and get a huge advantage over consumer CPUs
- immibis 1y agoWhy do you have to protect it? Have you suffered any actual problem, or are you being overly paranoid? I think only a few people have actually received DDoS-level traffic, and the rest are being paranoid.
- lofaszvanitt 1y agoYeah where does this absolute bullshit comes from and why would anyone target your blog?
- h4ck_th3_pl4n3t 1y agoThe problem's cause really is about _who_ has to pay for the traffic, and currently that's the hosting end. If you turn that model around, suddenly AI web scrapers have to behave and all the issues that we currently have are kind of solved(?), because there is no incentive to scrape datasets anymore that were put together by others, and there automatically will be a payment incentive instead to buy high-quality datasets.
- account42 1y agoBut I don't want to make human users of my website pay for the traffic just like I also donate to real world charities that I believe in.
- johnnienaked 1y agoHow can someone small protect their IP? It's called copyright law and it's been around for a long ass time just for some reason big tech gets a pass and can steal and control whatever the fuck they want without limit
- raxxorraxor 1y agoThis isn't even an issue, you have made that problem up. I host a blog and there are some AI bots coming around. Big deal. Most of them do respect a robots.txt. Some don't. Not a big deal as well. In contrast trying to change the infrastructure of the net, which previously was quite resistant to censorship is quite a big deal. This sounds exactly like a crazy preacher warning about the dangers of rock music. A completely made up threat. And we need the protection of god against these evil AI bots. Wow, a bot that disrespected a robots.txt. How can the internet survive... Also, OpenAI already has the data. You want to ensure they will never get competitors by putting up barriers now. It makes no sense...