18 ms·
AI-Shunning robots.txt
- vouaobrasil 3y agoNice. Let's all contribute to this...ideally, web-hosts should provide this sort of thing by default so we can starve AI companies from training data and combine it with other strategies to put them out of business for good.
- andybak 3y agoHow about AI from non-companies? Or genuinely non-profit or open projects? Also - out of curiosity - do you use any AI yourself?
- vouaobrasil 3y ago> How about AI from non-companies? Or genuinely non-profit or open projects? AI from any project will allow AI to be used commercially, and thus I oppose it. Moreover, I oppose AI on various other princincples even independent of this: it further isolates people and can be used to develop other technologies that are too powerful for us to handle. In short, I believe human beings en mass are too stupid to use AI. > Also - out of curiosity - do you use any AI yourself? I do not, or at least I try my best not too. In fact, I hate AI with a passion. Obviously, there may be products here and there that have used AI that I in turn use. What can you do? But I attempt to minimize any contact I have with AI: I don't use Grammarly, any form of auto-suggest, I use an ancient phone (and I RARELY use it, I hate smartphones), I don't use AI features in software such as AI-noise reduction, I turn off all automatic features in software that may have some AI behind it. If I find out a website uses AI for content generation, I ban it and never visit again. The other day I downloaded a text editor that looked cool but I deleted it because I realized it has an AI-console (even though I never used it). I also work for a business and I convinced them not to use AI. We're an online magazine and it turns out the vast majority of our readers supported that decision. In short, I am against AI because I believe it provides virtually no benefits to humanity, only detriments.
- rideontime 3y agoLikewise, I've unsubscribed from multiple paid Patreons and Substacks as soon as they started using AI to generate content. I'd rather see an amateur MSPaint scribble than some dall-e monstrosity at the head of a newsletter.
- vouaobrasil 3y agoExactly! Even if a person isn't that great at drawing or painting, it's so much more interesting to view their attempts because they reflect something about the person. The AI-generated fluff might look nice (even I'll admit some look interesting), but it's devoid of human soul. And there's really not much point in art if it has no communicative value that originates from living beings.
- rideontime 3y agoYeah, I should be clear that I don't use the word "monstrosity" to suggest that all AI-generated images are ugly, but rather that they're inherently, viscerally disgusting regardless of their visual appeal. No amount of progress in the visual quality of their output will change how I feel about it.
- nerdjon 3y agoI think this is an interesting situation where “AI” has become just a general term that it has lost so much meaning thanks to things like chatgpt. Video game AI is obviously in a different league than ChatGPT but uses the same label. Some AI is machine learning and some isn’t. I agree with a lot of what you are saying, but I think there are valid use cases of AI (before chatgpt) that is actually a benefit. You don’t use a smart phone, but auto correct is a genuinely great addition. It doesn’t remove anything the human does and improves the usability. On its own autocorrect isn’t going to write a story. Even the suggestions that have been added in recent years are more for human usability than anything else. Handwriting reading models, fall detection models, etc. I do think we need to separate generative AI that replaces humans from traditional AI that offers assistance. I know someone is going to argue, well chatgpt is augmenting me by checking my code, emails, etc. and that may be true right now but we are kidding ourselves if that will be the situation long term.
- belter 3y agoOr redirect them to poisoned material?
- vouaobrasil 3y agoThat is a good idea. Maybe redirect them to massive datasets to cause the company mass embarrassment. There are already some image-modifying programs that generate poison images, and the bots could be redirected to such images...
- nerdjon 3y agoI am curious, do we have any evidence that AI is adhering to robots.txt and isn’t ignoring it since they are not technically crawling in the traditional sense? Even if they are right now it would be a quick switch for them to just ignore it.
- andybak 3y agoThis is about crawling for training data by the look of things. Not sure if the CHatGPT browsing mode uses a different user-agent but most of the entries in that list look like crawlers.
- nerdjon 3y agoI had assumed this is related to sites like chatgpt going out and searching with a specific request. Regardless, my original question is still valid. The companies have already shown a lack of care about the data they train off of. So if ethics have already gone out the window, what is to stop them from ignoring this file if they are not already.
- mrkramer 3y agoInternet Archive's crawler is not respecting robots.txt because they want to archive everything not just parts of the Web. But if you are actively breaking robots.txt then your crawler will have a bad reputation and you will have an army of webmasters trying to block your crawler by any means. You can see crawling requests in your sever logs, that's how you know if they are respecting it or not. Imo, they best solution would be to license your content so crawlers pay a fee for crawling and using your content.
- nerdjon 3y agoWell TIL that IA does not respect robots.txt. Does IA themselves block crawlers? It doesn't look like it according to their robots.txt, even going so far as to say "Please crawl our files." What would stop an actor from maliciously complying with a robots.txt file by just going to the internet archive instead.
- andybak 3y agoAs someone who uses and benefits from the results of AI crawlers, I would only want to block crawls under very specific circumstances. I would back a general move to block crawlers from non-open models (whatever that means and if such a thing was practical) as it might be a strong lever to encourage good behaviour.
- jddj 3y agoThe named source, https://darkvisitors.com https://darkvisitors.com, is interesting.
- gavinhking 3y agoI made this, let me know if you have questions or feedback.
- tbeseda 3y agoThanks for the work on this! I automated my site's robots.txt[0] by scraping your site. It would be extra nice if darkvisitor.com exposed a plain text version or JSON representation of the list. [0] https://tbeseda.com/blog/automating-my-robots-txt-to-block-ai-user-agents https://tbeseda.com/blog/automating-my-robots-txt-to-block-a...
- glynnormington 3y agoThat was one of the points of the new repo. A plain text version of the file is https://raw.githubusercontent.com/ai-robots-txt/ai.robots.txt/main/robots.txt https://raw.githubusercontent.com/ai-robots-txt/ai.robots.tx... The other point was to make this community maintained rather than rely on one source to provide all the inputs.
- tbeseda 3y agoDefinitely! And I'll likely use that raw URL in my tooling going forward. So thanks for starting the repo. I'm thinking it also helps to bring up a feature request on the source material so we can all limit the drift. ie. if darkvisitors.com had a sort of plain text API, your repo could check for new entries via GH Actions and create issues or even PRs.
- cabirum 3y agoThe crawlers can simply stop identifying themselves via custom user agent, can't they? Also why are "AI" crawlers are worse than "normal" crawlers? Either way, this is an exercise in futility.
- karaterobot 3y ago> Also why are "AI" crawlers are worse than "normal" crawlers? A search engine will index your content to bring people to it through search. An AI crawler will take your content to recapitulate it and sell it to others. Obviously it's more complicated than this, but this is how one might see it who wishes to use this file. > Either way, this is an exercise in futility. Not necessarily disqualifying. Laws against theft are also futile, in the sense that honest people don't need them and dishonest people don't follow them, and history since at least Hammurabi has been replete with examples of such laws not stopping theft. And yet. Seems worth the calories it costs to say "for the record, I do not give my consent for what you're doing".
- cabirum 3y agoSearch engines are not the beacons of holiness - they sell ads, they sell data on who searched what, they manipulate results. Search engines and AI things are typically owned by the same company. AIs are fed with the data collected by a search engine. The only difference is whether AI gets the data in realtime or waits for the search engine to collect another data dump. Fighting windmills as I see it.
- nunez 3y agoSearch engines manipulate results way less aggressively than LLMs do.
- vouaobrasil 3y ago> Either way, this is an exercise in futility. Is it really? Every drop of opposition towards AI in my book is a good thing. This robots.txt thing is a small drop maybe, but over time public hatred for AI can build and it might in fact be taken down. Especially outside the tech bubble, many people are ambivalent towards AI. Yes, in modern society were are taught to value innovation and ignore its downsides, but the more vocal opponents are against it, the more those downsides will become apparent. Hopefully, it will bring the ruin of all AI companies and research.
- bakugo 3y agoThis makes complete sense because, as we all know, AI companies are very concerned with respecting the rights of the people they steal data from, and totally won't just ignore this.
- frizlab 3y agoAt least you show intent and can then potentially prove they are not respecting your wishes. It’s better than doing nothing.
- internetter 3y agoThis is missing a couple, one that comes to mind is `FriendlyCrawler`, which is most definitely not friendly, and very likely for AI
- glynnormington 3y agoFeel free to submit a PR. :-)
- rocky_raccoon 3y agoNot that I'm arguing for or against preventing access from AI crawlers, but wouldn't it make more sense to block them at a higher level, e.g. the webserver, and not even give them the choice to obey/disobey robots.txt?
- rideontime 3y agoHow would you propose doing so?
- adrianN 3y agoWe could repurpose the evil bit.
- rideontime 3y agoOne second, let me google this. e: Okay, this is funny.
- gtirloni 3y agoWeb servers can check the user-agent and block the request. E.g. nginx $http_user_agent
- rocky_raccoon 3y agoOff the top of my head: - Cloudflare - Webserver-level user-agent blocking (Apache, nginx) - Application-level user-agent blocking (`if request.user_agent == 'OpenAI'`) None of them are ideal since you can simply change your user agent, but all of them seem like better options than robots.txt to me.
- natch 3y agoWe need AIs to know more, not less. If many people block AIs from reading their sites, AIs will just be stuffed with biased information from people pushing agendas.
- nerdjon 3y agoSo the value of them will plummet? That sounds like a win for society.
- natch 3y agoWhy would the value of AIs plummet if they know more? Or did you mean sites? Information wants to be free. If AI is trained only on data provided by those with agendas, you won’t want to live in that world.
- nerdjon 3y agoI am saying the opposite, if they have less data the value of AI's will plummet and hopefully their use will plummet. That is a good thing.
- natch 3y agoMarket dynamics. Since use of better AIs confers advantages, they will be improved. No set of players will be able to stop this because the incentives are so strong for others to continue using and developing them. The best we can hope for is AIs that are not misled. Sometimes you have to work with the tide, because fighting it is futile and even self defeating.
- starbugs 3y ago> Sometimes you have to work with the tide, because fighting it is futile and even self defeating. Said the fish before approaching the waterfall. While I agree that the incentive structure is set up in a strong way for AI to be further improved and rolled out, what's the endgame here? Who can build the most powerful centralized AI so that nearly everyone else is out of a job? And who is that going to benefit? I just don't get it. Have we all decided to "just play the game" and ignore how dumb it is?
- CalRobert 3y agoGiven how intertwined AI and search engines are it's hard to see how this helps aside from _maybe_ making things easier for Google, Microsoft, etc., unless you also don't want to be indexed by search engines.
- deleted 3y ago[deleted]