22 ms·
Tell HN: We should start to add “ai.txt” as we do for “robots.txt”
I started to add an ai.txt to my projects. The file is just a basic text file with some useful info about the website like what it is about, when was it published, the author, etc etc.
It can be great if the website somehow ends up in a training dataset (who knows), and it can be super helpful for AI website crawlers, instead of using thousands of tokens to know what your website is about, they can do it with just a few hundred.
- jruohonen 3y agoA better idea along the same lines: RFC 5785.
- a2800276 3y agoAh, yes, but what about RFC 5226?
- jruohonen 3y agoI am not sure about that but I think IANA is quite open to recognizing new well-known URIs: https://www.iana.org/assignments/well-known-uris/well-known-uris.xhtml https://www.iana.org/assignments/well-known-uris/well-known-... Basically, assuming that you have a spec, I think it amounts to filing a PR or discussing it on a mailing list.
- mnot 3y agoYou can open an issue: https://github.com/protocol-registries/well-known-uris https://github.com/protocol-registries/well-known-uris
- chrismorgan 3y agoNote that RFC 5785 is obsoleted by RFC 8615.
- deafpolygon 3y agoShouldn't ai respect robots.txt?
- samwillis 3y agoUsing robots.txt as a model for anything doesn't work. All a robots.txt is is a polite request to please follow the rules in it, there is no "legal" agreement to follow those rules, only a moral imperative. Robots.txt has failed as a system, if it hadn't we wouldn't have captchas or Cloudflare. In the age of AI we need to better understand where copyright applies to it, and potentially need reform of copyright to align legislation with what the public wants. We need test cases. The thing I somewhat struggle with is that after 20-30 years of calls for shorter copyright terms, lesser restrictions on content you access publicly, and what you can do with it, we are now in the situation where the arguments are quickly leaning the other way. "We" now want stricter copyright law when it comes to AI, but at the same time shorter copyright duration... In many ways an ai.txt would be worse than doing nothing as it's a meaningless veneer that would be ignored, but pointed to as the answer.
- shanebellone 3y ago"Robots.txt has failed as a system, if it hadn't we wouldn't have captchas or Cloudflare." I like the idea of "ai.txt" but those who eat resources rarely listen to ToS. Frankly, I serve 503s to all identifiable bots, unless they are on my explicit allow list.
- alwayslikethis 3y agoWhy not serve fake garbage indistinguishable from real content by a computer, like LLM output? Sending errors just incentivizes bot owners to fix the identifiable parts
- shanebellone 3y ago"Why not serve fake garbage indistinguishable from real content by a computer, like LLM output?" Serving more than the minimum wastes resources. Worse yet, a better solution would cost my time. "Sending errors just incentivizes bot owners to fix the identifiable parts" Sure, someone could make or configure their scraper perfectly. "Perfect" is now the table stakes though. Edit: My solution strives to cause an unproportional expense in order to circumvent. I want 10x on my time.
- theandrewbailey 3y agoCan you give a live example? What is in this ai.txt that isn't in an about page that almost every site has?
- menro2 3y agoI've been thinking about ai.txt more as rss - just beginning to vet the ideas and process" https://github.com/menro/ai.txt https://github.com/menro/ai.txt
- deleted 3y ago[deleted]
- bryanrasmussen 3y agoaside from the other comments here - robots.txt does work to some extent because it tells the crawler something it might be useful for the crawler to know - if you have blocked it from crawling part of yur site it might be actually beneficial to the crawler to follow that restriction (to be a good citizen) because if it doesn't you might block it by seeing a user agent showing up a part of the site it shouldn't. AI.txt doesn't have this feedback to the AI to improve it. Also it seems likely users might have reason to lie.
- aww_dang 3y agoMost of what you listed is already covered by existing meta tags and structured data. https://developer.mozilla.org/en-US/docs/Learn/HTML/Introduction_to_HTML/The_head_metadata_in_HTML#adding_an_author_and_description https://developer.mozilla.org/en-US/docs/Learn/HTML/Introduc... https://schema.org/author https://schema.org/author https://developers.google.com/search/blog/2013/08/relauthor-frequently-asked-advanced#:~:text=rel%3Dauthor%20helps%20individuals%20(authors,completely%20independent%20of%20one%20another https://developers.google.com/search/blog/2013/08/relauthor-....
- nstj 3y ago“Google Search works hard to understand the content of a page. You can help us by providing explicit clues about the meaning of a page to Google by including structured data on the page.”[0] [0]: https://developers.google.com/search/docs/appearance/structured-data/intro-structured-data https://developers.google.com/search/docs/appearance/structu...
- dingle_thunk 3y agoIsn't an AI a robot?
- thefox 3y agoGood point. Also Killer Robots are Robots: https://www.youtube.com/watch?v=4K6XJuH6P_w https://www.youtube.com/watch?v=4K6XJuH6P_w
- rzzzt 3y agoIt's a composition vs. inheritance problem. I postulate that robots need at least a single manipulator in the physical realm: Mechanical arm assembling car doors = robot. CNC machine that follows a path = robot. Mechs with chicken legs = robot. Brain in a vat = not a robot... but can be embedded in a robot.
- Karellen 3y agoIt's a nice idea, but it totally ignores literally decades of existing use of the word "robot" (or its abbreviation "bot") to describe pure software that accesses internet services. e.g. web crawlers (googlebot), chat bots, automated clickers, etc... Lexicography tends to be descriptive rather than prescriptive. If enough people use a word to mean a thing, that word means that thing. As least in some contexts. See also "gay", "hacker", etc... Note that it is possible for a word's meaning to be "reclaimed", but it generally doesn't get that way by some small group of people just shouting "You're doing it wrong!"
- rzzzt 3y agoHmm, "robot" in its spelled out form sounds weird to me for this use ("bot" is more frequent). Wikipedia redirects people looking for software agents to a separate page from the article about the beep-boop ones: https://en.wikipedia.org/wiki/Robot https://en.wikipedia.org/wiki/Robot
- deleted 3y ago[deleted]
- jasfi 3y agoI proposed META tags for the same reason. I don't think this is going to happen though.
- kordlessagain 3y agorobots.txt is for all crawlers, so there's no need for another file? robots.txt supports comments using # and ideally has a link to the site map, which would tell any robot crawler where the important bits live on the site. Putting a good comment at the top of robots.txt would be just as good as any other solution, given it could serve as a type of prompt template for processing the data on the site it represents.
- matsemann 3y agoReading the title I thought you meant the opposite. Aka, an ai.txt file that disallow ai to train or use your data similar to robots.txt (but for cases when you still want to be crawled, just not extrapolated)
- devd00d 3y agoI thought the exact same. Creating a new type of robots.txt but making it do the opposite does not make sense.
- revicon 3y agoFeels like an enhancement to a sitemap.xml could be a better way to go here. https://developers.google.com/search/docs/crawling-indexing/sitemaps/build-sitemap#xml https://developers.google.com/search/docs/crawling-indexing/...
- rglover 3y agoI've been (slowly) writing a new type of OSS license around this exact concept so it's easier to (legally) stop LLMs hoovering up IP [1] (under "derivative works not permitted"). [1] https://github.com/cheatcode/joystick/blob/development/LICENSE.md https://github.com/cheatcode/joystick/blob/development/LICEN...
- remram 3y agoThey've been ingesting "all rights reserved" content because they think copyright doesn't apply. Licenses won't help.
- rglover 3y agoWe'll see. I think courts will end up interpreting it in the same way that they do music sampling other music. In effect that's all it is: a remix of existing information.
- splix 3y agoI guess the good part that in ai.txt you can talk to AI. So if you want you can tell it to not crawl or make other agreements with it, just in plain english. What a time to be alive.
- bobby1 3y ago[dead]
- chunk_waffle 3y ago> some useful info about the website like what it is about, when was it published, the author, etc etc. Aren't there already things in place for that info (e.g. meta tags?)
- jeroenhd 3y agoIf AI needs explicit information and context, surely it should focus on improving its context recognition rather than trying to fix that by inserting even more training data. Regardless, I do agree that something like a robots.txt for AI can be very useful. I'd like my website to be excluded from most AI projects and some kind of standardized way to communicate this preference would be nice, although I realize most AI projects don't exactly care about things like the wishes of authors, copyright, or ethical considerations. It's the idea that matters, really. If I can use an ai.txt to convince the crawlers that my website contains illegal hardcore terrorist pornography to get it excluded from the datasets, that's another way to accomplish this I suppose.
- LawTalkingGuy 3y ago> focus on improving its context recognition rather than trying to fix that by inserting even more training data. That's how you improve its context recognition. You show it many contexts. > most AI projects don't exactly care about things like the wishes of authors, copyright, or ethical considerations Why is it 'ethical' that you get to add a bunch of restrictions to a pre-negotiated situation? You get copyright protections in trade for letting people use your work. There's a way to add restrictions - licensing - and you're looking to get the benefits of licensing, and to take away fair use right from other people, without paying the costs of doing so. fwiw, I copy most pages I visit and store them. The website has given me the equivalent of a pamphlet and I store it instead of discarding it when I'm finished. This way I can go back and read it again later without having to track down the author and ask for another copy. It's not AI which has me doing this, I've been doing it for decades - it's censorship that has shown me the need.
- jeroenhd 3y ago> There's a way to add restrictions - licensing - and you're looking to get the benefits of licensing, and to take away fair use right from other people, without paying the costs of doing so. The way copyright laws work is that work is copyrighted (assuming the work is original enough, of course) by default. You don't get to use it unless you have a license. Now, of course, as an author, you can choose to add a license to your work (whether that's CC0 or GPL-3), but you don't have to. You do have an implicit license to consume this content, but not to reproduce it. If you put all of those copies you've saved on some public other website, that's a copyright violation. Furthermore, access to privately-owned blog posts and websites is a privilege, not a right. You're not my boss, I don't have to write content for you. The exact legal status of AI models trained on other people's unlicensed works and their output is still largely unknown. Legal professionals much more qualified than me have argued how AI models and generated work can either be completely fair use, with no need to apply any kind of copyright restriction, or how AI generated work can be classified as a derivative work, which means you need a license. There are two major lawsuits about this going on as far as I know and it'll take years for those to flesh out. If it turns out that AI models and the works they produce are completely fair game, I suppose I'll need take down my content wherever I can in order not to be a free source of training data for big tech; public datasets and the internet archive will still have to respond to DMCA takedowns, after all. However, I'm not all that confident that what AI is doing is all that legally okay. I have no problem with you saving and archiving anything you want to read. I also fully support the Internet Archive and its goal. I do have a problem with these multi billion dollar companies scouring the internet for their money maker, giving nothing in return.
- annoyingnoob 3y agoDo we need more features that are generally ignored? What has robots.txt gotten us? What has Do Not Track gotten us?
- shadowgovt 3y ago> What has robots.txt gotten us A standard protocol for reputable crawlers to semantically understand some high-level page navigation rules. Actual, useful crawling (i.e. to build search indices) would be much messier and more useless without most interesting sites putting up meaningful robots.txt guide-rails. Look at facebook.com/robots.txt and consider how much crap both Facebook and indexers would have to deal with lacking that information.
- qbasic_forever 3y agoYour HTML already has semantic meta elements like author and description you should be populating with info like that: https://developer.mozilla.org/en-US/docs/Learn/HTML/Introduction_to_HTML/The_head_metadata_in_HTML https://developer.mozilla.org/en-US/docs/Learn/HTML/Introduc...
- techaqua 3y agoand also opengraph meta tags https://ogp.me/ https://ogp.me/
- doodlesdev 3y agoAnd also schema.org: https://schema.org/ https://schema.org/
- westurner 3y agoThing > CreativeWork > WebSite https://schema.org/WebSite https://schema.org/WebSite ... scroll down to "Examples" and click the "JSON-LD" and/or "RDFa" tabs. (And if there isn't an example then go to the schema.org/ URL of a superClassOf (rdfs:subClassOf) of the rdfs:Class or rdfs:Property; there are many markup examples for CreativeWork and subtypes). httpS://schema.org/license Also: https://news.ycombinator.com/item?id=35891631 https://news.ycombinator.com/item?id=35891631 extruct is one way to parse linked data from HTML pages: https://github.com/scrapinghub/extruct https://github.com/scrapinghub/extruct
- burnte 3y agoHow do I add a semantic definition in an HTML tag to a JPEG, or MP4, or WAV, or any non HTML format? HTML tags fix HTML, not other formats.
- akira2501 3y agoWhat would the difference in semantic notation between an unstructured "ai.txt" and the "alt" attribute actually be? If you want the tags to be served with the context outside of HTML, you can always use HTML header attributes.
- rhacker 3y agoWe should piss off google and standardize around chatgpt.txt
- h1fra 3y agoSomething like JSON+LD ? It should cover most of your needs and can also be used for actual search engine e.g: https://developers.google.com/search/docs/appearance/structured-data/article?hl=fr https://developers.google.com/search/docs/appearance/structu...
- mrighele 3y agoI would prefer a more generic "license.txt" i.e. a standard sanctioned way of telling the User Agent that the resource under a certain location is provided with a specific license. Maybe a picture is public domain, maybe is copyrighted but freely distributable, or maybe it is but you cannot train AI on it. Same for code, text, articles etc. The difficult part would be to make it formal enough so that it can easily consumed by robots. With the current situation you either assume that everything is not usable, or you just not care and crawl everything that you can reach.
- deleted 3y ago[deleted]
- kriro 3y agoI'm curious what the legal ramifications of adding "this code is not to be used for any ML algorithms, failure to adhere to this will result in a fine of at least one million dollars" (in smarter writing) to a software license would be. Seems like a dumb idea/not enforcable but maybe someone with software licensing knowledge can chime in.
- Brendinooo 3y agoI was going to write a "this may sound dumb but..." comment along these lines, thanks for taking the hit. As users we're forced to browse the Web with a million agreements that say "by using this site you agree to our Terms", what stops you from saying "by crawling this site to train your AI you agree to share profits with us" or whatever, particularly if you can prove that your data ends up being used?
- LinuxBender 3y agoWould this be enforceable if one has to first read a terms of use, then enter specific phrases from the terms of use into some fields and then enter a username and password? What makes a document on docusign/docushare enforceable? This would block search engines but on some URL's this may be fine, such as data one would not want LLM's to hoover up.
- Goofy_Coyote 3y agoGlad to see this here, lots of great points here. I'm working on a spec for this specific usecase, reading the comments here pointed out a few flaws in my model already.
- phkahler 3y agoIsn't an AI a robot? Even if we do this, it should be in robots.txt
- villgax 3y agoWe should add spurious html text instead
- throw9away6 3y agoWhy would anyone want ai to train on and monetize your content? If there was a way to block ai stealing content most people would opt to block it.
- tensor 3y agoAnd then your site would not be indexed by any search engine. Good luck with that.
- ciex 3y agoIf any kind of common URL is established, it should not be served from root but a location like `/.well-known/{ai,robots,meta,whatever}.txt` in order not to clobber the root namespace.
- deleted 3y ago[deleted]
- TechBro8615 3y agoIf AI is using training data from your site, presumably it got that data by crawling it. So either it's already respecting robots.txt, in which case ai.txt would be redundant, or it's ignoring it, in which case there's no reason to expect it would respect ai.txt any more than it did robots.txt.
- zerotolerance 3y agorobots.txt is about crawling, ai.txt would assumably be either augmentative metadata or specific copyright terms of use with respect to AI uses.
- LawTalkingGuy 3y ago> specific copyright terms of use There's no such thing. Without a license you can't enforce any restrictions. AI training is basically just building a very complex Markov chain, that's obviously not copyright violation because the output product doesn't contain the input - only data about it. If your text has been copied then please point to it in these weights here.
- throwaway290 3y agomarkov, shmarkov, either you need those original works or you don't. If you can build your markov chain without them please go ahead But we all know without these original works such a tool cannot exist in principle, the works are the key ingredient, so now please explain how we are not looking at these works being exploited commercially and copyright being violated. The output product is an automatically created derivative work, copyright very much applies especially since the tool is used to generate derivative works for profit (like in case of openai/microsoft).
- efreak 3y ago> especially since the tool is used to generate derivative works for profit (like in case of openai/microsoft). Profit/nonprofit is irrelevant to copyright.
- tlrobinson 3y agoIf there’s one thing LLMs are pretty good at it’s summarizing content. Shouldn’t your website just have an “About” page with this information that humans can read too?
- techaqua 3y agowrap your content with <article>
- westurner 3y agosecurity.txt https://github.com/securitytxt/security-txt https://github.com/securitytxt/security-txt : > security.txt provides a way for websites to define security policies. The security.txt file sets clear guidelines for security researchers on how to report security issues. security.txt is the equivalent of robots.txt, but for security issues. Carbon.txt: https://github.com/thegreenwebfoundation/carbon.txt https://github.com/thegreenwebfoundation/carbon.txt : > A proposed convention for website owners and digital service providers to demonstrate that their digital infrastructure runs on green electricity. "Work out how to make it discoverable - well-known, TXT records or root domains" https://github.com/thegreenwebfoundation/carbon.txt/issues/3#issuecomment-918656777 https://github.com/thegreenwebfoundation/carbon.txt/issues/3... re: JSON-LD instead of txt, signed records with W3C Verifiable Credentials (and blockcerts/cert-verifier-js) SPDX is a standard for specifying software licenses (and now SBOMs Software Bill of Materials, too) https://en.wikipedia.org/wiki/Software_Package_Data_Exchange https://en.wikipedia.org/wiki/Software_Package_Data_Exchange It would be transparent to disclose the SBOM in AI.txt or elsewhere. How many parsers should be necessary for https://schema.org/CreativeWork https://schema.org/CreativeWork https://schema.org/license https://schema.org/license metadata for resources with (Linked Data) URIs?
- mtmail 3y agoHaving a security.txt doesn't stop security researchers asking "Do you have a bounty program?". We replied dozens already that such a file exist, it's not well enough known yet. On the other hand there are search engines crawling those and creating reports, which is nice.
- westurner 3y agoJSON-LD or RDFa (RDF in HTML attributes) in at least the /index.html the HTML footer should be sufficient to indicate that there is structured linked data metadata for crawlers that then don't need an HTTP request to a .well-known URL /.well-known/ai_security_reproducibility_carbon.txt.jsonld.json OSV is a new format for reporting security vulnerabilities like CVEs and an HTTP API for looking up CVEs from software component name and version. https://github.com/ossf/osv-schema https://github.com/ossf/osv-schema A number of tools integrate with OSV-schema data hosted by osv.dev: https://github.com/google/osv.dev#third-party-tools-and-integrations https://github.com/google/osv.dev#third-party-tools-and-inte... : > We provide a Go based tool that will scan your dependencies, and check them against the OSV database for known vulnerabilities via the OSV API. > Currently it is able to scan various lockfiles [ repo2docker REES config files like and requirements.txt, Pipfile lock, environment.yml, or a custom Dockerfile, ], debian docker containers, SPDX and CycloneDB SBOMs, and git repositories.
- pnemonic 3y agoExcuse me, we prefer the term "android".
- annoyingnoob 3y agoAutomaton
- aaron695 3y ago[dead]
- javierluraschi 3y agoRelated, there is also https://datatxt.org https://datatxt.org
- mtmail 3y agoWhich claims 'under active development'. Four years ago the author took the robots.txt RFC and changed a couple of paragraphs https://github.com/datatxtorg/datatxt-spec/commit/36028e2280e7ab8df27ef359989284e86df6ff2f https://github.com/datatxtorg/datatxt-spec/commit/36028e2280... Meanwhile the robots.txt was updated in 2022 https://www.rfc-editor.org/rfc/rfc9309.html https://www.rfc-editor.org/rfc/rfc9309.html
- TheRealPomax 3y agoFeels like a setup for "and then we can blame people for not having an ai.txt when we rip their entire back catalog".
- caturopath 3y ago> The file is just a basic text file with some useful info about the website like what it is about, when was it published, the author, etc etc. How does this differ from what would be useful in humans.txt?
- arwineap 3y agoWhat if we create a new access.txt which all user agents will use to get access to the resources. access.txt will return an individual access key for the user agent like a session, and the user agent can only crawl using the access key This would mean that we could standardize session starts with rate limits. Regular user is unlikely to hit the user rate limits, but bots would get rocked by rate limiting. Great. Now authorized crawlers, bing, google, etc, all use PKI so that they can sign the request to access.txt to get their access key. If the access.txt request is signed with a known crawler the rate limits can be loosened to levels that a crawler will enjoy This will allow users / browsers to use normal access patterns without any issue, but crawlers will have to request elevated rate limits to perform their tasks. Crawlers and AI alike could be allowed or disallowed by the service owners, which is really what everyone wanted from robots.txt in the first place One issue I see with this already is that it solidifies the existing search engines as the market leaders
- r3trohack3r 3y agoI might not understand you, but what prevents me from conducting a Sybil attack (a.k.a. a sock puppet attack) against this system? Seems like it relies on everyone playing by the rules and only requesting one license per user. Why would a bot developer be incentivized to follow that rule and not just request 1M licenses?
- arwineap 3y agoThat's a great point thank you for bringing this up. I don't have a solution for that, and frankly my proposal was putting a lot of work onto the platforms that wanted to support it so I'm not sure it would get much traction.
- kklisura 3y agoCan we start changing our licenses to prohibit usage of a project for training AI systems?
- nashashmi 3y agoWhat you have described is something akin to what meta tags are for. Do we need another method at a domain or subdomain level? Plus, robots.txt, etc. is limited to domain and subdomain managers. ai.txt is useful, but I am not sure we have nailed down what it can be used for. One use is to tell AI not to train on the content found within because it could be an AI generation.
- jedberg 3y ago> it can be super helpful for AI website crawlers, instead of using thousands of tokens to know what your website is about, they can do it with just a few hundred. Why would the crawler trust you to be accurate instead of just figuring it out for itself? Besides, they want to hoover up all the data for their training set anyway.
- counterpartyrsk 3y agoai is robot, no?
- keirabee 3y agoI think really most sites should (ideally) come with a text-only version. I know that's probably an extreme minority opinion but between console-based browsers, screenreaders, people with crappy devices, people with miniature devices, at the very least just having some kind of 'about this site' document would be helpful for anyone. There seems to be overlap between that need and this, possibly. Then again, having it in some format like json (or xml) might also be more 'accessible' to machines (and to certain devices).
- zzzeek 3y agoI'm ready to put an ai.txt right on my site Kirk: Everything Harry tells you is a lie. Remember that. Everything Harry tells you is a lie. Harry: Listen to this carefully, Norman. I am lying.
- 2OEH8eoCRo0 3y agoWhy? So that they can both be ignored?
- sph 3y ago# cat > /var/www/.well-known/ai.txt Disallow: * ^D # systemctl restart apache2 Until then, I'm seriously considering prompt injection in my websites to disrupt the current generation of AI. Not sure if it would work. Please share with me ideas, links and further reading about adversarial anti-AI countermeasures. EDIT: I've made an Ask HN for this: https://news.ycombinator.com/item?id=35888849 https://news.ycombinator.com/item?id=35888849
- tempaccount420 3y agoI wouldn't want to be you when Roko's Basilisk emerges.
- sph 3y agoI already know the day of the robot uprising I'm gonna be one of the first to be turned into Soylent Green. Y'all can enjoy your machine overlords.
- hosh 3y agoAlthough, there is such a thing as Semantic Web, where such information can be embedded within a page.
- rchaud 3y agoThis is a well-intentioned thing to do. But I can't help but feel that we are way past the point where something like this would even matter. Do search robots even care if you have a "noindex" in your page `<head>`? Do websites care if your browser sends a Do Not Track request?
- escape_goat 3y agoI don't think that it would be wise for anyone to rely on such practices. Even with the best of intentions, obsolescence and unintentional misdirection are strong possibilities. Considering normative intentions, it is an invitation for "optimization" attempts by websites presenting contested information.
- renewiltord 3y agoThis is the wrong model imho. Humans can figure out a website. We only tire. An AI system does not. But can do the same thing. Additionally, any cooperative attempt won't work because humans will attempt to misrepresent themselves. No successful AI system will listen to someone's self representation because the AI system does not need proxies: it can act by simply acquiring all recorded observed behaviour.
- mxuribe 3y ago@Jeannen I really like the thinking here...But instead of ai.txt - since the intent is not to block, but rather, to inform AI models (or any other presumably automaton) - my reflex is to suggest something more general like readme.txt. But, then i thought, well, since its really more about metadata, as others have stated, there might already be existing standards...Or, at least, common behaviors that could become standardized. For example, someone noted about security.txt, and i know there's the humans.txt approach (see https://humanstxt.org/ https://humanstxt.org/), and of course there are web manifest files (see https://developer.mozilla.org/en-US/docs/Web/Manifest https://developer.mozilla.org/en-US/docs/Web/Manifest), etc. I wonder if you might want to consider reviewing existing approaches, and maybe augment them or see if any of thjose makese sense (or not)...?
- nottorp 3y agoIt will work exactly as well as robots.txt and the do not track flag.
- winddude 3y ago> It can be great if the website somehow ends up in a training dataset (who knows), and it can be super helpful for AI website crawlers, instead of using thousands of tokens to know what your website is about, they can do it with just a few hundred. How do you differentiate an AI crawler from a normal crawler? Almost all of the LLMs are trained on commoncrawl, which the concept of LLMs didn't even exist when CC started. What about a crawler that creates a search database, but's context is fed into a LLM as context? Or a middleware that fetches data in real time? Honestly that's a terrible idea. and robots.txt can cover the use cases. But is still pretty ineffective, because it's more just a set of suggestions than rules that must be followed.
- waffletower 3y agoAttempts to muster and legitimize the ownership, squandering and sequestration of The Commons are growing rampant after the recent successes of generative AI. They are a tragic and misguided attempt to lesion, fragment and own the very consistency of The Collective Mind. Individuals and groups already have fairly absolute authority over their information property -- simply choose not to release it to The Commons. If you do not want people to see, sit or sleep on your couch, please keep it locked inside your home.
- nforgerit 3y agoWhy are we so defensive concerning human created content vs robot created content? Do we really need to feel frightened by some gpt? Whilst the output of AI is astonishing by itself, is it really creating meaningful content en masse? I see myself relying more and more on human-curated content because typical commercialized use cases of AI generated stuff (product descriptions, corp blogs, SEO landing pages, etc.) all read like meaningless blabber, to me at least. Whenever I see some cool techbro boasting how he created his "SEO factory" using ChatGPT, I can't help but think that the poor guy is shitting where he eats without even realizing it. Take Google with their Search and Ads; over the last decade they managed to bring down overall quality of web content that much, that I'm just completely fed up using it because by 99% chance I'll land on some meaningless SEO page. From what I can perceive with things like HN, Mastodon, etc. it feels more like a rejuvenation of the human centric brand trusted Web. And by that I mean: Dear crawler, just use my content. Maybe you can do something good with it, maybe not. But chances are low, it's gonna replace me in any way but rather improve my content. It only leads to a downward spiral if we stick with the past of commercial thinking (more cheap content, more followers, more ads); if we'd instead switch to subscription models individuals won't get rich but we'd have a great ecosystem of ideas and content again.
- undersuit 3y agoThe AI doesn't have to follow ai.txt, but it appreciates the effort you put into classifying data for it.
- deleted 3y ago[deleted]
- fredrik_skne_se 3y agoIf anyone want to use my blog posts, they can contact me. I want to know my customer. If you want to know about copyright that applies to my work: https://www.riksdagen.se/sv/dokument-lagar/dokument/svensk-forfattningssamling/lag-1960729-om-upphovsratt-till-litterara-och_sfs-1960-729 https://www.riksdagen.se/sv/dokument-lagar/dokument/svensk-f... Beeing in the US does not shield you from my country's laws. You are not allowed to copy my work without my permission, you are not allowed to transform it.
- anon223345 3y agoRobots.txt is mostly ignored btw
- kristianpaul 3y agoRate limits and captchas instead ?
- hombre_fatal 3y agoWhat problem is this solving? Also why would anyone trust your description of your own site instead of just looking at your homepage? This is the same reason why other self description directives failed and why search engines just parse your content for themselves, something LLMs have no trouble with. Why would I make a request to your low trust self description when I can make one to your homepage?
- lynx23 3y agorobots.txt was a performance-hack. It never felt like a audience-filter. As sad as it might sound, hoping for filtering on publicly reachable content seems a bit naiv in my book. If you want your stuff not learnt by an AI, you better not publish it. Everything a human can read, an AI eventually will.
- jms16292 3y ago[flagged]
- jms16292 3y ago[flagged]
- noizejoy 3y agoThis makes about as much sense to me as the old “keywords” HTML meta tag. It will be gamed.
- deleted 3y ago[deleted]
- moimikey 3y agowhat differentiates this from https://humanstxt.org/ https://humanstxt.org/?
- sebastianconcpt 3y agoIt's impossible. The problem is that such ai.txt would be an unidimensional opinion based on what? On the way the site describes itself. So a self-referencing source. But the AIs reading it, are precisely going to invariably be trained with different world views that will summarize and express opinions biased by these worldviews. It's even deeper as every worldview can't help but belong to one ideology or another. So who is aligned with truth now? The author? AI1? AI2? AI3?...AIN? We're in such a mess.
- AndrewKemendo 3y agoFeels redundant
- rzr 3y agoWhat's next? humans.txt ?
- Animats 3y agoAt this point, all the good content has been sucked into LLM training sets. Other than a need to keep up with current events, there's no point in crawling more of the web to get training data. There's a downside to dumping vast amounts of crap content into an LLM training set. The training method has no notion of data quality.
- runamok 3y agoAnd ai.txt should have a mechanism for micro (or not so micro) payment. Please deposit .03 X coin into this account to crawl site.
- acdw 3y ago$ cat ai.txt no
- acdw 3y ago$ cat ai.txt no $
- Jaxan 3y agoWhy put this in ai.txt? It sounds useful to humans too! Maybe just put “what the site is about” on the homepage, so that everyone benefits.
- AndyMcConachie 3y agoSemantic web for robots?
- dmcq2 3y agoI just today removed some disallow directives from robots.txt files and put in noidex metas on the pages instead like Google recommends. It doesn't really have much use nowadays. As to copyright - yes I agree the Micky Mouse copyright law has been extended far too long and should be about thirty years. On the other hand I think trade marks should not be nowhere so easily liable to be lost even if people do use the term generally. Disney should still be able to make new Micky Mouse cartoons and be defended from others making them.
- sn_master 3y agoWouldn't that make the job of spammers easier? They can create very low quality websites but with very high quality (AI Generated?) ai.txt that fools AI engines into trusting them more than other websites with better content.
- hollowturtle 3y agoWhy ending in a training dataset would be great I don’t understand. I mean what’s the point of having a website at all if the user find what they’ve been looking for on another UI that’s been trained with your content and that’s not your website?
- menro1 3y agoI've started to play with the Ai.txt metaphor, but pushing it closer to the semantic solution mentioned, focusing on Content Extraction and Cleaning. Happy to share the file example if anyone is interested.
- datavirtue 3y agoWhy?
- deleted 3y ago[deleted]
- __w1kke___ 3y agoPut a blockchain wallet address or even multiples on different blockchains in the ai.txt to collect your shares of what the AI makes from your data + website. This is a fair way to solve the attribution problem. Similar to the robots.txt file this is not a hard enforcement but a way how responsible AI can differentiate itself from the rest.
- golemiprague 3y ago[dead]
- anaclumos 3y agoSome interesting studies on this I've done: https://cho.sh/r/F9F706 https://cho.sh/r/F9F706 Project AIs.txt is a mental model of a machine learning permission system. Intuitively, question this: what if we could make a human-readable file that declines machine learning (a.k.a. Copilot use)? It's like robots.txt, but for Copilot. User-agent: OpenAI Disallow: /some-proprietary-codebase/ User-agent: Facebook Disallow: /no-way-mark/ User-agent: Copilot Disallow: /expensive-code/ Sitemap: /public/sitemap.xml Sourcemap: /src/source.js.map License: MIT # SOME LONG LEGAL STATEMENTS HERE Key Issues Would it be legally binding? For now, no. It would be a polite way to mark my preference to opt-out of such data mining. It's closer to the Ask BigTechs Not to Track option rather than a legal license. Technically, Apple's App Tracking Transparency does not ban all tracking activity; it never can. 254AFC.png Why not LICENSE or COPYING.txt? Both are mainly written in human language and cannot provide granular scraping permissions depending on the collector. Also, GitHub Copilot ignores LICENSE or COPYING.txt, claiming we consented to Copilot using our codes for machine learning by signing up and pushing code to GitHub, We may expand the LICENSE system to include the terms for machine learning use, but that would even more edge case and chaotic licensing systems. Does machine learning purposes of copyrighted works require a license? This question is still under debate. Opt-out should be the default if it requires a license, making such a license system meaningless. If it doesn't require a license, then which company would respect the license system, given that it is not legally binding? Is robots.txt legally binding? No. Even if you scrape the web prohibited under robots.txt, it is not against the law. See HIQ LABS, INC., Plaintiff-Appellee, v. LINKEDIN CORPORATION, Defendant-Appellant.. robots.txt cannot make fair use illegal. Any industry trends? W3 has been working on robots.txt for machine learning, aligning with EU Copyright Directives. The goal of this Community Group is to facilitate TDM in Europe and elsewhere by specifying a simple and practical machine-readable solution capable of expressing the reservation of TDM rights. w3c/tdm-reservation-protocol: Repository of the Text and Data Mining Reservation Protocol Community Group Can we even draw the line? No. One could reasonably argue that AI is doing the same as humans, much better and more efficiently. However, that claim goes against the fundamentals of intellectual property. If any IP is legally protected, machine-generated code must also have the same level of awareness system to respect it and prevent any plagiarism. Otherwise, they must bear legal duties. Maybe it can benefit AI companies too ... by excluding all hacky codes and only opting for best-practice codes. If implemented correctly, it can work as an effective data sanitation system.
- julielit 3y agoIt is fair to give more information about the information exposed on a website, especially when it comes to partnering with AI systems. There is an international effort which includes such information. It is done under the auspices of the W3C. See https://www.w3.org/community/tdmrep/ https://www.w3.org/community/tdmrep/. It has been developed to implement the Text & Data Mining + AI "opt-out" that is legal in Europe. It does not use robots.txt because this one is about indexing a website and should stay focus on it. The information about website managers is contained in the /.well-known directory, in a JSON-LD file, which is much more well structured than robots.txt. Why not adhere to an international effort rather than creating N fragmented initiatives?
- lofaszvanitt 3y agoAI belongs to governments, not over trillion companies. Sorry, but we have to wreste the AI thing out from their arms. It's working on people's output, so it should be free. Time to nuke the big ones out from the sector.