3 ms·
Does any of the LLM providers actually use llms.txt? If I remember correctly this "standard" was setup by someone but without involvement of any of the major A
by rickette 4mo ago
Does any of the LLM providers actually use llms.txt?
If I remember correctly this "standard" was setup by someone but without involvement of any of the major AI players.
- HermanMartinus 4mo agoI can definitively say llms.txt is not used by any AI players. I run a blogging platform with around 80k blogs and /llms.txt is not requested by anything (other than humans checking to see if there's an llms.txt path). All regular pages are aggressively scraped to the extent it's a problem I have to consistently manage, but not llms.txt.
- isaachinman 4mo agoHow is a static blog being scraped a problem? Do you not use a CDN?
- the_real_cher 4mo agoAre all blogs static though?
- johannes1234321 4mo agoVery few blogs require frequent updates. Even with user comments.
- nickserv 4mo ago> a blogging platform with around 80k blogs But nah, I'm sure OP doesn't know about CDNs.
- nickserv 4mo agoI'm seeing quite a bit of request for these on my work's GitBook documentation site. But perhaps these are developers specifically targeting these pages to feed whatever LLM they are using.
- 0123456789ABCDE 4mo ago> I can definitively say llms.txt is not used by any AI players. https://developers.openai.com/llms.txt https://docs.anthropic.com/llms.txt https://geminicli.com/llms.txt https://github.com/llms.txt https://docs.aws.amazon.com/llms.txt https://openrouter.ai/docs/llms.txt
- m4tthumphrey 4mo agoOP clearly meant that the AI players are not reading and/or honouring llms.txt of other websites when scraping.
- 0123456789ABCDE 4mo agoi stand corrected, but what was clear to you, obviously was not clear to me.
- sunshine-o 4mo agoAmazing, I didn't know. So it get even stranger, I am the only one reading those /llms.txt ...
- 0123456789ABCDE 4mo agoyes, they do. anyone who's, even slightly, clued into how agents access documentation, has been making changes to their pages. ex: https://searchtxt-web.fly.dev/search?q=aws https://searchtxt-web.fly.dev/search?q=aws
- solumos 4mo agoNo, requesting "Accept: text/markdown" in the headers and returning markdown is the more agreed upon standard at this point.[0] [0] - https://acceptmarkdown.com/ https://acceptmarkdown.com/
- christoff12 4mo agoThis is interesting. I should start incorporating this -- it couldn't hurt to do both.
- kamma4434 4mo agoNow, it would be super cool to get markdown and zero javascript bundles…
- solumos 4mo agoIf you want to see what that looks like, I one-shot a browser with Claude that does it[0]. Docs pages are early adopters to this[1][2], so that AI agents can better handle tasks. [0] - https://github.com/solumos/md-browse https://github.com/solumos/md-browse [1] - https://docs.stripe.com https://docs.stripe.com [2] - https://vercel.com/docs https://vercel.com/docs
- sunshine-o 4mo agoI just found out Cloudflare supports real-time html to md conversion [0] - [0] https://blog.cloudflare.com/markdown-for-agents/#convert-html-to-markdown-automatically https://blog.cloudflare.com/markdown-for-agents/#convert-htm...