3 ms·
Except that officially, that is not what they do? It's a dick move not to put the two usecases under separate user agents, but their documentation says you're f
by Deathmax 2mo ago
Except that officially, that is not what they do? It's a dick move not to put the two usecases under separate user agents, but their documentation says you're free to block Google-Extended via robots.txt which is used for training and grounding, while still being included in the search index.
> Google-Extended does not impact a site's inclusion in Google Search nor is it used as a ranking signal in Google Search.
https://developers.google.com/crawling/docs/crawlers-fetchers/google-common-crawlers#google-extended https://developers.google.com/crawling/docs/crawlers-fetcher...
Exclusion from grounding does mean that your site won't get sourced in the AI overview, but I'm not sure what the click through rates are like on those.
- mysterydip 2mo agoAI crawlers, famous for respecting robots.txt ;)
- dannyw 2mo agoThe western AI crawlers are generally very-well behaved. GPTBot and ClaudeBot are nice crawlers and I see them respecting my robots.txt. There are of course plenty of mystery, obfuscated/camouflaged scrapers/crawlers from god knows whom. Thankfully they are easy to spot and ban, although I've definitely thought about deliberating sending them poisoned data.