AI Visibility FactorsFR
All Factors >

AI crawler directives in robots.txt

Do dedicated AI crawler rules influence visibility in AI?

Impact
High
Influences
Discovery & Crawling, Understanding & Retrieval, Mention & Recommendation
Proof
High
Consensus
Strong
Category
Technical
Last reviewed
2026-09-30

AI crawler directives in robots.txt explained

AI crawler directives are dedicated robots.txt entries that name AI crawlers or platform control tokens. This factor checks whether those entries are present, which services they target, and whether their rules allow or block the intended content. File availability and general User-agent: * rules are covered in robots.txt.

Record the presence of dedicated entries separately from permission to crawl. Under the Robots Exclusion Protocol, crawlers use matching groups, falling back to User-agent: * when no specific group matches. If neither applies, no rules apply. An absent dedicated entry is therefore not automatically a problem.

Impact details

Impact: High when dedicated directives change effective search access. Major AI platforms document that blocking their search bots can exclude a site's content from search-based answers or limit its visibility in search results, depending on the service. Search permissions are distinct from training permissions; opting out of a training bot does not by itself opt a site out of search.

The impact rating describes the possible consequence of a restrictive or mistaken rule. It does not describe the incremental benefit of inserting an AI bot name into a file that already permits it. No primary evidence reviewed establishes a citation or recommendation bonus from redundant explicit permission.

The influence stages are interpreted as follows:

Assess a restriction in a dedicated AI entry here. Assess a restriction inherited from the default group under the general robots.txt factor; do not count the same block twice.

Proof & consensus details

Proof: High for permissions and documented product consequences. Major AI platforms document the role of robots.txt in allowing or excluding their bots, distinguish search from training uses, and explain the consequences of blocking access. This provides direct evidence for the access-control mechanism and its stated effects, rather than proof of a visibility gain from allowing a bot. See the documentation on crawler permissions and roles and bot exclusion and its consequences.

Consensus: Strong for distinguishing effective search permission from training permission. This does not imply universal adherence by every automated client. User-requested fetchers have different documented behavior, and platforms can use other sources. The rating is an editorial assessment of agreement on the mechanism, not a survey result or a quantified ranking weight.

There is insufficient evidence here to conclude that allowing training collection increases future brand recommendations. Likewise, an assistant mentioning a blocked site does not alone prove that it crawled that site against its rules: the answer could draw on other material.

Recommendation

Review any dedicated AI bot entries already in robots.txt. Check that their names and rules match your intended permissions, and fix accidental blocks. Decide separately whether to allow bots that collect content for training.

You do not need to add individual entries for AI bots and crawlers if your existing rules already permit their access. Add specific entries only when you need different permissions for a particular bot. If you do, keep any intended path restrictions in that entry: named bots do not automatically inherit the general User-agent: * rules. RFC 9309.

AI platforms

ChatGPT: distinguish OAI-SearchBot from training crawler GPTBot; their permissions are independent. OpenAI says opting out of OAI-SearchBot excludes a site from ChatGPT search answers, although navigational links remain possible. ChatGPT-User handles user actions; its requests may not be governed by robots.txt, and it does not determine Search inclusion. OpenAI.

Claude: distinguish search agent Claude-SearchBot, user-request agent Claude-User, and training agent ClaudeBot. Anthropic states that its bots honor robots.txt. Disabling Claude-SearchBot prevents indexing for search optimization and may reduce visibility in search results. Anthropic.

Perplexity: PerplexityBot supports search, not foundation-model training. Perplexity-User performs user-requested fetching and generally ignores robots.txt. Perplexity.

Google AI Overviews and AI Mode: Googlebot governs Search crawling; there are no additional AI-specific technical requirements. Google Search Central.

Gemini and specified Vertex AI uses: Google-Extended controls future Gemini training and grounding in Gemini Apps and Grounding with Google Search on Vertex AI. It is a control token, not a separate HTTP crawler, and does not affect Google Search inclusion or ranking. Google crawler reference.

Bing and Copilot: audit bingbot permissions for the Bing search pathway. Bing documents its crawler identities and robots.txt controls; a named AI-specific entry should not replace the Bingbot check. Bing crawler reference, Bing robots.txt guidance.

Audit instructions

  1. Use the robots.txt file checked under the general factor for each relevant hostname. List the AI services you want to assess and their documented bot names or control tokens.
  2. Record whether each has a dedicated entry: present or absent. Check names against the platform documentation rather than counting recognizable bot names.
  3. For each present entry, test its rules against representative public URLs and intentionally excluded paths. Consider repeated matching groups together, using the platform's documented behavior.
  4. Record allowed, blocked, or uncertain, together with the rule responsible. When an entry is absent, check the default rules from the general factor; absence alone is not a failure.
  5. Compare dedicated permissions with the owner's intended search, user-requested retrieval, and training choices. Flag accidental differences and unnecessary entries; keep intentional opt-outs separate from errors.