Your GEO Score
78/100
Analyze your website →

7 Rules for robots.txt: AI Bots to Allow in 2026

7 Rules for robots.txt: AI Bots to Allow in 2026

If you want to appear in AI answers, allow the search and retrieval bots in your robots.txt: Googlebot, Bingbot, OAI-SearchBot (ChatGPT search), Claude-SearchBot and PerplexityBot. Decide separately on the training bots (GPTBot, ClaudeBot, Google-Extended, Applebot-Extended, CCBot), because blocking them keeps your content out of future model training but does not remove you from search. User-triggered fetchers such as ChatGPT-User and Perplexity-User are a special case: their operators state that robots.txt may not apply to them.

The seven rules below are based on the official crawler documentation of each operator. User-agent names and purposes change, so check the linked sources before you edit your file.

AI bots at a glance: purpose and recommended action

User-agentOperatorPurpose (per operator docs)Recommendation for AI visibility
GooglebotGoogleCrawling for Google Search, including AI Overviews and AI ModeAllow
Google-ExtendedGoogleProduct token (not a separate crawler) that controls use of content for training Gemini models and for grounding in Gemini Apps and Vertex AI. Does not affect Google Search.Your choice
BingbotMicrosoftCrawling for Bing searchAllow
OAI-SearchBotOpenAISurfaces websites in ChatGPT search featuresAllow
GPTBotOpenAICrawls content that may be used to train OpenAI foundation modelsYour choice
ChatGPT-UserOpenAIFetches pages for actions a ChatGPT user initiates; robots.txt rules may not applyAllow (blocking is unreliable)
Claude-SearchBotAnthropicIndexes content to improve Claude's search resultsAllow
Claude-UserAnthropicRetrieves pages when a Claude user asks a questionAllow
ClaudeBotAnthropicCollects web content that may be used for model trainingYour choice
PerplexityBotPerplexitySurfaces and links websites in Perplexity search; not used for foundation model trainingAllow
Perplexity-UserPerplexityVisits pages for user questions; generally ignores robots.txtAllow (blocking is unreliable)
Applebot-ExtendedAppleDoes not crawl; controls whether Applebot data is used to train Apple foundation modelsYour choice
CCBotCommon CrawlBuilds the open Common Crawl web archive, which third parties can reuse, including for AI trainingYour choice

Sources: OpenAI crawler overview, Anthropic crawler FAQ, Perplexity crawlers, Google common crawlers, Apple on Applebot, Common Crawl CCBot.

Rule 1: Separate search bots from training bots

"AI bot" is not one category. Search and retrieval bots decide whether your pages can be found and cited when someone asks ChatGPT, Claude or Perplexity a question today. Training bots collect content that may end up in future models. The two decisions have different consequences:

  • Blocking a search bot (OAI-SearchBot, Claude-SearchBot, PerplexityBot) can remove you from that engine's cited sources. Anthropic writes that disabling Claude-SearchBot may reduce your visibility in user search results.
  • Blocking a training bot (GPTBot, ClaudeBot, Google-Extended, Applebot-Extended, CCBot) is a licensing and content-policy decision. It does not take you out of the respective search products.

Many sites allow everything. Publishers with paid content often allow search bots but block training bots. Both are legitimate, as long as the choice is deliberate and not the side effect of a blanket Disallow: /.

Rule 2: Do not block Googlebot or Bingbot to "stop AI"

Google is explicit: AI Overviews and AI Mode are part of Search, and robots.txt rules for Googlebot are the control for them. If you want to limit how your content is shown in AI features, use nosnippet, data-nosnippet, max-snippet or noindex. Blocking Googlebot removes you from Search entirely.

Google-Extended is often misunderstood. According to Google, it only governs use of your content for training Gemini models and for grounding in Gemini Apps and the Vertex AI API. It does not affect inclusion or ranking in Google Search, and it is not the switch for AI Overviews. The same logic applies to Bingbot: blocking it removes you from Bing search.

Rule 3: Know which agents ignore robots.txt

User-triggered fetchers work differently from crawlers. OpenAI states that for ChatGPT-User, robots.txt rules may not apply because the action is initiated by a user. Perplexity says Perplexity-User generally ignores robots.txt for the same reason. If you need to keep these agents out of certain areas, robots.txt is the wrong tool. Use authentication or server-side rules.

Anthropic handles this differently and lists Claude-User as a robots.txt-controllable agent. Blocking it prevents Claude from retrieving your page when a user asks about it.

Rule 4: Write one group per agent with precise paths

robots.txt is standardized as RFC 9309. A crawler follows the group that matches its user-agent most specifically and ignores the generic * group once it finds its own. That means rules in User-agent: * do not carry over to a bot that has its own group. Within a group, the most specific (longest) matching path wins.

A starting template for a site that wants AI search visibility but no model training:

# Search and retrieval: allowed
User-agent: Googlebot
User-agent: Bingbot
User-agent: OAI-SearchBot
User-agent: Claude-SearchBot
User-agent: Claude-User
User-agent: PerplexityBot
Disallow: /admin/
Disallow: /cart/

# Model training: blocked
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
User-agent: Applebot-Extended
User-agent: CCBot
Disallow: /

# Everyone else
User-agent: *
Disallow: /admin/
Disallow: /cart/

Sitemap: https://www.example.com/sitemap.xml

Two details that often go wrong: the Crawl-delay field is not supported by Google, so it will not throttle Googlebot. And a missing trailing slash changes meaning: Disallow: /private also blocks /private-offers/. If you prefer not to write the file by hand, our robots.txt generator lets you tick which AI crawlers may read your site and copy the finished blocks.

Rule 5: Treat robots.txt as a request, not access control

robots.txt tells cooperative crawlers where not to go. It does not protect anything. The file is public, so listing /internal-reports/ tells everyone where to look. Google also notes that a disallowed URL can still be indexed and shown without a snippet if other pages link to it. Consequences:

  • Protect private areas, staging sites and APIs with authentication, not robots.txt.
  • To keep a page out of search results, allow crawling and set noindex. A bot that cannot fetch the page never sees the noindex.
  • For scrapers that ignore your rules, use rate limiting or blocking at the server or CDN level.

Rule 6: Test before and after every change

A robots.txt change can silently remove your site from an AI engine. Test in three layers:

  1. Fetch the live file: curl -s https://www.example.com/robots.txt. Check that it returns status 200 and plain text, not an HTML error page or a redirect to a login.
  2. Simulate the agents: Paste your file into our AI crawler check to see which AI crawlers it fully blocks. In Google Search Console, the robots.txt report shows which version Google last fetched and any parsing errors.
  3. Check the server logs: After a week, look for the user-agents you changed. Did blocked training bots stop? Do search bots still get 200 responses on your key pages?

A bot can be blocked even when robots.txt allows it, for example by a firewall, CDN bot rule or JavaScript-only content. We describe the typical causes in AI crawler blocked despite robots.txt: 3 hidden causes, and how to make client-rendered pages readable in making JavaScript sites accessible to AI crawlers.

Rule 7: Verify bots and review the list regularly

Anyone can send a request with "GPTBot" or "CCBot" in the user-agent header. Before you whitelist or block based on logs, verify the source. Common Crawl notes that fake CCBot requests exist and recommends a reverse DNS check (genuine requests resolve to *.crawl.commoncrawl.org); it also publishes its IP ranges. Google offers the same kind of verification for Googlebot.

Operators also add new agents over time. OpenAI's overview, for example, now lists OAI-AdsBot for validating ad landing pages next to its search and training bots. Put a quarterly review in the calendar: compare your robots.txt with the operators' documentation, check your logs for unknown AI user-agents and document each change with a dated comment in the file.

If you want a full check of crawler access, rendering and citation readiness across your site, our GEO audit covers robots.txt as one part of the review.

FAQ

Which AI bots should I allow in robots.txt?

For visibility in AI search, allow Googlebot, Bingbot, OAI-SearchBot, Claude-SearchBot, Claude-User and PerplexityBot. Whether to allow training bots such as GPTBot, ClaudeBot, Google-Extended, Applebot-Extended and CCBot is a separate content-policy decision.

Does blocking Google-Extended remove my site from AI Overviews?

No. Google states that Google-Extended does not affect inclusion or ranking in Google Search. AI Overviews are part of Search and are controlled through Googlebot and snippet directives such as nosnippet.

What are the most important AI bots?

The most relevant AI user-agents come from OpenAI (GPTBot, OAI-SearchBot, ChatGPT-User), Anthropic (ClaudeBot, Claude-SearchBot, Claude-User), Perplexity (PerplexityBot, Perplexity-User), Google (Google-Extended as a control token) and Common Crawl (CCBot).

Is robots.txt still used?

Yes. robots.txt was formalized as an internet standard in RFC 9309 in 2022, and all major search and AI operators document how their crawlers read it. It remains the main way to tell crawlers which areas they may fetch.

Is robots.txt legally binding?

robots.txt is a technical convention that cooperative crawlers follow voluntarily. It is not an access control mechanism. Whether ignoring it has legal consequences depends on jurisdiction and circumstances and is a question for legal counsel, not for the file itself.

How can I tell if a visitor is an AI bot?

Check the user-agent string in your server logs, then verify it with a reverse DNS lookup or the IP ranges the operator publishes. User-agent strings alone can be faked.

To create an optional Markdown overview of your key pages, use our free llms.txt generator.

Ready for better AI visibility?

Test now for free how well your website is optimized for AI search engines.

Start Free Analysis

Share Article

Make this blog a preferred source

One click and Google will prioritise articles from geo-tool.com in Top Stories and Discover. The setting applies to your account only and can be undone at any time.

Add as a preferred source on Google

About the Author

Gorden WübbeG

AI Search Evangelist | Founder of geo-tool.com | Co-founder of famefact

Gorden Wübbe measures whether AI systems such as ChatGPT, Perplexity, Gemini, and Google AI Mode recommend a company, and shows how it earns a place on that shortlist. When OpenAI opened up GPTs, he built a GEO tool right away and secured the geo-tool.com domain. It grew into one of the first GEO tools in the German-speaking market.

As co-founder of the Berlin agency famefact, he has been building marketing tools since 2011. He tests new GEO hypotheses on his own portfolio of more than 200 domains before applying them to client projects. His conviction: rankings are no longer the goal. What matters is whether AI names a company when a buyer asks.

Husband. Father of three. Slowmad.

GEO Quick Tips
  • Structured data for AI crawlers
  • Include clear facts & statistics
  • Formulate quotable snippets
  • Integrate FAQ sections
  • Demonstrate expertise & authority