If you want to appear in AI answers, allow the search and retrieval bots in your robots.txt: Googlebot, Bingbot, OAI-SearchBot (ChatGPT search), Claude-SearchBot and PerplexityBot. Decide separately on the training bots (GPTBot, ClaudeBot, Google-Extended, Applebot-Extended, CCBot), because blocking them keeps your content out of future model training but does not remove you from search. User-triggered fetchers such as ChatGPT-User and Perplexity-User are a special case: their operators state that robots.txt may not apply to them.
The seven rules below are based on the official crawler documentation of each operator. User-agent names and purposes change, so check the linked sources before you edit your file.
AI bots at a glance: purpose and recommended action
| User-agent | Operator | Purpose (per operator docs) | Recommendation for AI visibility |
|---|---|---|---|
| Googlebot | Crawling for Google Search, including AI Overviews and AI Mode | Allow | |
| Google-Extended | Product token (not a separate crawler) that controls use of content for training Gemini models and for grounding in Gemini Apps and Vertex AI. Does not affect Google Search. | Your choice | |
| Bingbot | Microsoft | Crawling for Bing search | Allow |
| OAI-SearchBot | OpenAI | Surfaces websites in ChatGPT search features | Allow |
| GPTBot | OpenAI | Crawls content that may be used to train OpenAI foundation models | Your choice |
| ChatGPT-User | OpenAI | Fetches pages for actions a ChatGPT user initiates; robots.txt rules may not apply | Allow (blocking is unreliable) |
| Claude-SearchBot | Anthropic | Indexes content to improve Claude's search results | Allow |
| Claude-User | Anthropic | Retrieves pages when a Claude user asks a question | Allow |
| ClaudeBot | Anthropic | Collects web content that may be used for model training | Your choice |
| PerplexityBot | Perplexity | Surfaces and links websites in Perplexity search; not used for foundation model training | Allow |
| Perplexity-User | Perplexity | Visits pages for user questions; generally ignores robots.txt | Allow (blocking is unreliable) |
| Applebot-Extended | Apple | Does not crawl; controls whether Applebot data is used to train Apple foundation models | Your choice |
| CCBot | Common Crawl | Builds the open Common Crawl web archive, which third parties can reuse, including for AI training | Your choice |
Sources: OpenAI crawler overview, Anthropic crawler FAQ, Perplexity crawlers, Google common crawlers, Apple on Applebot, Common Crawl CCBot.
Rule 1: Separate search bots from training bots
"AI bot" is not one category. Search and retrieval bots decide whether your pages can be found and cited when someone asks ChatGPT, Claude or Perplexity a question today. Training bots collect content that may end up in future models. The two decisions have different consequences:
- Blocking a search bot (OAI-SearchBot, Claude-SearchBot, PerplexityBot) can remove you from that engine's cited sources. Anthropic writes that disabling Claude-SearchBot may reduce your visibility in user search results.
- Blocking a training bot (GPTBot, ClaudeBot, Google-Extended, Applebot-Extended, CCBot) is a licensing and content-policy decision. It does not take you out of the respective search products.
Many sites allow everything. Publishers with paid content often allow search bots but block training bots. Both are legitimate, as long as the choice is deliberate and not the side effect of a blanket Disallow: /.
Rule 2: Do not block Googlebot or Bingbot to "stop AI"
Google is explicit: AI Overviews and AI Mode are part of Search, and robots.txt rules for Googlebot are the control for them. If you want to limit how your content is shown in AI features, use nosnippet, data-nosnippet, max-snippet or noindex. Blocking Googlebot removes you from Search entirely.
Google-Extended is often misunderstood. According to Google, it only governs use of your content for training Gemini models and for grounding in Gemini Apps and the Vertex AI API. It does not affect inclusion or ranking in Google Search, and it is not the switch for AI Overviews. The same logic applies to Bingbot: blocking it removes you from Bing search.
Rule 3: Know which agents ignore robots.txt
User-triggered fetchers work differently from crawlers. OpenAI states that for ChatGPT-User, robots.txt rules may not apply because the action is initiated by a user. Perplexity says Perplexity-User generally ignores robots.txt for the same reason. If you need to keep these agents out of certain areas, robots.txt is the wrong tool. Use authentication or server-side rules.
Anthropic handles this differently and lists Claude-User as a robots.txt-controllable agent. Blocking it prevents Claude from retrieving your page when a user asks about it.
Rule 4: Write one group per agent with precise paths
robots.txt is standardized as RFC 9309. A crawler follows the group that matches its user-agent most specifically and ignores the generic * group once it finds its own. That means rules in User-agent: * do not carry over to a bot that has its own group. Within a group, the most specific (longest) matching path wins.
A starting template for a site that wants AI search visibility but no model training:
# Search and retrieval: allowed
User-agent: Googlebot
User-agent: Bingbot
User-agent: OAI-SearchBot
User-agent: Claude-SearchBot
User-agent: Claude-User
User-agent: PerplexityBot
Disallow: /admin/
Disallow: /cart/
# Model training: blocked
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
User-agent: Applebot-Extended
User-agent: CCBot
Disallow: /
# Everyone else
User-agent: *
Disallow: /admin/
Disallow: /cart/
Sitemap: https://www.example.com/sitemap.xml
Two details that often go wrong: the Crawl-delay field is not supported by Google, so it will not throttle Googlebot. And a missing trailing slash changes meaning: Disallow: /private also blocks /private-offers/. If you prefer not to write the file by hand, our robots.txt generator lets you tick which AI crawlers may read your site and copy the finished blocks.
Rule 5: Treat robots.txt as a request, not access control
robots.txt tells cooperative crawlers where not to go. It does not protect anything. The file is public, so listing /internal-reports/ tells everyone where to look. Google also notes that a disallowed URL can still be indexed and shown without a snippet if other pages link to it. Consequences:
- Protect private areas, staging sites and APIs with authentication, not robots.txt.
- To keep a page out of search results, allow crawling and set
noindex. A bot that cannot fetch the page never sees the noindex. - For scrapers that ignore your rules, use rate limiting or blocking at the server or CDN level.
Rule 6: Test before and after every change
A robots.txt change can silently remove your site from an AI engine. Test in three layers:
- Fetch the live file:
curl -s https://www.example.com/robots.txt. Check that it returns status 200 and plain text, not an HTML error page or a redirect to a login. - Simulate the agents: Paste your file into our AI crawler check to see which AI crawlers it fully blocks. In Google Search Console, the robots.txt report shows which version Google last fetched and any parsing errors.
- Check the server logs: After a week, look for the user-agents you changed. Did blocked training bots stop? Do search bots still get 200 responses on your key pages?
A bot can be blocked even when robots.txt allows it, for example by a firewall, CDN bot rule or JavaScript-only content. We describe the typical causes in AI crawler blocked despite robots.txt: 3 hidden causes, and how to make client-rendered pages readable in making JavaScript sites accessible to AI crawlers.
Rule 7: Verify bots and review the list regularly
Anyone can send a request with "GPTBot" or "CCBot" in the user-agent header. Before you whitelist or block based on logs, verify the source. Common Crawl notes that fake CCBot requests exist and recommends a reverse DNS check (genuine requests resolve to *.crawl.commoncrawl.org); it also publishes its IP ranges. Google offers the same kind of verification for Googlebot.
Operators also add new agents over time. OpenAI's overview, for example, now lists OAI-AdsBot for validating ad landing pages next to its search and training bots. Put a quarterly review in the calendar: compare your robots.txt with the operators' documentation, check your logs for unknown AI user-agents and document each change with a dated comment in the file.
If you want a full check of crawler access, rendering and citation readiness across your site, our GEO audit covers robots.txt as one part of the review.
FAQ
Which AI bots should I allow in robots.txt?
For visibility in AI search, allow Googlebot, Bingbot, OAI-SearchBot, Claude-SearchBot, Claude-User and PerplexityBot. Whether to allow training bots such as GPTBot, ClaudeBot, Google-Extended, Applebot-Extended and CCBot is a separate content-policy decision.
Does blocking Google-Extended remove my site from AI Overviews?
No. Google states that Google-Extended does not affect inclusion or ranking in Google Search. AI Overviews are part of Search and are controlled through Googlebot and snippet directives such as nosnippet.
What are the most important AI bots?
The most relevant AI user-agents come from OpenAI (GPTBot, OAI-SearchBot, ChatGPT-User), Anthropic (ClaudeBot, Claude-SearchBot, Claude-User), Perplexity (PerplexityBot, Perplexity-User), Google (Google-Extended as a control token) and Common Crawl (CCBot).
Is robots.txt still used?
Yes. robots.txt was formalized as an internet standard in RFC 9309 in 2022, and all major search and AI operators document how their crawlers read it. It remains the main way to tell crawlers which areas they may fetch.
Is robots.txt legally binding?
robots.txt is a technical convention that cooperative crawlers follow voluntarily. It is not an access control mechanism. Whether ignoring it has legal consequences depends on jurisdiction and circumstances and is a question for legal counsel, not for the file itself.
How can I tell if a visitor is an AI bot?
Check the user-agent string in your server logs, then verify it with a reverse DNS lookup or the IP ranges the operator publishes. User-agent strings alone can be faked.
To create an optional Markdown overview of your key pages, use our free llms.txt generator.
Ready for better AI visibility?
Test now for free how well your website is optimized for AI search engines.
Start Free AnalysisRelated GEO Topics
Share Article
About the Author
- Structured data for AI crawlers
- Include clear facts & statistics
- Formulate quotable snippets
- Integrate FAQ sections
- Demonstrate expertise & authority


