Should you block AI crawlers in robots.txt?

The right setting depends on whether you want to stop training, search access or a visit requested by a user.

Steen Stones · Reviewed

Choose which use you want to restrict before blocking an AI crawler, because the settings do different jobs. OpenAI separates training from search, while Anthropic also identifies a bot for visits requested by users. Google-Extended covers Gemini uses without affecting Google Search inclusion or ranking, so the right choice depends on the provider.

What does robots.txt control?

A robots.txt file tells search crawlers which page addresses they can access on your website. Those instructions depend on the crawler respecting them, rather than the file enforcing a block by itself.

Robots.txt cannot do the job of protecting information that needs to stay private on your website. Google recommends other controls, such as password protection, to keep private files secure from crawlers.

The instructions in robots.txt files cannot enforce crawler behavior to your site; it's up to the crawler to obey them.

developers.google.com, Robots.txt Introduction and Guide, Google Search Central, Documentation, Google for Developers (opens in a new tab)

Can you block OpenAI training without blocking ChatGPT search?

OpenAI lets you allow OAI-SearchBot for ChatGPT search while disallowing GPTBot for training crawling. These are independent choices, so a decision against training does not require closing off search access.

Blocking OAI-SearchBot has a different consequence: OpenAI says the site will not appear in ChatGPT search answers. It may still appear as a navigational link, which is an explicit exception in the documentation.

ChatGPT-User visits pages in response to actions people initiate, rather than automatically crawling the web. OpenAI says robots.txt rules may not apply, so a search crawler setting does not necessarily cover these visits.

Does blocking Google-Extended block AI Overviews?

Google-Extended does not affect whether your site is included or ranked in Google Search. Googlebot controls crawling for Search, including its AI features, so Google-Extended is not the switch for AI Overviews.

Google-Extended covers training future Gemini models and grounding answers in Gemini Apps and Vertex AI. Grounding supplies content from the Google Search index when the model answers, which makes this more than a training control.

Googlebot preferences affect Google Search and its features, which makes a broad block a wider decision. Google also provides nosnippet, data-nosnippet, max-snippet and noindex controls for limiting information shown from pages in Search.

What happens when you block Claude crawlers?

Blocking ClaudeBot tells Anthropic to exclude future material from your site from its model training datasets. ClaudeBot collects content that may contribute to training, which is a different job from answering a search.

Blocking Claude-SearchBot stops Anthropic indexing your content for search, while blocking Claude-User stops retrieval for a user query. Anthropic says these choices may reduce visibility in the relevant search results, so they need separate consideration from training.

Anthropic says its bots honour robots.txt directives, including the bots used for search and user-requested access. That differs from OpenAI, whose documentation says robots.txt rules may not apply to ChatGPT-User actions.

Does robots.txt remove a page from Google?

Blocking a page in robots.txt does not necessarily remove its address from Google results. Google says a blocked page address can still appear, although the result will not have a description.

The distinction matters when the aim is to hide information rather than simply control crawler access. Password protection is a separate control for private files; robots.txt instructions alone depend on the crawler following them.

What should you check before blocking a crawler?

Check whether the named crawler handles training, search or user-requested access before changing its rule. OpenAI separates search from training, Anthropic identifies distinct bots, and Google-Extended covers Gemini uses beyond training.

For ChatGPT search, also check whether the website host allows traffic from OpenAI’s published searchbot IP addresses. OpenAI names that alongside allowing OAI-Searchbot, so the robots.txt setting is not the whole access check.

Common questions

Can you refuse AI training while staying available to ChatGPT search?
OpenAI lets website owners allow OAI-SearchBot for search while disallowing GPTBot for training crawling.
Is Google-Extended only a training control?
No. Google-Extended also covers grounding in Gemini Apps and Grounding with Google Search on Vertex AI. It does not affect Google Search inclusion or ranking.
Is blocking ClaudeBot the same as blocking Claude search?
No. ClaudeBot collects potential training content, while Claude-SearchBot handles search indexing and Claude-User retrieves content for user requests.
Will robots.txt keep private information secure?
Robots.txt relies on crawlers obeying its instructions. Google recommends other protections, such as password-protecting private files, when information must remain secure.

AI can't recommend what it doesn't know.

Check if AI recommends you

Start free. No commitment.

Check if AI recommends you