Should I block AI crawlers from my website?
Short answerFor most businesses that want customers, no. Block the search crawlers (OAI-SearchBot, PerplexityBot, Claude-SearchBot, Googlebot, Bingbot) and ChatGPT, Perplexity and Google cannot recommend you. The training crawlers (GPTBot, ClaudeBot, Google-Extended, Applebot-Extended) are a separate choice: blocking them keeps your content out of future model training without removing you from AI search.
Key points
- AI companies run separate bots for training models and for search, and you can allow one while blocking the other.
- Blocking the search bots removes you from AI answers; blocking the training bots does not.
- Google-Extended has no effect on Google Search or AI Overviews, which use ordinary Googlebot.
- robots.txt is a request that reputable crawlers honour, not a lock, and your hosting or CDN may be blocking AI bots without you knowing.
The question usually arrives as a worry: AI companies are taking content without asking, so should you shut them out? The answer depends on which bot you mean, because the same company can run one crawler that learns from your pages and another that decides whether you appear when a buyer asks for a recommendation. Get that distinction right and the decision becomes simple.
What are the three kinds of AI bot?
A crawler, or bot, is software that visits web pages automatically and reads them. The AI companies now run three distinct kinds.
- Training crawlers collect pages to help build future AI models. What they read may shape what the model knows, but they do not send you visitors.
- Search crawlers build the index an AI tool searches when it answers a question with sources. These decide whether you can be named and linked in an answer.
- User-triggered fetchers visit a page because a person asked the AI tool to look at it, for example by pasting your link into a chat.
You control each one separately in robots.txt, the plain-text file at yoursite.com/robots.txt that tells crawlers which pages they may visit.
Which bot does what?
These are the names that matter, with what each company documents about them.
| Bot | Company | Purpose | If you block it |
|---|---|---|---|
| OAI-SearchBot | OpenAI | Search: surfaces sites in ChatGPT search | You will not be shown in ChatGPT search answers |
| GPTBot | OpenAI | Training future models | Your content is not used for training; search is unaffected |
| ChatGPT-User | OpenAI | Visits pages when a user asks | OpenAI says robots.txt rules may not apply, as the user started it |
| PerplexityBot | Perplexity | Search: finds and links sites in Perplexity answers | You lose citations in Perplexity; Perplexity says it is not used for training |
| Perplexity-User | Perplexity | Visits pages when a user asks | Perplexity says it generally ignores robots.txt for user requests |
| Claude-SearchBot | Anthropic | Search: improves Claude's search results | Less visibility in Claude's search answers |
| ClaudeBot | Anthropic | Training future models | Future content excluded from training |
| Claude-User | Anthropic | Visits pages when a user asks | Claude cannot fetch your pages for users |
| Googlebot | Google Search, including AI Overviews and AI Mode | You disappear from Google Search entirely | |
| Google-Extended | Controls use in Gemini model training and grounding in Gemini apps | No effect on Google Search or ranking | |
| Applebot | Apple | Siri, Spotlight and Safari search features | You leave Apple's search features |
| Applebot-Extended | Apple | Controls use of Applebot's pages for training Apple's AI models | Pages can still appear in Apple search results |
The sources are each company's own documentation: OpenAI, Perplexity, Google and Apple, with Anthropic's in its help centre. OpenAI is explicit that the settings are independent: you can allow OAI-SearchBot to appear in search while disallowing GPTBot.
Blocking a search bot opts you out of being recommended by that tool. Blocking a training bot does not.
What does Google-Extended actually control?
This is the most misunderstood line in robots.txt. Google says Google-Extended manages whether content it crawls may be used to train future Gemini models and for grounding, meaning supplying facts at answer time, in Gemini apps and Google's Vertex AI service. Google also states it does not affect inclusion in Google Search and is not a ranking signal.
AI Overviews and AI Mode are part of Google Search, so they use ordinary Googlebot. Blocking Google-Extended does not remove you from them. Google's only controls for those features are the standard ones: nosnippet, max-snippet and noindex. Each of those also limits how you appear in ordinary results, so there is no clean way to stay in Google Search while leaving AI Overviews. For a business that wants enquiries, you would not want one anyway. The guide on how to appear in Google AI Overviews covers the other direction.
Should you block the training bots?
This is a genuine choice, and reasonable businesses make it differently.
Block them if your content is the product: a publisher, a course provider, a photographer, a firm whose paid reports or original research are what clients buy. Letting models learn from that work for free may undercut you.
Leave them open if your website exists to win customers for a service. Your pages describe what you do, where and for whom. There is little to protect, and there is a plausible upside: models learn what they know about the world partly from training data, so being in it may help a tool know your business exists even when it is not searching. That benefit is unproven, and no AI company documents it, so treat it as a possibility rather than a reason.
Either way, the decision does not affect whether AI search tools can recommend you, as long as the search bots stay allowed.
If you do decide to block training, remember the bots that are not run by AI companies. Common Crawl's crawler, CCBot, builds a free public archive of the web that many AI models have been trained on. Blocking GPTBot and ClaudeBot while leaving CCBot open only does half the job.
Sample business: Hartwell Joinery. Its website holds service descriptions, prices, a process page and photographs of finished kitchens. Nothing on it is sold as content, and every buyer question it answers is one Hartwell wants repeated. The sensible setting is to allow every bot, search and training alike. A design studio selling paid kitchen-planning guides would weigh it differently: allow the search bots, block GPTBot, ClaudeBot, Google-Extended, Applebot-Extended and CCBot, and keep the paid guides behind a login.
What can robots.txt not do?
- It is a request, not a lock. Reputable crawlers follow it. Badly behaved ones ignore it, and some disguise themselves. A real block needs a firewall or your CDN, the network service many sites sit behind.
- User-triggered visits may bypass it. OpenAI and Perplexity both say their user-initiated fetchers may not follow robots.txt.
- It is not retroactive. Blocking a training bot today does not remove anything collected before.
- It takes time. OpenAI says changes can take around 24 hours to reach its systems.
The reverse problem is more common than people expect. Some hosting, security and CDN services now offer settings that block AI bots, and some have switched them on by default for new sites. OpenAI notes that your host or CDN must allow traffic from its published search bot addresses for you to be included. A site can have a perfect robots.txt and still be invisible to ChatGPT because of a firewall setting nobody remembers enabling.
What setup suits most businesses?
For a business that wants enquiries, the simplest answer is to allow everything, as this site does. Its robots.txt lists each AI bot by name and allows it, so there is no doubt about intent. If you want to stay in AI search but keep your content out of training, use these lines:
| Line in robots.txt | Effect |
|---|---|
User-agent: OAI-SearchBot, User-agent: PerplexityBot, User-agent: Claude-SearchBot, then Allow: / | Search bots may read the whole site |
User-agent: GPTBot, User-agent: ClaudeBot, User-agent: Google-Extended, User-agent: Applebot-Extended, then Disallow: / | Training bots are asked to stay out |
Sitemap: followed by your sitemap address | Points every crawler to your full page list |
Each User-agent line goes on its own line, followed by the rule that applies to that group. Leave Googlebot and Bingbot alone unless you have a specific reason. Microsoft Copilot is built on Bing, and Bing is widely understood to be among the search providers ChatGPT draws on, so Bingbot matters to AI visibility more than its name suggests.
How do you check what your site does today?
- Open
yoursite.com/robots.txtin a browser. Look forDisallow: /under any search bot named above, or underUser-agent: *, which applies to every bot without its own rule. - Ask whoever manages your hosting or CDN whether any "block AI bots" or bot-fighting setting is on.
- Check the robots.txt report in Google Search Console to confirm Google can read the file. If Google cannot see your site at all, start with why your business is not showing on Google.
- Run the free test in the AI search visibility guide. If ChatGPT or Perplexity cannot describe your business at all, access is the first thing to rule out, and getting cited by Perplexity explains what comes next.
If you are tidying robots.txt anyway, adding an llms.txt file takes another few minutes. It will not open any door robots.txt has closed, but it gives the tools you have let in a cleaner summary.
Straight answers.
The follow-up questions owners ask most.
- If I block GPTBot, will ChatGPT stop recommending me?
- No. OpenAI states its settings are independent: GPTBot covers training, while OAI-SearchBot decides whether you can appear in ChatGPT search answers. You can block one and allow the other.
- Can I stay in Google Search but stay out of AI Overviews?
- Not cleanly. Google's controls for AI features are the same snippet and indexing controls used for ordinary results, so limiting AI Overviews also limits your normal listing.
- Will blocking AI crawlers remove my content from models already trained?
- No. A robots.txt change applies to future crawling. It does not remove what was collected before.
- How quickly does a robots.txt change take effect?
- It varies by company. OpenAI says its systems can take around 24 hours to adjust after you update robots.txt.