| Topic | Managing AI crawler access |
|---|---|
| Best for | Site owners and developers |
| Skill level | Beginner to intermediate |
| Related tool | Search Readiness Auditor |
| Reading time | 5 minutes |
If an AI product cannot fetch your pages, it cannot read them when answering a question. Many sites block AI crawlers by accident, through a blanket rule, a security plugin or a template copied from another site. Others block them on purpose. Either way, the decision should be deliberate, and it starts with knowing which crawler does what.
Crawlers have different jobs
The names below are the ones most site owners meet. Their behaviour and documentation change over time, so confirm details with each company before relying on them.
- GPTBot is used by OpenAI to collect content that may be used to train its models.
- OAI-SearchBot is used by OpenAI to surface sites in ChatGPT search results.
- ChatGPT-User fetches pages when a person using ChatGPT asks it to look something up.
- ClaudeBot is Anthropic's crawler, with separate agents for user-requested fetching and search.
- PerplexityBot builds Perplexity's search index.
- Google-Extended is a control token that lets you opt content out of use for certain Google AI models. It does not change how your pages rank in Google Search.
- CCBot belongs to Common Crawl, whose open dataset is widely used in AI training.
Training and search are different decisions
The most useful distinction is between crawlers that collect training data and crawlers that power search or live answers. Blocking a training crawler stops your content from being collected for training. Blocking a search crawler can remove you from the product's answers altogether. Many businesses want to be found in AI search while limiting training use, and the separate user agents make that possible.
How rules are written
A robots.txt file lives at the root of your domain. Each group begins with a user agent and lists what it may or may not fetch. A group that blocks everything for a specific crawler looks like this:
User-agent: GPTBot Disallow: /
And a group that allows a crawler access to the whole site:
User-agent: OAI-SearchBot Allow: /
If a crawler has no group of its own, it follows the rules for the wildcard user agent. This is where accidents happen. A wildcard rule that disallows the whole site blocks every crawler that lacks its own group.
Common mistakes
- Leaving a staging rule in place after launch, so the live site disallows everything.
- Blocking all AI crawlers because of an article about training, without considering search visibility.
- Relying on a plugin or CDN setting that blocks bots you did not intend to block.
- Assuming robots.txt hides content. It is a request that well-behaved crawlers follow, not a security control. Private content needs authentication.
Checking what you currently have
Open your own domain followed by a slash and robots.txt in a browser and read the file. Look first for a wildcard group beginning with User-agent followed by an asterisk, because it applies to every crawler that does not have its own group. If it contains Disallow followed by a single slash, the whole site is closed to those crawlers. Then look for groups naming the specific AI crawlers listed above and read what each allows.
Remember that robots.txt is only one layer. A web application firewall, a bot management service or a hosting setting can also refuse requests from known crawler addresses or user agents even when your file allows them. If your robots.txt looks correct but an assistant still cannot see your pages, ask your host or security provider whether bot protection is blocking those user agents.
Deciding on training access
Whether to allow training crawlers is a business decision with no universally correct answer. Reasons to allow it include wanting your expertise represented in future models and the possibility that a more widely learned brand is mentioned more often. Reasons to block it include concern about your content being reused without compensation and a wish to control how your work is used.
Whatever you choose, make it a choice and write it down. Many sites are in their current position because a developer pasted a rule list years ago. Review the decision with whoever owns your content and legal position, and revisit it as the policies of AI companies evolve.
Testing after a change
After editing the file, load it in a browser to confirm it is saved and readable at the root of the domain. Check that it returns a normal success response and is not redirected or served as a web page. If you run several subdomains, remember that each one has its own robots.txt. Rerun a test with the Search Readiness Auditor and keep a copy of the old file in case you need to revert.
Finally, set a reminder to recheck the file after any redesign, platform migration or change of host. Those events are when rules most often change without anyone noticing.
A sensible starting policy
For a business that wants to be recommended, allow the search and user-request crawlers for the assistants your customers use. Decide separately whether to allow training crawlers, based on your own view of how your content may be used. Review the file when you launch, redesign or change hosting.
You can check your current rules with the Search Readiness Auditor, which reads your robots.txt and reports whether eight AI crawlers are blocked. Once access is sorted, the next step is making sure the pages themselves are worth citing, covered in writing content that AI assistants cite.
Frequently asked questions
Will blocking GPTBot remove me from ChatGPT answers?
Not by itself. GPTBot relates to training collection. Search visibility in ChatGPT involves OAI-SearchBot, and live lookups involve ChatGPT-User. Check the provider's current documentation.
Does blocking Google-Extended hurt my Google rankings?
Google states that this token does not affect inclusion or ranking in Google Search. It controls use of your content for certain AI models.
Is robots.txt enough to protect private pages?
No. It is a polite request, not an access control. Use authentication for anything private.
How do I know if I am already blocking AI crawlers?
Open yourdomain.com/robots.txt and read it, or run the auditor, which tests each crawler name against your rules.