ChatGPT-User and robots.txt: What You Can—and Cannot—Block

Separate OpenAI search, training and user-triggered requests, then apply robots, verification, firewall and access controls at the right layer.

Sonar the Answer Whale routes OAI-SearchBot, GPTBot, and ChatGPT-User through separate publisher control gates.

Direct answer: robots.txt can express different choices for OpenAI’s search and training crawlers, but it is not an access-control system. OpenAI documents OAI-SearchBot for search discovery, GPTBot for content that may support model training, and ChatGPT-User for some user-initiated requests. OpenAI says robots rules may not apply to ChatGPT-User because the visit is initiated by a user.

The practical rule is simple: use robots.txt for crawler preferences; use authentication, authorization, rate limits and firewall policy for content that must remain protected. Allowing a crawler makes a page eligible for a use—it does not guarantee retrieval, citation, inclusion or traffic.

Start with the three documented OpenAI agents

The agent name, job and control layer should stay separate
AgentDocumented jobPrimary controlBoundary
OAI-SearchBotSurface websites in ChatGPT search featuresrobots.txt plus delivery checksAccess does not guarantee a citation
GPTBotCrawl content that may be used to improve foundation modelsrobots.txt training preferenceIts rule is separate from search discovery
ChatGPT-UserSome actions initiated by a ChatGPT or Custom GPT userApplication, origin, WAF and rate-limit policyOpenAI says robots rules may not apply

OpenAI says the OAI-SearchBot and GPTBot settings are independent. A publisher can allow potential search discovery while disallowing potential training use. OpenAI also notes that one crawl may serve multiple allowed uses to avoid duplicate fetching, so a request count cannot reveal the eventual downstream purpose.

Choose the policy outcome before writing directives

A useful policy starts with the content class and business decision, not a copied blocklist. Public reporting, documentation, subscriber previews, licensed archives, account pages and private tools do not share the same access promise.

  • Public editorial pages: decide separately whether search discovery and potential training are acceptable.
  • Paid or licensed content: keep the protected version behind authorization; a voluntary crawler directive does not enforce an entitlement.
  • Account and private pages: require authentication and verify that caches, feeds and error pages do not expose protected data.
  • Public tools and APIs: use quotas, input limits, abuse controls and clear terms because user-triggered fetching can create real load.

Use the broader AI crawler field guide when your policy covers several vendors. Use the allow-or-block decision guide when the unresolved issue is commercial rather than technical.

Use the smallest robots.txt rule that matches the decision

This example allows OpenAI’s documented search crawler while disallowing its documented training crawler:

User-agent: OAI-SearchBot
Allow: /

User-agent: GPTBot
Disallow: /

That file does not make a private URL safe, and it does not guarantee appearance in ChatGPT. It expresses a crawler preference. Test it against the exact public hostname and path you serve, including CDN behavior. If you block a crawler from fetching a page, do not expect it to discover a noindex directive placed only inside that blocked page.

Do not add a ChatGPT-User rule and assume the problem is solved. When the concern is expensive endpoints, licensed documents or account data, enforce the boundary where the content is served.

Verify the request before trusting the name

A user-agent string is easy to copy. Preserve the complete user agent, source IP, timestamp, hostname, path, method, response status, bytes, request rate and authentication context. Then compare the source with OpenAI’s current published IP ranges and account for reverse proxies that may hide the original address.

Record both identity evidence and the response actually served
EvidenceQuestionFailure signal
Claimed agentDoes the full string match the documented pattern?Only a copied product name appears
NetworkIs the address inside the current official range?The source is unrelated or obscured
BehaviorAre method, rate and paths consistent with the job?Credential probing or destructive requests
Delivered responseWhat status and content class did the origin return?Protected data escaped through a cache or fallback

Our crawler identity verification workflow covers CIDR matching, failure handling and last-known-good data. When identity remains uncertain, prefer a narrow challenge, observation rule or rate limit over a permanent site-wide exception.

Run four tests, not one

  1. Policy test: fetch /robots.txt publicly and confirm the intended group and path precedence.
  2. Delivery test: request a representative public page and confirm a stable 200 response without a challenge page.
  3. Protection test: request protected content without authorization and confirm the origin denies it even when the user agent changes.
  4. Failure test: simulate unavailable identity data and excessive rate; the system should fall back to the ordinary public policy, not grant privilege.

OpenAI says its search systems may take about 24 hours to adjust after a robots change. Treat that as a product timing note, not a promise that a page will appear. Record the change time, edge tested, response and recheck time so a later observation is attributable.

Download the OpenAI crawler control decision ledger. Every sample row is marked EXAMPLE-REMOVE; replace it before using the file as a site record.

Know what the evidence cannot prove

A verified request and 200 response prove delivery to that requester at that moment. They do not prove that a later answer used the page, that a citation was earned, or that a visibility change came from the robots edit. Measure discovery, citations, explicit referrals and outcomes as separate layers. The ChatGPT SEO operating guide shows how to keep those observations separate.

Verification note: we rechecked the linked OpenAI documentation and IP-range publication on September 3, 2026. Vendor names, roles and network ranges can change; attach a retrieval date to every operational rule.

Primary documentation

Keep learning

Continue this topic

Community discussion

Discuss: ChatGPT-User and robots.txt: What You Can—and Cannot—Block

Have a question, a useful example, or a different perspective? Join the discussion, share evidence, and help other readers reach a better answer.

0 replies Moderated
No replies yet.

Be the first to ask a focused question, share a practical example, or add useful evidence.

Ask a question or join the discussion

Share evidence, a useful example, or a clear question. Be specific, stay on topic, and challenge ideas without attacking people. First-time replies may be held for moderation.