OpenAI Publisher Controls: Crawl Access, Citations and Referrals Are Different Outcomes
Choose OpenAI publisher controls by outcome: search discovery, training collection, title-and-link suppression, referrals, or authenticated access.
OpenAI publisher controls cover different outcomes. OAI-SearchBot concerns search discovery. GPTBot concerns potential model training. A user-directed fetch is another event. Links or citations in ChatGPT are appearances, while utm_source=chatgpt.com can identify some referral visits.
Do not report one of these signals as if it proved all the others. A crawler request does not prove a citation, and a citation does not prove a visit.
Four separate outcomes
| Signal | What it can show | What it cannot prove |
|---|---|---|
| OAI-SearchBot access | Declared permission for search discovery requests | Indexing, citation, referral, or ranking |
| GPTBot access | Declared permission relating to model training collection | Search discovery or a ChatGPT citation |
| Link or citation | The page was referenced in a particular output | A stable rank, click, conversion, or complete retrieval history |
| Tagged referral | A visit with the documented ChatGPT source parameter | Every ChatGPT visit or the answer that produced it |
Search and training are separate
OpenAI says publishers should allow OAI-SearchBot if they want their sites included in ChatGPT search summaries and snippets. GPTBot is the separately named crawler for potential training use. A publisher can therefore document different preferences for search discovery and training.
OpenAI also notes that ChatGPT Atlas may discover a link and title through another search or discovery source even when OAI-SearchBot is blocked. Publishers that do not want this title-and-link appearance should use a noindex meta tag. The crawler must be allowed to fetch the page to read that directive, so a robots disallow and a readable noindex are not interchangeable controls.
Keep private material behind authentication. As the crawler policy guide explains, robots.txt is not access authorization.
Test the control that matches the outcome
A publisher can want a page available to people but absent from ChatGPT Atlas title-and-link results. In that case, the page must remain fetchable long enough for the relevant crawler to see noindex. If the material is private, authentication is the stronger boundary because robots and index directives are not access controls.
- Choose the exact outcome: no training collection, no search summary, no title-and-link appearance, or no public access.
- Apply the matching control to a representative URL and save the robots, meta, or response-header state.
- Fetch the public response without browser cookies and verify that the intended crawler can read the directive.
- Recheck the named ChatGPT surface after a reasonable recrawl window. A valid control does not promise immediate removal from cached or third-party discovery.
Measure access with logs
- Save the robots.txt version and deployment timestamp.
- Verify the client using OpenAI’s current documentation rather than trusting the user-agent string alone.
- Record requested URL, timestamp, response code, bytes, and cache status.
- Separate OAI-SearchBot, GPTBot, and user-directed activity where the evidence permits.
- Watch crawl volume and errors after a policy change.
- Keep the result at the access layer: “verified requests occurred” or “no verified requests were observed.”
Absence from a limited log window is not proof that a page can never appear. Cached results, other discovery providers, filters, retention, and low request volume can all affect the observation.
Measure referrals with boundaries
OpenAI says ChatGPT referral URLs include utm_source=chatgpt.com. Create a documented analytics segment for that parameter and preserve the raw landing page, timestamp, campaign fields, consent state, and on-site outcome. Keep an additional referrer-based segment for diagnosis, but do not assume it captures every visit.
Report sessions or visits according to the analytics system’s definition. Do not rename them citations. A visitor may follow a link without the publisher being able to reconstruct the exact answer, prompt, or retrieval path.
Use the five-layer measurement crosswalk to keep access, appearance, referral, and outcome rows separate.
A publisher control card
- Desired outcome: search discovery, no training collection, public user access.
- Controls: OAI-SearchBot allowed, GPTBot disallowed, public pages indexable, private pages authenticated.
- Evidence: dated robots file, verified logs, index directives, analytics segment.
- Limit: no claim that access guarantees a citation or that tagged referrals represent every ChatGPT visit.
- Review trigger: OpenAI documentation change, crawler incident, site migration, or measurement change.
This card makes the policy auditable without promising control over selection. It also gives legal, editorial, analytics, and infrastructure owners a shared record of what the site intends.
Primary documentation
Keep learning
Continue this topic
Next in this topic
ChatGPT Search Prompt Families: Run a 31-Field Visibility Study
Earlier in this topic
llms.txt and Google Search: What the File Does Not Change
AEO & AI Search
Ask a question or join the discussion