ChatGPT Search Prompt Families: Run a 31-Field Visibility Study

Measure controlled ChatGPT Search runs with prompt families, frozen environments, repeated samples, citation review and separate referral evidence.

Sonar separates hidden prompt-rewrite currents from measurable mention, citation and referral outcomes in a controlled run.

Direct answer: an exact-prompt tracker can record what appeared in a controlled ChatGPT Search run, but it cannot treat the prompt string as the full retrieval query. OpenAI documents that ChatGPT Search may rewrite one question into one or more targeted searches. It may also use general location, optional device location and, when enabled, relevant memory. A defensible brand-visibility study therefore needs prompt families, frozen environment fields, repeated runs and separate coding for mentions, citations and referrals.

This guide supplies that measurement contract. It does not claim to reconstruct OpenAI’s candidate set, weighting or rejected sources. The downloadable 31-field prompt-family run ledger contains example rows marked EXAMPLE-REMOVE; replace them before using the file for a real study.

Why one saved prompt is not the complete search input

OpenAI’s current ChatGPT Search documentation, retrieved August 29, 2026, says the product may translate a question into multiple targeted queries and may issue more specific searches after reviewing initial results. The same page explains that approximate location can come from an IP address, precise device location is optional, and relevant memory may influence a rewritten query when memory is enabled.

Those are documented inputs to retrieval, not a public ranking formula. A tracker can preserve the visible question, environment and answer. It usually cannot observe every rewritten query, the complete candidate pool, the scoring process or the reasons a URL was rejected. Report the layer you observed instead of assigning a stable “ChatGPT rank” to hidden machinery.

Evidence layers in a ChatGPT Search visibility run
Layer Recordable evidence Boundary
Input Prompt, prompt-family ID, conversation state and run time Prior context or memory may be incomplete unless deliberately controlled
Environment Interface, plan, signed-in state, language, location controls and memory state Product experiments and server-side changes may remain unknown
Answer Visible text, sources, links, order, caveats and errors A cited URL does not expose the full candidate set or guarantee recurrence
Referral Publisher analytics for tagged visits such as utm_source=chatgpt.com A citation can occur without a click; analytics does not count every appearance

Define the reader task before writing prompts

Start with the decision a real user is trying to make. “Best accounting software” and “how to reconcile an invoice overpayment” are not interchangeable prompts about the same brand. One is a comparison task; the other is troubleshooting. Group prompts by task before comparing source visibility.

Give each family a stable ID and one explicit scope. Useful families include discovery, comparison, verification, troubleshooting, local and transactional tasks. Do not build a family by changing only punctuation or swapping one synonym. The variants should represent natural ways capable readers express the same underlying need.

Build paraphrase families instead of keyword replicas

A small family can include a direct question, a constraint-led version, a scenario and a follow-up. For example, a comparison family might ask for options, then narrow by team size, data residency or budget. Save each wording as its own prompt ID while keeping the family-level reader job constant.

Prompt-family design example
Variant What changes What stays fixed Use in analysis
Direct Plain statement of the task Reader job and eligibility constraints Baseline wording
Constraint-led Plan, geography, format or risk appears first Decision being made Tests whether a material constraint changes sources
Scenario Task is expressed through a realistic situation Required outcome Tests vocabulary and contextual variation
Follow-up Runs after a declared opening question Saved conversation sequence Measures context-dependent behavior separately

Keep prompt families small enough to review manually. More prompts create more rows, not automatically better coverage. Add a variant because it represents a real expression of the task, not because a dashboard rewards volume.

Freeze the environment that can change the result

Record the product surface, device class, signed-in state, plan, language, approximate country or city boundary, device-location setting, memory state and conversation mode. OpenAI’s current Memory FAQ explains that memory can be disabled and that Temporary Chats do not use or create memories. Do not store a precise address, personal memory content or account identifier in a public artifact. “Location allowed: city level” is normally enough for a method record.

Run local and non-local prompt families separately. OpenAI documents that location can shape relevant searches and local answers, so mixing cities in one denominator makes the result hard to interpret. If memory cannot be reliably disabled or audited in the chosen environment, label that run uncontrolled rather than pretending it matches a clean baseline.

Repeat runs without manufacturing a ranking score

Choose the number of runs before reviewing answers. Preserve errors and missing answers in the denominator. A useful statement is: “The brand appeared in 7 of 10 completed runs for this prompt family and frozen setup.” “The brand ranks at 70% in ChatGPT” invents a stable ranking unit that OpenAI does not document.

Record a run ID, exact timestamp and raw-output reference. A regenerated answer is a new run. A changed prompt, location, memory state or conversation sequence is a new condition. Do not replace a failed row with a successful rerun while keeping the original denominator.

Code mentions, citations and referrals as separate outcomes

A brand name in the prose, a clickable citation to the brand’s domain, a citation to a third-party page about the brand and a referral visit are four different observations. Store them in separate columns. This follows the same evidence boundary used in the AI visibility measurement crosswalk and the OpenAI publisher-controls guide.

Outcome coding and safe claims
Observed outcome Record Safe interpretation Do not infer
Brand mention Name, variant and answer span The answer mentioned the brand in this run That the brand’s site was retrieved
Owned-domain citation Resolved URL, source label and answer position The response linked the owned domain in this run A stable rank, endorsement or future recurrence
Third-party citation Resolved URL, publisher and relationship A third party supplied visible evidence about the brand That the owned page was eligible or considered
Referral Landing page, source/medium, time window and consent-safe outcome A tagged visit reached the site All mentions or citations generated visits

Audit whether a citation supports the adjacent claim

Open each cited URL and record its resolved destination. Check whether it directly supports the nearby statement, supports only part of it, conflicts with it or is inaccessible. Search citations can be incomplete or outdated; OpenAI explicitly tells users to open sources and verify important claims.

Keep the answer capture and claim-support review separate. A tracker that counts URLs but never opens them measures visible sourcing, not evidence quality. For a deeper review, use the evidence-led publishing guide to classify facts, observations and inferences.

Use the 31-field run ledger

The CSV’s unit of analysis is one prompt run. It joins the family, prompt, environment, visible answer, source coding, publisher-referral check, reviewer decision and method limitation. The two illustrative rows are labeled EXAMPLE-REMOVE and are not observations from SearchEngineAnswer.

Keep raw outputs in a restricted evidence store when they contain conversation context or personal data. The public ledger should contain a safe reference, a short evidence note and aggregate results. It should not expose account IDs, precise locations, memory contents, private prompts, personal identifiers or authentication data.

Compare tools only after comparing their contracts

Two vendors may use the same label for different methods. Before comparing scores, check prompt-family design, geography, signed-in state, memory handling, run frequency, repetitions, source parsing, error treatment and denominator. The SEO and GEO tool-score evaluation guide provides a broader buyer-side review.

If a vendor cannot document those fields, treat its score as a product-specific indicator. It may still be useful for directional monitoring, but it is not comparable to a separately sampled study merely because both results are percentages.

Turn changes into investigations, not guarantees

A prompt-family tracker is most useful as an alert. Investigate when an owned source disappears across repeated runs, a third party repeatedly supplies stale evidence, an important product fact is absent, or citations move to a page that answers the task more completely. Check crawl eligibility, factual clarity, page freshness and competing evidence before editing copy.

Do not paste tracked prompts into headings or promise inclusion after a rewrite. The justified outcome is a reproducible observation with a declared setup. The next measurement should use the same contract, preserve product drift and explain any change to the sample before comparing periods.

Method note: This page is a measurement protocol, not a completed ChatGPT visibility study. It was rebuilt from current OpenAI documentation retrieved August 29, 2026. SearchEngineAnswer did not receive hidden query logs or ranking data from OpenAI.

Keep learning

Continue this topic

Community discussion

Discuss: ChatGPT Search Prompt Families: Run a 31-Field Visibility Study

Have a question, a useful example, or a different perspective? Join the discussion, share evidence, and help other readers reach a better answer.

0 replies Moderated
No replies yet.

Be the first to ask a focused question, share a practical example, or add useful evidence.

Ask a question or join the discussion

Share evidence, a useful example, or a clear question. Be specific, stay on topic, and challenge ideas without attacking people. First-time replies may be held for moderation.