ChatGPT Long Pastes: Build a 31-Field Evidence Register
Preserve source versions, composer state, location references, extraction checks and privacy boundaries when ChatGPT converts a long paste into an attachment.
Direct answer: when ChatGPT converts a long paste into an attachment, preserve the source version, character and byte counts, conversion state, location scheme and extraction checks as one evidence record. The content may be the same, but the handling path is different.
OpenAI’s release notes currently document automatic attachment conversion above 10,000 characters across ChatGPT plans. That threshold is a dated product rule, not a durable file standard. This guide supplies a 31-field register for workflows that need quotations, comparisons or auditable source boundaries.
What changed and when
As accessed August 29, 2026, OpenAI’s ChatGPT release notes document three relevant dates. On March 25, 2026, pastes over 5,000 characters began converting for Plus, Pro and Business. On June 22, OpenAI raised the threshold to 10,000 characters and extended the behavior to Free and Go. On August 4, it added Enterprise and Education, making the 10,000-character behavior available across plans.
The notes also document a “Show in text field” control that moves the content back into the composer. Record the access date, plan and client surface because product behavior can change after this article is updated.
Conversion is not evidence loss by itself
The release note describes a composer-handling change. It does not state that ChatGPT discards the pasted content, changes its meaning or guarantees identical processing on every client. A visible attachment therefore proves the composer selected the attachment path; it does not prove complete extraction or a specific retention outcome.
Test those adjacent questions separately. Preserve the source, inspect representative passages and record the workspace policy that applies. Do not turn a product-interface observation into a claim about model attention, privacy or long-term storage.
Map the four evidence layers
| Layer | Record | Question answered |
|---|---|---|
| Source | Version, hash, character count and byte count | What content did the operator intend to submit? |
| Composer | Client, plan, conversion result and attachment alias | How did ChatGPT package the input? |
| Extraction check | Location scheme, sampled references and transformations | Could selected source passages be located correctly? |
| Response | Prompt, answer evidence reference and review decision | What did ChatGPT return for this case? |
A response can look correct even when a source boundary is unclear. A complete source file can also exist while the prompt asks an ambiguous question. The register prevents one successful answer from standing in for all four checks.
Count characters and bytes before submission
Record both the character count and byte count before pasting. OpenAI’s note defines the trigger in characters, while engineering pipelines often measure bytes or tokens. Non-ASCII text can make those values diverge, and token counts depend on the tokenizer and model context.
Keep the original line endings and encoding in the source artifact. If you normalize the file, produce a new version and hash. Do not edit the source after a response and keep the old case ID; the evidence chain should show exactly which version entered the workflow.
Give the source a stable location scheme
Long evidence needs references that survive the attachment path. For plain text, create line numbers in a separate review copy. For HTML or Markdown, use stable section IDs plus line references. For a PDF, record the file hash and page number, while noting that printed and PDF page numbers can differ.
Do not ask ChatGPT to invent line numbers for an unnumbered source and then treat them as evidence. Supply the location scheme before submission, or map response passages back to the canonical source during human review.
Use a three-point extraction check
| Check | Prompt action | Failure signal |
|---|---|---|
| Beginning | Locate a distinctive early reference | Wrong version or leading content omitted |
| Middle | Locate a distinctive central reference | Partial extraction or section confusion |
| End | Locate a distinctive final reference | Truncation or incomplete source handling |
| Boundary | State the first and last available references | Unclear source extent |
This is a smoke check, not a measurement of complete recall. If the task depends on every row, clause or citation, use a deterministic parser or dedicated document workflow and compare its output with the source.
Record “Show in text field” as a separate path
OpenAI documents that users can move the attachment back into the message with “Show in text field.” Treat that as a new composer state. Record whether the control was checked, the resulting character count when observable, and whether formatting or references changed.
Do not assume the attachment and direct-paste paths are operationally identical because the visible words match. If the distinction matters to your workflow, run a paired case with the same source version and prompt, then compare observable outputs without claiming the model’s internal process.
Protect private and regulated material
The long-paste release note does not define the privacy, training, retention or compliance rules for every workspace. Apply the current policy for the actual plan and organization before submitting material. Record the workspace type, Temporary Chat state and data-control state only at the level the review needs.
Do not place confidential prompts, personal identifiers, credentials, unpublished contracts, health records or regulated data in the public register. Use aliases and private evidence references. If the task cannot be completed without exposing material outside its approved system, stop the workflow rather than masking the risk with a generic disclaimer.
Compare source claims, not answer fluency
A smooth summary can omit a qualifier, merge two sections or attribute a statement to the wrong source. Define the claims that matter, locate their supporting spans in the canonical document and compare the response with those spans.
| Code | Meaning | Decision |
|---|---|---|
| Direct | The cited source span supports the claim as written | Accept with the recorded reference |
| Qualified | The source supports a narrower claim | Edit the answer to preserve the qualifier |
| Conflicting | The source contradicts the response | Reject or correct |
| Absent | No supporting span was located | Hold for review |
The citation-ready passage test provides a complementary method for keeping claims and supporting passages close enough to inspect.
Run a bounded paired test
- Freeze one source version and compute its SHA-256 hash.
- Create a numbered review copy without changing the canonical source.
- Record the plan, workspace, client surface and documented threshold on the access date.
- Submit one case through the automatic attachment path.
- When available, create a paired case through “Show in text field.”
- Run beginning, middle and end extraction checks.
- Review a declared set of consequential claims against the source.
- Record omissions, transformations, privacy boundary and workflow decision.
Do not call a two-case check an accuracy benchmark. It verifies that a specific evidence workflow survived two recorded composer paths under a dated product state.
Download the 31-field register
Download the ChatGPT attachment evidence register. One row represents one source submission through one recorded composer path. Remove both EXAMPLE-REMOVE rows and replace every placeholder before using it.
The register connects source identity, composer behavior, sampled extraction evidence, privacy state and the final workflow decision. It deliberately stores aliases and evidence references rather than raw private material. Validate that every row has 31 columns.
Set re-audit and stop conditions
Repeat the attachment check when OpenAI changes the documented threshold, composer controls, supported file behavior or plan availability, or when your organization changes workspace policy. Start a new case when the source version, client surface, plan or prompt changes; do not overwrite the earlier record.
Stop the workflow when the source hash is missing, the attachment cannot be mapped to the intended version, a consequential section fails the extraction sample, or the workspace is not approved for the material. A fluent response does not waive any of those conditions.
Choose the workflow by evidence risk
Use the attachment path for ordinary source work after the sampled extraction checks pass. Add deterministic parsing and claim-level review when missing one clause, row or qualification would change the decision. Do not submit material that the workspace is not authorized to process.
Use the evidence-led publishing guide to preserve claim-source relationships, and the retrieval provenance guide when a later workflow combines uploaded evidence with live search or tools. The 10,000-character threshold explains when the composer changes paths; the register determines whether your evidence process remains auditable.
Keep learning
Continue this topic
Next in this topic
GPT-5.6 Long-Context Pricing: What Changes Above 272K Tokens
Earlier in this topic
Claude inference_geo: Separate Inference Location from Data at Rest
Tools & Workflows
Ask a question or join the discussion