2026 Retrieval-Augmented Search Patents: A Claims-First Research Protocol
A claims-first protocol for searching USPTO, Espacenet, the European Patent Register, and WIPO PATENTSCOPE without confusing applications, families, grants, ownership, or legal status.
Scope: This article defines a method for building a 2026 patent landscape around retrieval-augmented search and grounded answer systems. It does not publish a completed landscape, patent count, ownership ranking, infringement opinion, freedom-to-operate conclusion, or legal-status conclusion.
Patent vocabulary is unstable. A relevant claim may use terms such as information retrieval, query reformulation, vector search, semantic index, external knowledge, passage selection, reranking, grounded generation, or source attribution without saying “RAG.” Discovery must be broad; inclusion must be claims-first and reproducible.
Define the technical boundary
Write inclusion criteria before searching. One defensible boundary is an independent claim that combines a user or system query with retrieval from a corpus or external source and uses the retrieved material to construct, ground, rank, verify, or cite a machine-generated response.
Exclude records that mention search only in the background, use retrieval solely for document display, or describe generic language-model generation without a claimed retrieval relationship. Keep near-misses in an exclusion ledger with the reason; they improve query development without inflating the final dataset.
Search multiple official systems
| System | Primary use | Record to preserve |
|---|---|---|
| USPTO Patent Public Search | U.S. publications and grants | Publication, claims, dates, applicants |
| Espacenet | International discovery and families | Bibliography, family links, classifications |
| European Patent Register | EP procedural and legal events | Dated register status and documents |
| WIPO PATENTSCOPE | PCT applications and documents | International publication and claims |
Save the exact query, fields, filters, system, date, result count, and exported identifiers. Patent search interfaces and collections change; a later rerun is a new observation.
Build vocabulary and classification searches
Start with known relevant records, technical papers, product documentation, and inventor vocabulary. Extract claim phrases, synonyms, assignee names, inventors, citations, and CPC or IPC classifications. Combine text and classification searches rather than relying on one keyword.
Use forward and backward citations as discovery leads, not automatic inclusion. Search continuations, divisionals, and family members. Preserve false positives so the team can refine exclusions without repeatedly reviewing the same record.
Review independent claims first
- Read the abstract for orientation, not inclusion.
- Read every independent claim in the relevant record.
- Map the query, corpus, retrieval, selection or ranking, generation, and citation elements.
- Record which elements are required by the claim and which appear only in examples.
- Classify the claim as include, exclude, or needs specialist review.
- Save a short paraphrase and claim number; avoid copying long protected text into the article.
Do not treat a patent title or abstract as the legal boundary. Do not infer that a product practices a patent merely because product language resembles a claim.
Normalize families, owners, and status
Keep application, publication, and grant numbers distinct. Group records into families under a documented rule and retain every family member. Separate original applicant, current assignee or owner, inventors, and subsidiaries. Corporate rebranding is not proof of legal transfer.
Record legal status as an official-register observation with jurisdiction and “as of” date. Pending, granted, expired, abandoned, lapsed, and ceased records have different meanings, and family members can have different outcomes. For legal decisions, use qualified patent counsel and the official file history.
Publish a reproducible landscape
Release the search log, inclusion rules, exclusion ledger, normalized family table, claim-element rubric, verification date, and unresolved records where licensing permits it. Report both record counts and family counts. Separate descriptive clusters from legal conclusions.
Candidate identifiers found during discovery should not enter the published landscape until their documents, family relationships, applicants, and status are verified in official systems. SearchEngineAnswer has not completed that verification set for this protocol.
The evidence-led publishing guide provides the claim ledger and skeptical review used for the eventual dataset.
Ask a question or join the discussion