2026 Retrieval-Augmented Answer Patents: 64 Claims-First Families

A capped 2026 review screened 344 unique US, WO and EP publications, read 173 independent-claim records and found 64 working families meeting a strict query–retrieval–answer test.

Sonar pulls a claims-review filter that turns 400 displayed patent records into 173 claim reviews and 64 included working families.

Result: four capped searches produced 400 displayed records and 344 unique US, WO and EP publications dated January 1 through August 31, 2026. Our title-and-snippet rule promoted 173 publications to independent-claim review. Sixty-four met the preregistered three-part test; 109 did not. The 64 included publications resolved to 64 working simple-family groups because the exact normalization rules found no second in-window member for any included group.

This is a bounded claims-first review, not a census of every retrieval-augmented generation patent. Each search stopped at its first 100 displayed results, and the prescreen could miss relevant claims hidden behind weak discovery text. Download the 173-row claim decision ledger (CSV) to inspect every inclusion, exclusion, claim number, original paraphrase, working family ID and official record link.

The result funnel

What each count represents
StageCountMeaning
Displayed discovery records400First 100 records from each of four frozen searches, before deduplication.
Unique publications344Unique US, WO or EP publication identifiers in the capped discovery set.
Claim-reviewed publications173Records promoted by the registered title-and-snippet signal rule.
Included publications64An English independent claim required the query, external retrieval and generated-answer relationship.
Excluded after claim review109The independent claim failed at least one registered element or lacked reviewable English claim text.
Working simple-family groups64Exact PCT links and exact metadata checks found no duplicate in-window included member.

The 64 is therefore both a publication count and a working-family count for this particular set. That equality should not be generalized: a broader date range or additional jurisdictions would often surface multiple publications from the same family.

What counted as a retrieval-grounded answer claim

We wrote the boundary before discovery. An in-scope publication entered the included set only when at least one independent claim expressly required all three elements:

  1. An input: a user or system query, request or prompt.
  2. External retrieval: search or retrieval over a corpus, knowledge base, index, document collection or another data source.
  3. A generated natural-language answer: retrieved material is used to generate, ground, verify, rank, cite or otherwise construct the answer or response.

A title containing “RAG” was not enough. Neither was an abstract describing search. The required relationship had to appear in an independent claim. Multimodal claims could qualify, but only when a generated natural-language component was required.

How discovery was bounded

We froze the publication window, authorities, query text, displayed order and stopping rule on August 31, 2026. Google Patents served only as the discovery layer; material claim decisions were checked against USPTO, WIPO or EPO publication records.

All queries exceeded the 100-record review cap
QueryConcept stringReported resultsRetained
Q1"retrieval augmented generation" OR "retrieval-augmented generation"1,954100
Q2("large language model" OR "generative language model") AND (retrieve OR retrieval) AND (answer OR response)8,471100
Q3(grounded OR grounding) AND "language model" AND (source OR corpus OR document)3,080100
Q4query AND ("knowledge base" OR corpus) AND ("generating a response" OR "generated response")1,897100

The system-reported totals overlap and are not added together. We retained exactly the first 100 displayed records from each query, then deduplicated publication identifiers. “400 displayed” describes the captured review set; it does not mean 400 distinct inventions.

How 344 unique publications became 173 claim reviews

Collecting every official claim document was possible but disproportionate for a capped discovery study, so we registered a prescreen amendment before individual claim-page collection. It used only the saved title and snippet;never applicant identity, perceived importance or a target result count.

A publication entered the claim-review queue when its discovery text showed either an explicit RAG term plus a response term; a model, retrieval and response term together; or the deliberately broader retrieval-plus-response combination. The third route protected older functional language that did not say “RAG” or “language model.” Records without one of those combinations were excluded at prescreen, not declared irrelevant in all possible claims.

All 173 queued publications were then reviewed against the independent-claim test. Every potential inclusion and every ambiguous exclusion received a second pass. No publication was promoted to reach a minimum count.

Results by publication authority

Publication-level decisions before working-family grouping
AuthorityReviewedIncludedExcludedIncluded share
US140538737.9%
WO / PCT2061430.0%
EP135838.5%
Total1736410937.0%

These shares measure this prescreened queue, not an authority’s overall patent activity or claim quality. The queue was dominated by US records because the capped search results and discovery ranking surfaced more of them.

What included independent claims actually required

The included set is broader than chat-style question answering. The examples below are original claim paraphrases, not quotations or infringement mappings. Follow each publication link to inspect the official claim text.

Representative claim architectures across the three authorities
PublicationApplicant in discovery recordClaimWhy it passed
US12619588B2SAP SE16A user prompt leads to knowledge-base or vector-store retrieval; retrieved context is combined with the prompt for an LLM response.
US12619501B2Cohesity, Inc.1A natural-language query retrieves from a filtered backup-data embedding index before a language-model response.
US12536774B1Intuit Inc.1An image-plus-instruction prompt retrieves known-object information and produces an enhanced natural-language image description.
WO2026054990A1Genesys Cloud Services, Inc.1A user query retrieves knowledge-base material through keyword and semantic indexes; the LLM answers from that data.
WO2026080122A1Microsoft Technology Licensing, LLC1A test question is answered from a source index and the generated response is evaluated.
EP4738167A1Siemens AG1A user prompt retrieves relevant documents and an LLM response is generated while restricted information is concealed.
EP4677456A1Microsoft Technology Licensing, LLC11A natural-language request uses queried data items to extract requested information and generate a cited response.

These examples show why a claims-first boundary matters. Vector retrieval, hybrid keyword-and-semantic retrieval, response evaluation, information controls, multimodal prompts and citation can all sit inside the same three-element research definition.

Why 109 claim-reviewed records were excluded

One primary reason was assigned to each excluded publication
Primary ledger reasonCountBoundary it failed
No generated natural-language answer30The claim produced retrieval, ranking, code, a graph, analytics or another output without requiring an answer in natural language.
No query, request or prompt27The independent claim lacked the input that starts the registered query-to-answer chain.
No claimed external retrieval18The claim generated a response from supplied context or model state without expressly requiring a search over an external source.
No English independent-claim text8The official WIPO publication language did not provide reviewable English claims for this study.
Output not required to be natural language8A model output or result existed, but the claim did not require a natural-language component.
Other registered reasons18Training-only, indexing-only, evaluation-only, retrieval-only, no live claim or another narrower failure recorded in the CSV.

The difference between “no generated natural-language answer” and “output not required to be natural language” is deliberate. The first group clearly claimed a different output. The second mentioned a model result but left its form open, so the study did not assume that it was text.

Family normalization changed no count

Publication records and patent families are different units. We joined an EP record to its PCT source when the official claim review used a linked WO publication, then tested exact priority-date-plus-applicant and exact normalized-title matches inside the included set. That reproducible pass produced 64 groups from 64 included publications and zero groups with two in-window members.

For linked EP records, the working family ID uses the PCT publication even when that WO member predates the study window. For other singleton records, the ID uses the reviewed publication. The CSV exposes both the ID and the normalization basis instead of hiding the rule.

This is a working research normalization, not a legal-family opinion. Exact metadata can miss relationships, continuations and divisionals; shared priority can also contain materially different claim sets. Anyone using the dataset for clearance, ownership or legal analysis should rebuild families from official priority chains and file histories with qualified patent counsel.

How to use the public ledger

The downloadable CSV contains all 173 claim-reviewed publications, not only the 64 inclusions. That makes it useful for auditing the boundary and finding near misses. Key fields include publication authority and number, date, applicant captured during discovery, priority date, decision, independent claim number, one primary reason, an original review note, working family ID, normalization basis and official record URL.

Start with decision, then read review_note beside the official claim. Treat blank family fields on excluded rows as intentional: family normalization was applied to the included set because the research question counts qualifying families. Do not treat the applicant field as proof of current ownership, and do not infer product use from a similar technical description.

For a follow-up study, keep the frozen ledger unchanged and add a new dated layer for extra jurisdictions, later publications, prosecution events or specialist claim review. Silent edits would make the original 2026 result impossible to reproduce.

What this review cannot establish

  • Not exhaustive: all four searches exceeded the 100-result cap, and ranking determines which records entered the frozen set.
  • Prescreen sensitivity: 171 unique discovery records did not proceed to claim review because their saved title and snippet lacked the registered combined signal. Relevant claim language could still exist in those documents.
  • Language boundary: an English independent-claim text was required. Eight queued WO records were excluded for unavailable English claim text.
  • Date boundary: the publication window ends August 31, 2026. Later publications and earlier family members fall outside the count.
  • One-reviewer design: one reviewer classified the set, then repeated the included and ambiguous decisions. This reduces simple drift but is not independent double coding.
  • No legal-status study: we did not use inclusion to declare a claim pending, granted, valid, enforceable, owned by a named company or practiced by a product.

The defensible statement is narrow: under this query set, cap, window, prescreen and independent-claim rule, 64 of 173 reviewed publications qualified and formed 64 working family groups.

Official sources and interpretation boundary

U.S. claim text was reviewed in official USPTO Patent Public Search publication PDFs. PCT records were reviewed in WIPO PATENTSCOPE. EP records and linked claim documents were reviewed through the EPO Publication Server. Google Patents supplied only the frozen discovery exports and was not treated as the final authority for a material claim decision.

Inclusion does not establish validity, enforceability, infringement, ownership, a licensing need, commercial importance or that any product implements the claim. A PCT publication is not a worldwide grant, and one family member’s outcome does not control another jurisdiction. The evidence-led publishing guide explains the skeptical claim-ledger method used here.

Claims-first reading

A patent family is a map of possible system design, not proof of a shipped feature

The 64-family landscape is most useful for finding repeated technical problems and vocabulary across jurisdictions.

  • Read the independent claims before the abstract or product speculation.
  • Group related filings into families so one invention is not counted repeatedly.
  • Compare claims with public product documentation before making an implementation inference.
SearchEngineAnswer retrieval-augmented answer patent audit showing the 64-family result
Original release QA capture of the completed claims-first patent-family audit.

My takeaway: I use patents to generate testable questions. I do not use them as evidence that a current search product works exactly as described.

Keep learning

Continue this topic

Community discussion

Discuss: 2026 Retrieval-Augmented Answer Patents: 64 Claims-First Families

Have a question, a useful example, or a different perspective? Join the discussion, share evidence, and help other readers reach a better answer.

0 replies Moderated
No replies yet.

Be the first to ask a focused question, share a practical example, or add useful evidence.

Ask a question or join the discussion

Share evidence, a useful example, or a clear question. Be specific, stay on topic, and challenge ideas without attacking people. First-time replies may be held for moderation.