2026 Retrieval-Augmented Answer Patents: 64 Claims-First Families
A capped 2026 review screened 344 unique US, WO and EP publications, read 173 independent-claim records and found 64 working families meeting a strict query–retrieval–answer test.
Result: four capped searches produced 400 displayed records and 344 unique US, WO and EP publications dated January 1 through August 31, 2026. Our title-and-snippet rule promoted 173 publications to independent-claim review. Sixty-four met the preregistered three-part test; 109 did not. The 64 included publications resolved to 64 working simple-family groups because the exact normalization rules found no second in-window member for any included group.
This is a bounded claims-first review, not a census of every retrieval-augmented generation patent. Each search stopped at its first 100 displayed results, and the prescreen could miss relevant claims hidden behind weak discovery text. Download the 173-row claim decision ledger (CSV) to inspect every inclusion, exclusion, claim number, original paraphrase, working family ID and official record link.
The result funnel
| Stage | Count | Meaning |
|---|---|---|
| Displayed discovery records | 400 | First 100 records from each of four frozen searches, before deduplication. |
| Unique publications | 344 | Unique US, WO or EP publication identifiers in the capped discovery set. |
| Claim-reviewed publications | 173 | Records promoted by the registered title-and-snippet signal rule. |
| Included publications | 64 | An English independent claim required the query, external retrieval and generated-answer relationship. |
| Excluded after claim review | 109 | The independent claim failed at least one registered element or lacked reviewable English claim text. |
| Working simple-family groups | 64 | Exact PCT links and exact metadata checks found no duplicate in-window included member. |
The 64 is therefore both a publication count and a working-family count for this particular set. That equality should not be generalized: a broader date range or additional jurisdictions would often surface multiple publications from the same family.
What counted as a retrieval-grounded answer claim
We wrote the boundary before discovery. An in-scope publication entered the included set only when at least one independent claim expressly required all three elements:
- An input: a user or system query, request or prompt.
- External retrieval: search or retrieval over a corpus, knowledge base, index, document collection or another data source.
- A generated natural-language answer: retrieved material is used to generate, ground, verify, rank, cite or otherwise construct the answer or response.
A title containing “RAG” was not enough. Neither was an abstract describing search. The required relationship had to appear in an independent claim. Multimodal claims could qualify, but only when a generated natural-language component was required.
How discovery was bounded
We froze the publication window, authorities, query text, displayed order and stopping rule on August 31, 2026. Google Patents served only as the discovery layer; material claim decisions were checked against USPTO, WIPO or EPO publication records.
| Query | Concept string | Reported results | Retained |
|---|---|---|---|
| Q1 | "retrieval augmented generation" OR "retrieval-augmented generation" | 1,954 | 100 |
| Q2 | ("large language model" OR "generative language model") AND (retrieve OR retrieval) AND (answer OR response) | 8,471 | 100 |
| Q3 | (grounded OR grounding) AND "language model" AND (source OR corpus OR document) | 3,080 | 100 |
| Q4 | query AND ("knowledge base" OR corpus) AND ("generating a response" OR "generated response") | 1,897 | 100 |
The system-reported totals overlap and are not added together. We retained exactly the first 100 displayed records from each query, then deduplicated publication identifiers. “400 displayed” describes the captured review set; it does not mean 400 distinct inventions.
How 344 unique publications became 173 claim reviews
Collecting every official claim document was possible but disproportionate for a capped discovery study, so we registered a prescreen amendment before individual claim-page collection. It used only the saved title and snippet;never applicant identity, perceived importance or a target result count.
A publication entered the claim-review queue when its discovery text showed either an explicit RAG term plus a response term; a model, retrieval and response term together; or the deliberately broader retrieval-plus-response combination. The third route protected older functional language that did not say “RAG” or “language model.” Records without one of those combinations were excluded at prescreen, not declared irrelevant in all possible claims.
All 173 queued publications were then reviewed against the independent-claim test. Every potential inclusion and every ambiguous exclusion received a second pass. No publication was promoted to reach a minimum count.
Results by publication authority
| Authority | Reviewed | Included | Excluded | Included share |
|---|---|---|---|---|
| US | 140 | 53 | 87 | 37.9% |
| WO / PCT | 20 | 6 | 14 | 30.0% |
| EP | 13 | 5 | 8 | 38.5% |
| Total | 173 | 64 | 109 | 37.0% |
These shares measure this prescreened queue, not an authority’s overall patent activity or claim quality. The queue was dominated by US records because the capped search results and discovery ranking surfaced more of them.
What included independent claims actually required
The included set is broader than chat-style question answering. The examples below are original claim paraphrases, not quotations or infringement mappings. Follow each publication link to inspect the official claim text.
| Publication | Applicant in discovery record | Claim | Why it passed |
|---|---|---|---|
| US12619588B2 | SAP SE | 16 | A user prompt leads to knowledge-base or vector-store retrieval; retrieved context is combined with the prompt for an LLM response. |
| US12619501B2 | Cohesity, Inc. | 1 | A natural-language query retrieves from a filtered backup-data embedding index before a language-model response. |
| US12536774B1 | Intuit Inc. | 1 | An image-plus-instruction prompt retrieves known-object information and produces an enhanced natural-language image description. |
| WO2026054990A1 | Genesys Cloud Services, Inc. | 1 | A user query retrieves knowledge-base material through keyword and semantic indexes; the LLM answers from that data. |
| WO2026080122A1 | Microsoft Technology Licensing, LLC | 1 | A test question is answered from a source index and the generated response is evaluated. |
| EP4738167A1 | Siemens AG | 1 | A user prompt retrieves relevant documents and an LLM response is generated while restricted information is concealed. |
| EP4677456A1 | Microsoft Technology Licensing, LLC | 11 | A natural-language request uses queried data items to extract requested information and generate a cited response. |
These examples show why a claims-first boundary matters. Vector retrieval, hybrid keyword-and-semantic retrieval, response evaluation, information controls, multimodal prompts and citation can all sit inside the same three-element research definition.
Why 109 claim-reviewed records were excluded
| Primary ledger reason | Count | Boundary it failed |
|---|---|---|
| No generated natural-language answer | 30 | The claim produced retrieval, ranking, code, a graph, analytics or another output without requiring an answer in natural language. |
| No query, request or prompt | 27 | The independent claim lacked the input that starts the registered query-to-answer chain. |
| No claimed external retrieval | 18 | The claim generated a response from supplied context or model state without expressly requiring a search over an external source. |
| No English independent-claim text | 8 | The official WIPO publication language did not provide reviewable English claims for this study. |
| Output not required to be natural language | 8 | A model output or result existed, but the claim did not require a natural-language component. |
| Other registered reasons | 18 | Training-only, indexing-only, evaluation-only, retrieval-only, no live claim or another narrower failure recorded in the CSV. |
The difference between “no generated natural-language answer” and “output not required to be natural language” is deliberate. The first group clearly claimed a different output. The second mentioned a model result but left its form open, so the study did not assume that it was text.
Family normalization changed no count
Publication records and patent families are different units. We joined an EP record to its PCT source when the official claim review used a linked WO publication, then tested exact priority-date-plus-applicant and exact normalized-title matches inside the included set. That reproducible pass produced 64 groups from 64 included publications and zero groups with two in-window members.
For linked EP records, the working family ID uses the PCT publication even when that WO member predates the study window. For other singleton records, the ID uses the reviewed publication. The CSV exposes both the ID and the normalization basis instead of hiding the rule.
This is a working research normalization, not a legal-family opinion. Exact metadata can miss relationships, continuations and divisionals; shared priority can also contain materially different claim sets. Anyone using the dataset for clearance, ownership or legal analysis should rebuild families from official priority chains and file histories with qualified patent counsel.
How to use the public ledger
The downloadable CSV contains all 173 claim-reviewed publications, not only the 64 inclusions. That makes it useful for auditing the boundary and finding near misses. Key fields include publication authority and number, date, applicant captured during discovery, priority date, decision, independent claim number, one primary reason, an original review note, working family ID, normalization basis and official record URL.
Start with decision, then read review_note beside the official claim. Treat blank family fields on excluded rows as intentional: family normalization was applied to the included set because the research question counts qualifying families. Do not treat the applicant field as proof of current ownership, and do not infer product use from a similar technical description.
For a follow-up study, keep the frozen ledger unchanged and add a new dated layer for extra jurisdictions, later publications, prosecution events or specialist claim review. Silent edits would make the original 2026 result impossible to reproduce.
What this review cannot establish
- Not exhaustive: all four searches exceeded the 100-result cap, and ranking determines which records entered the frozen set.
- Prescreen sensitivity: 171 unique discovery records did not proceed to claim review because their saved title and snippet lacked the registered combined signal. Relevant claim language could still exist in those documents.
- Language boundary: an English independent-claim text was required. Eight queued WO records were excluded for unavailable English claim text.
- Date boundary: the publication window ends August 31, 2026. Later publications and earlier family members fall outside the count.
- One-reviewer design: one reviewer classified the set, then repeated the included and ambiguous decisions. This reduces simple drift but is not independent double coding.
- No legal-status study: we did not use inclusion to declare a claim pending, granted, valid, enforceable, owned by a named company or practiced by a product.
The defensible statement is narrow: under this query set, cap, window, prescreen and independent-claim rule, 64 of 173 reviewed publications qualified and formed 64 working family groups.
Official sources and interpretation boundary
U.S. claim text was reviewed in official USPTO Patent Public Search publication PDFs. PCT records were reviewed in WIPO PATENTSCOPE. EP records and linked claim documents were reviewed through the EPO Publication Server. Google Patents supplied only the frozen discovery exports and was not treated as the final authority for a material claim decision.
Inclusion does not establish validity, enforceability, infringement, ownership, a licensing need, commercial importance or that any product implements the claim. A PCT publication is not a worldwide grant, and one family member’s outcome does not control another jurisdiction. The evidence-led publishing guide explains the skeptical claim-ledger method used here.
Claims-first reading
A patent family is a map of possible system design, not proof of a shipped feature
The 64-family landscape is most useful for finding repeated technical problems and vocabulary across jurisdictions.
- Read the independent claims before the abstract or product speculation.
- Group related filings into families so one invention is not counted repeatedly.
- Compare claims with public product documentation before making an implementation inference.

My takeaway: I use patents to generate testable questions. I do not use them as evidence that a current search product works exactly as described.
Keep learning
Continue this topic
Next in this topic
chat-latest Is a Moving Target: Freeze Model IDs for Reproducible Tests
Earlier in this topic
Cloudflare Content Format Chart: What 25 Requests Actually Returned
Research
Ask a question or join the discussion