Top-500 Retrieval Still Missed Useful Evidence

Re:CAP recovered relevant documents missed by flat retrieval, including gaps left by a 1,500-candidate ensemble. That changes how retrieval coverage should be audited.

Sonar the Answer Whale dives below a full top-500 retrieval net to recover glowing missed evidence cards

A top-500 retrieval cutoff can still leave useful evidence outside the evaluation set. Re:CAP, a new retrieval-coverage audit, recovered benchmark-relevant documents missed by conventional flat retrieval, including misses that remained after three different retrievers each contributed 500 results.

This matters for AI-search monitoring because “not retrieved” can mean two different things: the source was genuinely outside the system’s reach, or the evaluator stopped searching too early.

The search can look complete while the evidence is still below the cutoff

Traditional evaluation starts with a query, retrieves a ranked list, and checks whether known relevant documents appear within a fixed depth. That works when relevant evidence resembles the original query closely enough to rank early.

It is weaker for multi-hop questions, specialized terminology and documents connected through intermediate concepts. A first-stage result may reveal the entity, vocabulary or relationship needed to find a second document that never matched the original wording.

Re:CAP treats those first results as clues. It expands the search through discovered concepts, judges candidate relevance, and repeats the process to estimate what the flat cutoff failed to cover.

Re:CAP turns retrieval into a coverage audit

  1. Start with the original query. Run the baseline retriever and keep the ranked set.
  2. Extract expansion clues. Use retrieved documents to identify related entities, terms and paths.
  3. Retrieve beyond the original wording. Search again with the new clues rather than increasing only the first cutoff.
  4. Judge candidate relevance. Filter expanded results against the information need.
  5. Repeat until gains flatten. Track newly recovered evidence at each round and stop when marginal coverage is no longer useful.

The method is an audit layer, not a replacement for the production retriever. It asks whether the test collection and retrieval depth are sufficient to support the performance claim.

The recovered share was material across benchmarks

The authors report that Re:CAP recovered between 9% and 29% of gold documents absent from a flat BM25 top-500 list, with recovery reaching 48% on TREC-COVID in one reported setting.

More retrieval did not eliminate the issue. An ensemble that took the top 500 results from BM25, dense and hybrid retrieval, up to 1,500 candidates before deduplication, still missed 21.2% of TREC-COVID gold documents later recovered by Re:CAP.

On MuSiQue, the method improved recall by 12.9 percentage points over flat hybrid top-500 retrieval while using less than half the document budget. That result is especially relevant to multi-hop work, where an intermediate clue can change the next search.

Human review suggests many additions were not redundant

Human raters judged 78.9% of structurally distinct added documents to contain new information in the benchmark sample, with Fleiss’ kappa of 0.79 across 123 judgments. In a production sample, the reported new-information rate was 73.9% across 180 judgments.

Three repeated runs produced coverage estimates within about one percentage point. That repeatability supports using the process as an audit. It does not prove that every added document should be cited or trusted.

Translate the idea into an AI-search coverage test

Source never appears

Ask whether related terminology can reach it. Expand with entities and phrases found in early results.

Source appears only after expansion

Document the bridge term, then compare its original and expanded ranking depth.

Source enters context but is not cited

Audit claim use and visible attribution separately to locate the bottleneck.

Expanded evidence conflicts with early results

Repeat the answer with the broader set and compare which claims change.

This is a diagnostic design, not a way to infer a commercial engine’s private index. Public answer systems expose only part of their retrieval process. The test can reveal instability and missed public evidence without proving the exact internal cause.

Retrieval coverage is not citation visibility

A document can be discoverable but never selected, used or cited. It can also be cited without supporting the nearby claim. Re:CAP addresses the coverage of candidate evidence, one layer earlier than citation incidence or referral traffic.

Keep these measurements separate:

  • Reachability: can a search process find the source?
  • Retrieval: does it enter the candidate or context set?
  • Use: does it influence the answer?
  • Citation: does the interface visibly credit it?
  • Action: does a reader open the source or complete a task?

Our AI citation failure-layer model uses the same separation to prevent a downstream citation change from being misdiagnosed as an indexing change.

Use the audit to find blind spots, not manufacture certainty

Re:CAP uses model-assisted relevance judgments and benchmark collections, plus one reported production dataset. The findings show that flat retrieval can undercount relevant material in those settings. They do not establish one universal cutoff for every index, language or answer engine.

The best use is comparative: hold the query set and corpus steady, add a structured expansion process, and measure which useful documents appear only after the expansion. If the recovered set changes an answer or evaluation, the original cutoff was not a sufficient coverage boundary.

Keep learning

Continue this topic

Community discussion

Discuss: Top-500 Retrieval Still Missed Useful Evidence

Have a question, a useful example, or a different perspective? Join the discussion, share evidence, and help other readers reach a better answer.

0 replies Moderated
No replies yet.

Be the first to ask a focused question, share a practical example, or add useful evidence.

Ask a question or join the discussion

Share evidence, a useful example, or a clear question. Be specific, stay on topic, and challenge ideas without attacking people. First-time replies may be held for moderation.