How to Test Tavily’s Language Boosting and Strict Filtering

Tavily's language boost and strict filter make different promises. Use this 12-row unmeasured fixture to test compliance, source overlap, availability and relevance.

Sonar the Answer Whale routes one multilingual query through a language boost channel and a strict language filter gate

Direct answer: Tavily’s language option and filter_by_language option are not interchangeable. The first boosts results in the selected language; the second excludes results that do not match it. A multilingual retrieval test must score them against different expectations.

This guide provides a 12-row unmeasured fixture for four languages and three request modes. It is a test design, not a benchmark. Run it with your own account, preserve the returned URLs and classify the actual result language before drawing conclusions.

Two controls make two different claims

Tavily’s Search API reference says the language parameter boosts results in the specified language. It accepts an ISO 639-1 code or an English language name, and Tavily recommends writing the query in the same language.

Tavily’s September 2026 product update introduced filter_by_language. When it is true, a language value is required and results that do not match are excluded. That is a compliance rule, not a stronger version of boosting.

Test each control against the promise it actually makes
ModeRequestExpected contractFailure signal
BaselineNo language parameterNo language preference assertedNot applicable
Boostlanguage=esSpanish results receive a preferenceA weak lift can be measured, but one foreign-language result is not a contract failure
Strict filterlanguage=es and filter_by_language=trueNon-Spanish results are excludedAny confidently classified non-Spanish result

Use the 12-row language-control fixture

Download the Tavily language-control test fixture. It contains matched baseline, boost and strict-filter requests for English, Spanish, French and German. Every row is marked EXAMPLE-REMOVE because no request result is included.

Each row preserves the query, query language, parameter value, strict-filter state, search depth, result limit and the fields an evaluator must complete. Keep the query meaning aligned across languages. A literal translation is not always the same information need, so have a fluent reviewer confirm the prompt before running the set.

Use a fixed account, region, date window, search depth and result limit for the matched comparison. If another parameter changes between rows, you no longer know whether the language control caused the difference.

Classify the result before calculating a rate

Do not infer language from the country-code domain. Open the result and classify the title, visible article language and any mixed-language condition. Record redirects and inaccessible pages separately.

One returned URL receives one review state
StateUse whenCounting rule
MatchThe primary result content is in the requested languageInclude in the matching-language numerator
Non-matchThe primary content is clearly in another languageCount as a strict-filter violation
MixedSubstantial content exists in more than one languageReport separately; do not force a pass
UnknownThe page is inaccessible or the language cannot be establishedExclude from the denominator and report the count

Language-identification libraries can speed the review, but short titles, proper nouns and code-heavy pages often need a human decision. Preserve the detector result and the final reviewer decision as separate fields if you automate the first pass.

Calculate three comparisons

Language compliance is the number of reviewed matching-language results divided by all reviewed results. Apply it to every mode, but use it as a hard pass-or-fail gate only for strict filtering.

Domain overlap compares the unique domains returned for baseline, boost and strict filtering. A Jaccard score can show how much the source set changed: intersection divided by union. Report the domain lists as well as the score, because a small sample can make one URL look disproportionately important.

Result availability records how many usable results remain after filtering. Perfect language compliance with very few relevant results may be a poor retrieval outcome. Relevance still needs a separate human judgment.

A language gate and a retrieval-quality review answer different questions
MeasureQuestion answeredCannot establish
Language complianceDid returned pages match the requested language?Relevance or citation quality
Domain overlapDid the source set change between modes?Which change improved the answer
Usable result countHow much retrieval remained after filtering?Coverage of the wider web
Manual relevance gradeDid the result satisfy the information need?Population-wide performance from a small fixture

Keep Safe Search and speed modes separate

Tavily documents safe_search as applying to both web and image results. It is not supported with fast or ultra-fast search depth. Do not add Safe Search to only one language row or silently switch depth to make a request pass.

If safety behavior matters, run a second fixture where Safe Search is the independent variable and use a supported depth for every row. The language fixture should remain focused on language selection.

Set decision rules before the run

  1. Fail a strict-filter row when a reviewer confirms one nonmatching result.
  2. Do not fail a boost row merely because a foreign-language result appears.
  3. Flag any row with more than 20 percent unknown results for rerun or manual inspection.
  4. Report relevance separately from language compliance.
  5. Repeat the fixture on another date before treating a difference as stable.

This guide does not report Tavily performance. It turns the documented parameter contract into a repeatable evaluation that can produce a defensible result. For a broader experiment design, use the small SEO experiment method. For source-level reporting boundaries, see the AI answer citation failure layers.

Keep learning

Continue this topic

Community discussion

Discuss: How to Test Tavily’s Language Boosting and Strict Filtering

Have a question, a useful example, or a different perspective? Join the discussion, share evidence, and help other readers reach a better answer.

0 replies Moderated
No replies yet.

Be the first to ask a focused question, share a practical example, or add useful evidence.

Ask a question or join the discussion

Share evidence, a useful example, or a clear question. Be specific, stay on topic, and challenge ideas without attacking people. First-time replies may be held for moderation.