How to Test Tavily’s Language Boosting and Strict Filtering
Tavily's language boost and strict filter make different promises. Use this 12-row unmeasured fixture to test compliance, source overlap, availability and relevance.
Direct answer: Tavily’s language option and filter_by_language option are not interchangeable. The first boosts results in the selected language; the second excludes results that do not match it. A multilingual retrieval test must score them against different expectations.
This guide provides a 12-row unmeasured fixture for four languages and three request modes. It is a test design, not a benchmark. Run it with your own account, preserve the returned URLs and classify the actual result language before drawing conclusions.
Two controls make two different claims
Tavily’s Search API reference says the language parameter boosts results in the specified language. It accepts an ISO 639-1 code or an English language name, and Tavily recommends writing the query in the same language.
Tavily’s September 2026 product update introduced filter_by_language. When it is true, a language value is required and results that do not match are excluded. That is a compliance rule, not a stronger version of boosting.
| Mode | Request | Expected contract | Failure signal |
|---|---|---|---|
| Baseline | No language parameter | No language preference asserted | Not applicable |
| Boost | language=es | Spanish results receive a preference | A weak lift can be measured, but one foreign-language result is not a contract failure |
| Strict filter | language=es and filter_by_language=true | Non-Spanish results are excluded | Any confidently classified non-Spanish result |
Use the 12-row language-control fixture
Download the Tavily language-control test fixture. It contains matched baseline, boost and strict-filter requests for English, Spanish, French and German. Every row is marked EXAMPLE-REMOVE because no request result is included.
Each row preserves the query, query language, parameter value, strict-filter state, search depth, result limit and the fields an evaluator must complete. Keep the query meaning aligned across languages. A literal translation is not always the same information need, so have a fluent reviewer confirm the prompt before running the set.
Use a fixed account, region, date window, search depth and result limit for the matched comparison. If another parameter changes between rows, you no longer know whether the language control caused the difference.
Classify the result before calculating a rate
Do not infer language from the country-code domain. Open the result and classify the title, visible article language and any mixed-language condition. Record redirects and inaccessible pages separately.
| State | Use when | Counting rule |
|---|---|---|
| Match | The primary result content is in the requested language | Include in the matching-language numerator |
| Non-match | The primary content is clearly in another language | Count as a strict-filter violation |
| Mixed | Substantial content exists in more than one language | Report separately; do not force a pass |
| Unknown | The page is inaccessible or the language cannot be established | Exclude from the denominator and report the count |
Language-identification libraries can speed the review, but short titles, proper nouns and code-heavy pages often need a human decision. Preserve the detector result and the final reviewer decision as separate fields if you automate the first pass.
Calculate three comparisons
Language compliance is the number of reviewed matching-language results divided by all reviewed results. Apply it to every mode, but use it as a hard pass-or-fail gate only for strict filtering.
Domain overlap compares the unique domains returned for baseline, boost and strict filtering. A Jaccard score can show how much the source set changed: intersection divided by union. Report the domain lists as well as the score, because a small sample can make one URL look disproportionately important.
Result availability records how many usable results remain after filtering. Perfect language compliance with very few relevant results may be a poor retrieval outcome. Relevance still needs a separate human judgment.
| Measure | Question answered | Cannot establish |
|---|---|---|
| Language compliance | Did returned pages match the requested language? | Relevance or citation quality |
| Domain overlap | Did the source set change between modes? | Which change improved the answer |
| Usable result count | How much retrieval remained after filtering? | Coverage of the wider web |
| Manual relevance grade | Did the result satisfy the information need? | Population-wide performance from a small fixture |
Keep Safe Search and speed modes separate
Tavily documents safe_search as applying to both web and image results. It is not supported with fast or ultra-fast search depth. Do not add Safe Search to only one language row or silently switch depth to make a request pass.
If safety behavior matters, run a second fixture where Safe Search is the independent variable and use a supported depth for every row. The language fixture should remain focused on language selection.
Set decision rules before the run
- Fail a strict-filter row when a reviewer confirms one nonmatching result.
- Do not fail a boost row merely because a foreign-language result appears.
- Flag any row with more than 20 percent unknown results for rerun or manual inspection.
- Report relevance separately from language compliance.
- Repeat the fixture on another date before treating a difference as stable.
This guide does not report Tavily performance. It turns the documented parameter contract into a repeatable evaluation that can produce a defensible result. For a broader experiment design, use the small SEO experiment method. For source-level reporting boundaries, see the AI answer citation failure layers.
Keep learning
Continue this topic
Next in this topic
SEOPress 10.3 Gives AI Assistants 27 SEO Tools, but Exposure Is Off by Default
Earlier in this topic
Cloudflare AI Search Is GA: What the New Retrieval Pricing Buys
Tools & Workflows
Ask a question or join the discussion