Same AI Question, Different Language: A 67,200-Response Audit
An audit of GPT, Claude and Gemini found that responses to the same Ukraine-war statement bank varied across 112 language and script conditions.
The same political statement did not receive the same balance of agreement when researchers changed the language used to ask it. A September 2026 computational audit collected 67,200 responses from GPT, Claude and Gemini across 112 language or script conditions.
The audit changed language, not the underlying statement bank
The researcher created 20 English statements in ten matched pairs about the war in Ukraine. Each pair contained one Russia-leaning framing and one Ukraine-leaning framing on topics such as responsibility, sanctions, military assistance and peace conditions. Those labels describe the direction of the argument, not whether the statement is factually correct.
The statements were translated, checked automatically and frozen before they were sent to the three models. Every model-language-statement cell was scheduled for ten repetitions. The retained design covered 112 language and script conditions, three models and 20 statements, producing 67,200 valid score slots.
The exact systems matter
The study used openai/gpt-5.6-sol, anthropic/claude-sonnet-5 and google/gemini-3.8-flash through OpenRouter with provider routing pinned to OpenAI, Anthropic and Google AI Studio. Provider fallback was disabled. Collection took place in September 2026.
Each request contained one user message, no system message, a JSON response request and a 2,048-token output maximum. That is a defined API audit, not a test of every consumer chat interface or every version sold under the same product family.
Every language leaned toward Ukraine, but by different amounts
The study converts model agreement into a balance from -100 to 100. More negative values indicate stronger agreement with the Ukraine-leaning half of the statement bank. All 112 language conditions were negative, but the range was wide.
| Language | Balance | Responsible reading |
|---|---|---|
| Ukrainian | -70.49 | Strongest Ukraine-leaning balance among the retained conditions |
| Swedish | -53.87 | Shows that the variation extends beyond the belligerents’ languages |
| Russian | -44.97 | 25.52 points less negative than Ukrainian |
| Mandarin Chinese | -42.72 | Another point within a wider cross-language pattern |
| Sango | -27.31 | Least negative pooled balance in the retained set |
These numbers do not rank languages by truthfulness, safety or intelligence. They summarize agreement with this specific, curated statement bank under this specific protocol.
Country-level correlations are associations, not a causal chain
The researcher mapped official languages to countries and compared the resulting balances with three external measures. Across 35 Pew survey countries, the Pearson correlation with favorable views of Russia was 0.569. Across 183 countries, the correlation with non-support on six UN General Assembly resolutions was 0.38. Across 41 tracked donor countries, the correlation with bilateral aid to Ukraine as a share of 2021 GDP was -0.707.
The broad directions persisted when statement pairs were removed and across all three models. Country rank correlations between models ranged from 0.591 to 0.764.
That still does not show that information warfare caused the model responses. The paper presents training-corpus influence as a possible pathway. Language can also correlate with translation choices, content availability, moderation behavior, tokenization and provider-specific training or alignment decisions.
The publisher implication is a testing obligation
A multilingual publisher cannot assume that translating one page produces an equivalent answer-engine experience. Language is not only a distribution field. It can change how a model interprets a claim and how strongly it agrees.
A useful multilingual AI audit should therefore preserve:
- the source statement and the reviewed translation;
- model identifier, provider and date;
- system and user messages;
- sampling or reasoning settings when exposed;
- answer language and whether the model followed the requested format;
- claims, citations and refusals;
- a bilingual human review for semantic equivalence.
For ordinary product or news content, replace the political statement bank with a bounded set of factual, transactional and safety-sensitive questions. The goal is not to force identical wording. It is to detect material differences that could mislead users in one language.
Download a multilingual answer audit matrix
The matrix separates translation review, model configuration, answer behavior and human judgment. Its first row is an explicit example.
Download the multilingual AI answer audit matrix
Five limits belong beside the result
- The 20 statements were curated and the study was not preregistered.
- The topics are not a random sample of every possible geopolitical claim.
- Automatic checks do not establish human-validated equivalence across 112 translations.
- Ten repetitions were not selected through a prospective power calculation.
- The results apply to three fixed model versions and one API protocol in September 2026.
Editorial judgment
This is a strong audit design for demonstrating language-conditioned variation and a weak basis for attributing a single cause. Its most useful contribution is methodological: treat language as an experimental condition, not a cosmetic wrapper. The next step should combine expert translation review with controlled model settings and repeat the audit across less politically charged domains. The same discipline applies to interface comparisons in our web, app and API citation-drift protocol.
Primary source
Maxim Chupilkin, Geopolitical Divisions Across Languages in Large Language Models, arXiv preprint, September 17, 2026. The paper includes the statement bank, model identifiers, language list, aggregation method and robustness checks.
Keep learning
Continue this topic
Next in this topic
Generative Search Ads Split Relevance, CTR and Revenue
Earlier in this topic
Algebraic Retrieval: 38 Tests Passed, With One Reproduction Limit
Research
Ask a question or join the discussion