Challenge: Recent Large Reasoning Models (LRMs) with thinking traces have shown strong performance on English reasoning tasks.
Approach: They evaluate two leading LRMs with thinking traces on established benchmark XReasoning and propose directions for future research.
Outcome: The proposed models often revert to English or produce fragmented reasoning in other languages, revealing a substantial gap in the capability of thinking in non-English languages.

Similar Papers

The Reasoning Lingua Franca: A Double-Edged Sword for Multilingual AI (2026.eacl-short)

Copied to clipboard

Challenge: Large Reasoning Models (LRMs) are highly effective on mathematical, scientific, and other question-answering tasks.
Approach: They compare an LRM's reasoning in English to that of a multilingual question . they find that English reasoning traces exhibit a substantially higher presence of cognitive behaviors .
Outcome: The LRMs generate reasoning sequences in English, but the language of the question is not.
A Comprehensive Evaluation of Multilingual Chain-of-Thought Reasoning: Performance, Consistency, and Faithfulness Across Languages (2026.findings-eacl)

Copied to clipboard

Challenge: Recent work has examined final-answer accuracy in multilingual settings, but the behavior of thinking traces, i.e., the intermediate steps that lead to the final answer, remains underexplored.
Approach: They propose to measure language compliance, answer accuracy, and answer consistency when LRMs are explicitly instructed or prompt-hacked to think in a target language.
Outcome: The proposed model improves in English and other high-resource languages while relying on traces to varying degrees.
Large Reasoning Models Are (Not Yet) Multilingual Latent Reasoners (2026.findings-acl)

Copied to clipboard

Challenge: Recent work shows that large reasoning models arrive at the correct answer before completing textual reasoning steps, indicating the presence of latent reasoning.
Approach: They conduct a systematic investigation of multilingual latent reasoning in large reasoning models across 11 languages.
Outcome: The proposed model arrive at the correct answer before completing the reasoning steps, indicating the presence of latent reasoning.
Think Natively: Unlocking Multilingual Reasoning with Consistency-Enhanced Reinforcement Learning (2026.acl-long)

Copied to clipboard

Challenge: Current Large Reasoning Models exhibit two critical limitations when processing non-English languages: (1) They struggle to maintain input-output language consistency; (2) They generally perform poorly with wrong reasoning paths and lower answer accuracy compared to English.
Approach: They propose a language-consistency reward and a cross-lingual thinking alignment reward to improve the model's interpretability and accuracy.
Outcome: The proposed model achieves nearly 100% language consistency and superior performance on two multilingual benchmarks (MMATH and PolyMath).
A Survey of Multilingual Reasoning in Language Models (2025.findings-emnlp)

Copied to clipboard

Challenge: This survey provides the first in-depth review of multilingual reasoning in Language Models.
Approach: This survey provides the first in-depth review of multilingual reasoning in LMs.
Outcome: The present study provides the first in-depth review of multilingual reasoning in LMs.
Best-of-L: Cross-Lingual Reward Modeling for Mathematical Reasoning (2026.findings-eacl)

Copied to clipboard

Challenge: Recent studies have focused on improving reasoning ability in English models, with multilingual models receiving comparatively little attention.
Approach: They propose a framework that ranks candidate reasoning traces across languages rather than within a single language.
Outcome: The proposed framework improves accuracy by up to 10 points in English compared to using reward modeling within a single language.
Eliciting Better Multilingual Structured Reasoning from LLMs through Code (2024.acl-long)

Copied to clipboard

Challenge: xSTREET exposes a gap in base LLM performance between English and non-English reasoning tasks.
Approach: They propose a multilingual structured reasoning and explanation dataset that covers four tasks across six languages and extends the English STREET benchmark to 5 additional diverse languages.
Outcome: The proposed models show improved multilingual performance on scientific commonsense reasoning subtasks and no regression on non-reasoning tasks.
Why Do Multilingual Reasoning Gaps Emerge in Reasoning Language Models? (2026.findings-acl)

Copied to clipboard

Challenge: Recent reasoning language models (RLMs) achieve strong performance on complex reasoning tasks, yet they still exhibit a multilingual reasoning gap.
Approach: They propose a strategy that incorporates an English translation into the initial reasoning trace when an understanding failure is detected.
Outcome: The proposed strategy incorporates an English translation into the initial reasoning trace when an understanding failure is detected.
EfficientXLang: Towards Improving Token Efficiency Through Cross-Lingual Reasoning (2025.findings-emnlp)

Copied to clipboard

Challenge: Despite recent advances in Reasoning Language Models, most research focuses solely on English, even though many models are pretrained on multilingual data.
Approach: They evaluate three open-source RLMs: DeepSeek R1, Qwen 2.5, and Qwend 3 across four math datasets and seven typologically diverse languages.
Outcome: The proposed model reduces token usage and preserves accuracy even after translation into English.
Multilingual Reasoning via Self-training (2025.naacl-long)

Copied to clipboard

Challenge: Recent studies have introduced eclectic strategies to improve reasoning beyond English, but these methods are related to specific language that is not always optimal for reasoning.
Approach: They propose a modular approach that instructs models to structure reasoning passages in a different problem space and then self-refines their capabilities to deliver step-wise reasoning passage.
Outcome: The proposed approach achieves significant improvements in multilingual reasoning of various models and task, with improved reasoning consistency across languages.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations