Challenge: Existing benchmarks test reasoning over culturally grounded premises, but translation-parallel benchmarks inherit English-centric scenarios.
Approach: They propose a template-first benchmark that factorizes reasoning type and cultural aspect across question languages.
Outcome: The proposed benchmark factorizes reasoning type and cultural aspect across question languages.

Similar Papers

Cultural Benchmarking of LLMs in Standard and Dialectal Arabic Dialogues (2026.acl-long)

Copied to clipboard

Challenge: Most benchmarks focus on short text snippets in Modern Standard Arabic (MSA), overlooking cultural nuances that naturally arise in dialogues.
Approach: They propose a culturally grounded conversational dataset covering 13 Arabic-speaking countries, in both Modern Standard Arabic (MSA) and each country’s respective dialect, spanning 12 daily-life topics and 54 fine-grained subtopics.
Outcome: The proposed model performs worse on all three tasks than the MSA benchmark.
mCoT: Multilingual Instruction Tuning for Reasoning Consistency in Language Models (2024.acl-long)

Copied to clipboard

Challenge: Existing models show low performance for lesser resourced languages, but they can achieve surprising performance on complex reasoning tasks in natural language processing (NLP).
Approach: They compile the first large-scale multilingual math reasoning dataset, *mCoT-MATH*, covering eleven diverse languages.
Outcome: The proposed model achieves impressive consistency across languages and comparable performance to close- and open-source models even of much larger sizes.
MMLU-ProX: A Multilingual Benchmark for Advanced Large Language Model Evaluation (2025.emnlp-main)

Copied to clipboard

Challenge: Existing large language model evaluation benchmarks focus on English, while current multilingual tasks lack parallel questions that specifically assess cross-lingual reasoning abilities.
Approach: They propose a comprehensive benchmark covering 29 languages, built on an English benchmark.
Outcome: The MMLU-ProX is a comprehensive benchmark covering 29 languages, built on an English benchmark.
MMATH: A Multilingual Benchmark for Mathematical Reasoning (2025.findings-emnlp)

Copied to clipboard

Challenge: a benchmark for multilingual complex reasoning spans 374 high-quality math problems across 10 typologically diverse languages.
Approach: They propose a benchmark for multilingual complex reasoning across 10 languages . they show reasoning in English and answering in target languages can enhance performance .
Outcome: The proposed benchmark demonstrates that models with high-quality reasoning can perform in multiple languages.
CRUXEVAL-X: A Benchmark for Multilingual Code Reasoning, Understanding and Execution (2025.acl-long)

Copied to clipboard

Challenge: Existing code benchmarks focus on code generation, while those for code reasoning are insufficient.
Approach: They propose a multi-lingual code reasoning benchmark that contains 19 programming languages and at least 600 subjects for each language.
Outcome: The proposed model trains on Python and achieves 34.4% Pass@1 in other languages, revealing the cross-language generalization of LLMs.
SeaEval for Multilingual Foundation Models: From Cross-Lingual Alignment to Cultural Reasoning (2024.naacl-long)

Copied to clipboard

Challenge: a new benchmark for multilingual foundation models is being developed . brittleness of foundation models in the dimensions of semantics and multilinguality is a key limitation .
Approach: They propose a benchmark for multilingual foundation models, SeaEval . they examine how well these models comprehend cultural practices, nuances, and values .
Outcome: The proposed model can be used to evaluate multilingual and multicultural scenarios.
MultiMUC: Multilingual Template Filling on MUC-4 (2024.eacl-long)

Copied to clipboard

Challenge: We present multilingual parallel template filling datasets for MUCs . systems were required to extract one template per incident, containing details about perpetrators, victims, weapons used .
Approach: They introduce MultiMUC, the first multilingual parallel corpus for template filling . they obtain automatic translations from a strong multilingual machine translation system .
Outcome: The proposed dataset includes translations of the classic MUC-4 template filling benchmark into Arabic, Chinese, Farsi, Korean, and Russian.
CulFiT: A Fine-grained Cultural-aware LLM Training Paradigm via Multilingual Critique Data Synthesis (2025.acl-long)

Copied to clipboard

Challenge: Large Language Models exhibit a specific cultural bias, neglecting values and differences of low-resource regions.
Approach: They propose a culturally-aware training paradigm that leverages multilingual data and fine-grained reward modeling to enhance cultural sensitivity and inclusivity.
Outcome: The proposed model achieves state-of-the-art in cultural alignment and general reasoning.
BenchMAX: A Comprehensive Multilingual Evaluation Suite for Large Language Models (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing multilingual benchmarks focus primarily on language understanding tasks.
Approach: They develop a multi-way multilingual benchmark that measures critical capabilities of large language models across languages.
Outcome: Extensive experiments on BenchMAX reveal uneven utilization of core capabilities across languages, emphasizing the performance gaps that scaling model size alone does not resolve.
NativQA: Multilingual Culturally-Aligned Natural Query for LLMs (2025.findings-acl)

Copied to clipboard

Challenge: Existing frameworks for QA datasets lack regional specificity and cultural specificity.
Approach: They propose a framework to quench native language QA datasets in native languages for LLM evaluation and tuning.
Outcome: The proposed framework is scalable, language-independent and can be used to build culturally and regionally aligned QA datasets in native languages.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations