Challenge: Existing multilingual vision-language (VL) benchmarks typically only cover a handful of languages, underscoring the need for evaluation data for low-resource languages.
Approach: They propose a multilingual vision-language benchmark that evaluates cross-modal and text-only topical matching across 205 languages.
Outcome: The proposed model performs better in cross-modal and text-only topical matching in lower-resource languages than the most multilingual benchmarks.

Similar Papers

VLURes: Benchmarking Long-Text Grounding and Cross-Lingual Robustness in Vision Language Models (2026.findings-acl)

Copied to clipboard

Challenge: ***VLURes** provides a practical testbed for long-text grounding and multilingual robustness in web-realistic agent settings.
Approach: They propose a multilingual benchmark for evaluating vision-language models under long-text grounding.
Outcome: ***VLURes** provides a testbed for long-text grounding and multilingual robustness in web-realistic agent settings.
M5 – A Diverse Benchmark to Assess the Performance of Large Multimodal Models Across Multilingual and Multicultural Vision-Language Tasks (2024.findings-emnlp)

Copied to clipboard

Challenge: Recent advances in Large Language Models and their multimodal counterparts have shown significant performance disparities across different languages and cultural contexts.
Approach: They propose to evaluate LLMs on diverse vision-language tasks within a multilingual and multicultural context using M5 benchmark.
Outcome: The proposed benchmarks highlight task-agnostic performance disparities between languages and cultural contexts.
Babel-ImageNet: Massively Multilingual Evaluation of Vision-and-Language Representations (2024.acl-long)

Copied to clipboard

Challenge: Vision-and-language models with separate encoders for each modality are limited in availability.
Approach: They propose a multilingual benchmark that offers (partial) translations of ImageNet labels to 100 languages, built without machine translation or manual annotation.
Outcome: The proposed model outperforms models on English and low-resource languages.
Centurio: On Drivers of Multilingual Ability of Large Vision-Language Model (2025.acl-long)

Copied to clipboard

Challenge: Existing models for large vision-language tasks are trained on English data, which makes them struggle to understand non-English input and fail to generate output in the desired target language.
Approach: They conduct multi-stage experiments on 13 vision-language tasks and 43 languages . they find that one can include as many as 100 training languages simultaneously with as little as 25-50% of non-English data .
Outcome: The proposed model outperforms existing models in 14 tasks and 56 languages.
Lost in Translation: Do LVLM Judges Generalize Across Languages? (2026.findings-acl)

Copied to clipboard

Challenge: MM-JudgeBench is the first large-scale benchmark for multilingual and multimodal judge model evaluation.
Approach: They propose a multilingual benchmark for multilingual and multimodal judge model evaluation that includes over 60K pairwise preference instances spanning 25 typologically diverse languages.
Outcome: The proposed benchmark includes over 60K pairwise preference instances spanning 25 languages.
Towards Cross-Lingual Explanation of Artwork in Large-scale Vision Language Models (2025.findings-naacl)

Copied to clipboard

Challenge: LVLMs are increasingly capable of responding in multiple languages . however, there is a lack of evaluation tools for LVLs that handle multiple languages.
Approach: They used an extended dataset in multiple languages to evaluate LVLMs' ability to generate explanations in multiple language combinations.
Outcome: The proposed dataset in multiple languages evaluates LVLMs' ability to generate explanations in other languages.
GlotEval: A Test Suite for Massively Multilingual Evaluation of Large Language Models (2025.emnlp-demos)

Copied to clipboard

Challenge: Existing evaluation frameworks focus on English and a handful of high-resource languages, thereby overlooking the realistic performance of large language models in multilingual and lower-resourced scenarios.
Approach: They propose a unified and lightweight framework that integrates 27 benchmarks under a standard ISO 639-3 language identifier system to enable seamless incorporation of new benchmarks.
Outcome: The proposed framework integrates 27 benchmarks under a standard ISO 639-3 language identifier system, allowing for seamless incorporation of new benchmarks.
BenchMAX: A Comprehensive Multilingual Evaluation Suite for Large Language Models (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing multilingual benchmarks focus primarily on language understanding tasks.
Approach: They develop a multi-way multilingual benchmark that measures critical capabilities of large language models across languages.
Outcome: Extensive experiments on BenchMAX reveal uneven utilization of core capabilities across languages, emphasizing the performance gaps that scaling model size alone does not resolve.
MMLU-ProX: A Multilingual Benchmark for Advanced Large Language Model Evaluation (2025.emnlp-main)

Copied to clipboard

Challenge: Existing large language model evaluation benchmarks focus on English, while current multilingual tasks lack parallel questions that specifically assess cross-lingual reasoning abilities.
Approach: They propose a comprehensive benchmark covering 29 languages, built on an English benchmark.
Outcome: The MMLU-ProX is a comprehensive benchmark covering 29 languages, built on an English benchmark.
An Examination of the Compositionality of Large Generative Vision-Language Models (2024.naacl-long)

Copied to clipboard

Challenge: Recent studies have focused on the compositionality of vision-language models (VLMs) however, the performance of GVLMs in multimodal compositional reasoning remains under-explored.
Approach: They propose a syntactical bias score to quantify GVLMs' syntaktical bias . they propose 'SADE' task to assess GVLs's robustness against inclination toward syntical correctness.
Outcome: The proposed benchmarks are based on evaluation metrics and current benchmarks.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations