Challenge: Large Language Models perform well on unseen tasks in English, but their abilities in non-English languages are less explored due to limited benchmarks and training data.
Approach: They propose to release a large dataset for context-grounded question answering in 11 major Indian languages.
Outcome: The Indic-QA Benchmark compared large datasets of large LLMs on extractive and abstractive tasks in 11 major Indian languages.

Similar Papers

Revisiting Evaluation of Question Answering Systems in Low-Resource Indic Languages: Bridging Human and Metric Alignment (2026.acl-short)

Copied to clipboard

Challenge: Evaluating Question Answering systems in low-resource Indic languages remains challenging due to the scarcity of annotated data and the lack of reliable evaluation metrics.
Approach: They propose a language-based multi-aspect evaluation framework for question answering systems . the framework integrates semantic similarity, factual completeness, numerical accuracy and contextual relevance .
Outcome: The proposed metric is evaluated across eight Indic-language QA tasks using multiple LLMs . Across all settings, it shows stronger agreement with human evaluation .
IndicGenBench: A Multilingual Benchmark to Evaluate Generation Capabilities of LLMs on Indic Languages (2024.acl-long)

Copied to clipboard

Challenge: IndicGenBench is the largest benchmark for evaluating large language models on user-facing generation tasks across a diverse set of 29 Indic languages .
Approach: They evaluate large language models on user-facing generation tasks across 29 languages . they use human curation to provide multi-way parallel evaluation data for many under-represented languages a github repository .
Outcome: IndicGenBench is the largest benchmark for evaluating LLMs on user-facing generation tasks across a diverse set of 29 Indic languages covering 13 scripts and 4 language families.
MILU: A Multi-task Indic Language Understanding Benchmark (2025.naacl-long)

Copied to clipboard

Challenge: Existing benchmarks focus on English, leaving substantial gaps in assessing LLM capabilities in low-resource and linguistically diverse languages.
Approach: They propose a multi-task indic language understanding benchmark to assess LLMs in low-resource languages.
Outcome: The new benchmark spans 8 domains and 41 subjects across 11 Indic languages, reflecting general and culturally specific knowledge.
PARIKSHA: A Large-Scale Investigation of Human-LLM Evaluator Agreement on Multilingual and Multi-Cultural Data (2024.emnlp-main)

Copied to clipboard

Challenge: Evaluation of multilingual Large Language Models is challenging due to a variety of factors including the lack of benchmarks with sufficient linguistic diversity, contamination of popular benchmarks into LLM pre-training data and lack of local, cultural nuances in translated benchmarks.
Approach: They evaluate 30 models across 10 Indic languages by conducting 90K human evaluations and 30K LLM-based evaluations.
Outcome: The proposed models perform best in most Indic languages, while the agreement drops for direct assessment especially for Bengali and Odia.
XLQA: A Benchmark for Locale-Aware Multilingual Open-Domain Question Answering (2025.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) have shown significant progress in Open-domain question answering (ODQA) but most evaluations focus on English and assume locale-invariant answers across languages.
Approach: They propose a benchmark specifically designed for locale-sensitive multilingual ODQA that uses 3,000 English seed questions expanded to eight languages.
Outcome: The proposed benchmarks are based on 3,000 English seed questions expanded to eight languages and a human-verified annotation distinguishing locale-invariant and locale-sensitive cases.
MLQA: Evaluating Cross-lingual Extractive Question Answering (2020.acl-main)

Copied to clipboard

Challenge: Question answering (QA) models have shown rapid progress enabled by the availability of large, high-quality benchmark datasets.
Approach: They present a multi-way aligned extractive QA evaluation benchmark in 7 languages . they evaluate state-of-the-art cross-lingual models and machine-translation-based baselines .
Outcome: The proposed model is based on MLQA, which has over 12K instances in english and 5K in each other language.
MKQA: A Linguistically Diverse Benchmark for Multilingual Open Domain Question Answering (2021.tacl-1)

Copied to clipboard

Challenge: Existing multilingual QA datasets lack linguistic diversity and comparable evaluation between languages.
Approach: They propose a multilingual question-answer evaluation set with 10k English queries and human translations of them into 25 additional languages and dialects.
Outcome: The proposed model is based on a multilingual knowledge questions and answers evaluation set with 26 languages.
How Accurate Are LLMs at Multi-Question Answering on Conversational Transcripts? (2025.emnlp-industry)

Copied to clipboard

Challenge: Large Language Models (LLMs) are used for question answering over long contexts . high computational costs and latency hinder the process .
Approach: They explore the capabilities of Large Language Models to answer multiple questions based on the same conversational context.
Outcome: The proposed models outperform proprietary and public models in question answering . their results show that they can be cost-effective and transparent .
Towards Leaving No Indic Language Behind: Building Monolingual Corpora, Benchmark and Models for Indic Languages (2023.acl-long)

Copied to clipboard

Challenge: Recent advances in Natural Language Understanding are driven by pretrained multilingual models, which can potentially reduce the performance gap between high-resource languages through zero-shot knowledge transfer.
Approach: They propose to create a human-supervised benchmark for Indic languages, IndicXTREME, with nine diverse NLU tasks covering 20 languages.
Outcome: The proposed model improves on the monolingual corpora, IndicCorp, and IndicBERT in Indic languages with 105 evaluation sets across languages and tasks.
NativQA: Multilingual Culturally-Aligned Natural Query for LLMs (2025.findings-acl)

Copied to clipboard

Challenge: Existing frameworks for QA datasets lack regional specificity and cultural specificity.
Approach: They propose a framework to quench native language QA datasets in native languages for LLM evaluation and tuning.
Outcome: The proposed framework is scalable, language-independent and can be used to build culturally and regionally aligned QA datasets in native languages.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations