Discrepancy Detection at the Data Level: Toward Consistent Multilingual Question Answering (2025.emnlp-main)
Copied to clipboard
Lorena Calvo-Bartolomé, Valérie Aldana, Karla Cantarero, Alonso Madroñal de Mesa, Jerónimo Arenas-García, Jordan Lee Boyd-Graber
| Challenge: | Multilingual question answering systems must ensure factual consistency across languages while also accounting for cultural variation in subjective responses. |
| Approach: | They propose a user-in-the-loop fact-checking pipeline to detect factual and cultural discrepancies in multilingual QA knowledge bases. |
| Outcome: | The proposed tool detects factual and cultural discrepancies in bilingual question answering systems. |
Similar Papers
HEAD-QA: A Healthcare Dataset for Complex Reasoning (P19-1)
Copied to clipboard
| Challenge: | Recent progress in question answering has been led by neural models, but current methods are too data intensive and weak. |
| Approach: | They propose a multi-choice question answering testbed to encourage research on complex reasoning. |
| Outcome: | The proposed dataset is useful as a benchmark for future work. |
TruthTrap: A Bilingual Benchmark for Evaluating Factually Correct Yet Misleading Information in Question Answering (2026.findings-eacl)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are increasingly used to answer factual, information-seeking questions (ISQs). |
| Approach: | They propose to use a dataset to evaluate large language models to generate human-like text on ISQs in two languages, English and Farsi, and then use it to evaluate nine LLMs. |
| Outcome: | The proposed dataset shows that accuracy drops by 25% when models encounter misleading yet factual hints. |
Towards more equitable question answering systems: How much more data do you need? (2021.acl-short)
Copied to clipboard
| Challenge: | Question answering datasets in English are relatively new, but lack of linguistic diversity in the field is a challenge. |
| Approach: | They propose to use translation and cross-lingual transfer to produce QA systems in multiple languages to improve their performance. |
| Outcome: | The proposed approaches take advantage of existing resources to produce QA systems in multiple languages. |
GuideQ: Framework for Guided Questioning for progressive informational collection and classification (2025.findings-naacl)
Copied to clipboard
| Challenge: | Using a new multilingual dataset, we examine how LLMs can be used to represent factual knowledge across languages. |
| Approach: | They propose a methodology to measure the extent of representation sharing across languages by repurposing knowledge editing methods. |
| Outcome: | The proposed model can answer a question consistently across languages and can store the answers in a shared representation for several languages. |
NativQA: Multilingual Culturally-Aligned Natural Query for LLMs (2025.findings-acl)
Copied to clipboard
Md. Arid Hasan, Maram Hasanain, Fatema Ahmad, Sahinur Rahman Laskar, Sunaya Upadhyay, Vrunda N Sukhadia, Mucahid Kutlu, Shammur Absar Chowdhury, Firoj Alam
| Challenge: | Existing frameworks for QA datasets lack regional specificity and cultural specificity. |
| Approach: | They propose a framework to quench native language QA datasets in native languages for LLM evaluation and tuning. |
| Outcome: | The proposed framework is scalable, language-independent and can be used to build culturally and regionally aligned QA datasets in native languages. |
DLAMA: A Framework for Curating Culturally Diverse Facts for Probing the Knowledge of Pretrained Language Models (2023.findings-acl)
Copied to clipboard
| Challenge: | a few benchmarking datasets have been released to evaluate the factual knowledge of pretrained language models. |
| Approach: | They propose a framework for curating factual triples from Wikidata that are culturally diverse. |
| Outcome: | The proposed framework is built of factual triples from three pairs of contrasting cultures with 78,259 triples. |
CulFiT: A Fine-grained Cultural-aware LLM Training Paradigm via Multilingual Critique Data Synthesis (2025.acl-long)
Copied to clipboard
| Challenge: | Large Language Models exhibit a specific cultural bias, neglecting values and differences of low-resource regions. |
| Approach: | They propose a culturally-aware training paradigm that leverages multilingual data and fine-grained reward modeling to enhance cultural sensitivity and inclusivity. |
| Outcome: | The proposed model achieves state-of-the-art in cultural alignment and general reasoning. |
Multilingual Summarization with Factual Consistency Evaluation (2023.findings-acl)
Copied to clipboard
| Challenge: | Abstractive summarization models generate factually inconsistent summaries, reducing their utility for real-world applications. |
| Approach: | They propose to use data filtering and controlled generation to detect hallucinations in machine generated summaries. |
| Outcome: | The proposed models detect factual inconsistencies in machine generated summaries, but they focus on English only. |
XLQA: A Benchmark for Locale-Aware Multilingual Open-Domain Question Answering (2025.emnlp-main)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have shown significant progress in Open-domain question answering (ODQA) but most evaluations focus on English and assume locale-invariant answers across languages. |
| Approach: | They propose a benchmark specifically designed for locale-sensitive multilingual ODQA that uses 3,000 English seed questions expanded to eight languages. |
| Outcome: | The proposed benchmarks are based on 3,000 English seed questions expanded to eight languages and a human-verified annotation distinguishing locale-invariant and locale-sensitive cases. |
Towards Automating Healthcare Question Answering in a Noisy Multilingual Low-Resource Setting (P19-1)
Copied to clipboard
| Challenge: | a study aims to automate a multilingual digital helpdesk service available via text messaging to pregnant and breastfeeding mothers in South Africa. |
| Approach: | They examine a multilingual digital helpdesk service available via text messaging to pregnant and breastfeeding mothers in South Africa. |
| Outcome: | The proposed model can accelerate response time by several orders of magnitude. |