Can LLM Generate Culturally Relevant Commonsense QA Data? Case Study in Indonesian and Sundanese (2024.emnlp-main)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are increasingly being used to generate synthetic data for training and evaluating models. |
| Approach: | They investigate the effectiveness of using Large Language Models to generate culturally relevant commonsense QA datasets for Indonesian and Sundanese languages using both LLMs and human annotators. |
| Outcome: | The proposed model generates 4.5K questions per language, compared with 4.5k for Indonesian and 4.5km for Sundanese. |
Similar Papers
NativQA: Multilingual Culturally-Aligned Natural Query for LLMs (2025.findings-acl)
Copied to clipboard
Md. Arid Hasan, Maram Hasanain, Fatema Ahmad, Sahinur Rahman Laskar, Sunaya Upadhyay, Vrunda N Sukhadia, Mucahid Kutlu, Shammur Absar Chowdhury, Firoj Alam
| Challenge: | Existing frameworks for QA datasets lack regional specificity and cultural specificity. |
| Approach: | They propose a framework to quench native language QA datasets in native languages for LLM evaluation and tuning. |
| Outcome: | The proposed framework is scalable, language-independent and can be used to build culturally and regionally aligned QA datasets in native languages. |
CaLMQA: Exploring culturally specific long-form question answering across 23 languages (2025.acl-long)
Copied to clipboard
| Challenge: | Despite rising global usage of large language models, their ability to generate *long-form* answers to *culturally specific* questions remains unexplored in many languages. |
| Approach: | They perform the first study of textual multilingual long-form QA by creating a dataset of culturally specific questions across 23 different languages. |
| Outcome: | The results show that the best models make critical surface-level errors for many languages and their understanding of diverse cultures. |
SynDARin: Synthesising Datasets for Automated Reasoning in Low-Resource Languages (2025.coling-main)
Copied to clipboard
| Challenge: | Question Answering datasets are scarce for languages other than English due to the cost and difficulties of collection and manual annotation. |
| Approach: | They propose a method for generating and validating QA datasets for low-resource languages . they use English data as context to generate synthetic multiple-choice (MC) question-answer pairs . |
| Outcome: | The proposed method maintains quality, reduces likelihood of factual errors, and circumvents costly annotation. |
LLMs as Cultural Archives: Cultural Commonsense Knowledge Graph Extraction (2026.eacl-long)
Copied to clipboard
| Challenge: | Large language models encode rich cultural knowledge, but it remains mostly implicit and unstructured, limiting its interpretability and use. |
| Approach: | They propose an iterative framework for constructing a Cultural Commonsense Knowledge Graph using a prompt-based framework. |
| Outcome: | The proposed framework improves cultural reasoning and story generation on non-English cultures. |
mCSQA: Multilingual Commonsense Reasoning Dataset with Unified Creation Strategy by Language Models and Humans (2024.findings-acl)
Copied to clipboard
| Challenge: | Currently, multilingual datasets are created through translation, which cannot evaluate such language-specific aspects. |
| Approach: | They propose to curate a dataset for language-specific knowledge and commonsense . they propose to use multilingual commonsensiaq to leverage language models for a more efficient construction . |
| Outcome: | The proposed method reduces the creation cost by using multilingual LMs to create QAs . the proposed approach is based on the construction process of CSQA but with language models . |
Rapidly Developing High-quality Instruction Data and Evaluation Benchmark for Large Language Models with Minimal Human Effort: A Case Study on Japanese (2024.lrec-main)
Copied to clipboard
| Challenge: | Recent advances in large language models (LLMs) have aimed to refine their capacity to accurately follow human instructions and navigate intricate scenarios. |
| Approach: | They propose a method that uses a set of instructions to translate English into Japanese and then generates Japanese instruction data using GPT-4. |
| Outcome: | The proposed method outperforms Japanese-Alpaca models in the evaluation benchmarks without human references. |
Understanding the Capabilities and Limitations of Large Language Models for Cultural Commonsense (2024.naacl-long)
Copied to clipboard
| Challenge: | Large language models (LLMs) have demonstrated substantial commonsense understanding through numerous benchmark evaluations. |
| Approach: | They conduct a comprehensive examination of the capabilities and limitations of several state-of-the-art LLMs in the context of cultural commonsense tasks. |
| Outcome: | The language used to query the LLMs can impact their performance on cultural-related tasks. |
LLM-powered Data Augmentation for Enhanced Cross-lingual Performance (2023.emnlp-main)
Copied to clipboard
| Challenge: | Existing training data for multilingual commonsense reasoning datasets is limited. |
| Approach: | They propose to use large language models for data augmentation in multilingual datasets . they use Dolly-v2, StableVicuna, ChatGPT, and GPT-4 to augment three datasets. |
| Outcome: | The proposed model outperforms larger general-purpose, zero-shot models when training in smaller models. |
QA Analysis in Medical and Legal Domains: A Survey of Data Augmentation in Low-Resource Settings (2025.acl-srw)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have revolutionized natural language processing, but their success remains limited to high-resource domains. |
| Approach: | They analyze the coverage and representativeness of specialized-domain QA datasets against large-scale reference datasets. |
| Outcome: | The proposed methods and evaluations highlight the challenges faced by LLMs in low-resource domains. |
Seed-Free Synthetic Data Generation Framework for Instruction-Tuning LLMs: A Case Study in Thai (2024.acl-srw)
Copied to clipboard
| Challenge: | Xue et al., 2024) have demonstrated that large language models can perform at human level across multitudes of tasks and domains. |
| Approach: | They propose a seed-free framework for generating synthetic instruction-tuning data that incorporates fluency, diversity, and cultural context. |
| Outcome: | The proposed framework achieves competitive performance using only 5,000 instructions compared to state-of-the-art Thai LLMs trained on hundreds of thousands of instructions. |