Challenge: Large Language Models (LLMs) are increasingly being used to generate synthetic data for training and evaluating models.
Approach: They investigate the effectiveness of using Large Language Models to generate culturally relevant commonsense QA datasets for Indonesian and Sundanese languages using both LLMs and human annotators.
Outcome: The proposed model generates 4.5K questions per language, compared with 4.5k for Indonesian and 4.5km for Sundanese.

Similar Papers

NativQA: Multilingual Culturally-Aligned Natural Query for LLMs (2025.findings-acl)

Copied to clipboard

Challenge: Existing frameworks for QA datasets lack regional specificity and cultural specificity.
Approach: They propose a framework to quench native language QA datasets in native languages for LLM evaluation and tuning.
Outcome: The proposed framework is scalable, language-independent and can be used to build culturally and regionally aligned QA datasets in native languages.
CaLMQA: Exploring culturally specific long-form question answering across 23 languages (2025.acl-long)

Copied to clipboard

Challenge: Despite rising global usage of large language models, their ability to generate *long-form* answers to *culturally specific* questions remains unexplored in many languages.
Approach: They perform the first study of textual multilingual long-form QA by creating a dataset of culturally specific questions across 23 different languages.
Outcome: The results show that the best models make critical surface-level errors for many languages and their understanding of diverse cultures.
SynDARin: Synthesising Datasets for Automated Reasoning in Low-Resource Languages (2025.coling-main)

Copied to clipboard

Challenge: Question Answering datasets are scarce for languages other than English due to the cost and difficulties of collection and manual annotation.
Approach: They propose a method for generating and validating QA datasets for low-resource languages . they use English data as context to generate synthetic multiple-choice (MC) question-answer pairs .
Outcome: The proposed method maintains quality, reduces likelihood of factual errors, and circumvents costly annotation.
LLMs as Cultural Archives: Cultural Commonsense Knowledge Graph Extraction (2026.eacl-long)

Copied to clipboard

Challenge: Large language models encode rich cultural knowledge, but it remains mostly implicit and unstructured, limiting its interpretability and use.
Approach: They propose an iterative framework for constructing a Cultural Commonsense Knowledge Graph using a prompt-based framework.
Outcome: The proposed framework improves cultural reasoning and story generation on non-English cultures.
mCSQA: Multilingual Commonsense Reasoning Dataset with Unified Creation Strategy by Language Models and Humans (2024.findings-acl)

Copied to clipboard

Challenge: Currently, multilingual datasets are created through translation, which cannot evaluate such language-specific aspects.
Approach: They propose to curate a dataset for language-specific knowledge and commonsense . they propose to use multilingual commonsensiaq to leverage language models for a more efficient construction .
Outcome: The proposed method reduces the creation cost by using multilingual LMs to create QAs . the proposed approach is based on the construction process of CSQA but with language models .
Rapidly Developing High-quality Instruction Data and Evaluation Benchmark for Large Language Models with Minimal Human Effort: A Case Study on Japanese (2024.lrec-main)

Copied to clipboard

Challenge: Recent advances in large language models (LLMs) have aimed to refine their capacity to accurately follow human instructions and navigate intricate scenarios.
Approach: They propose a method that uses a set of instructions to translate English into Japanese and then generates Japanese instruction data using GPT-4.
Outcome: The proposed method outperforms Japanese-Alpaca models in the evaluation benchmarks without human references.
Understanding the Capabilities and Limitations of Large Language Models for Cultural Commonsense (2024.naacl-long)

Copied to clipboard

Challenge: Large language models (LLMs) have demonstrated substantial commonsense understanding through numerous benchmark evaluations.
Approach: They conduct a comprehensive examination of the capabilities and limitations of several state-of-the-art LLMs in the context of cultural commonsense tasks.
Outcome: The language used to query the LLMs can impact their performance on cultural-related tasks.
LLM-powered Data Augmentation for Enhanced Cross-lingual Performance (2023.emnlp-main)

Copied to clipboard

Challenge: Existing training data for multilingual commonsense reasoning datasets is limited.
Approach: They propose to use large language models for data augmentation in multilingual datasets . they use Dolly-v2, StableVicuna, ChatGPT, and GPT-4 to augment three datasets.
Outcome: The proposed model outperforms larger general-purpose, zero-shot models when training in smaller models.
QA Analysis in Medical and Legal Domains: A Survey of Data Augmentation in Low-Resource Settings (2025.acl-srw)

Copied to clipboard

Challenge: Large Language Models (LLMs) have revolutionized natural language processing, but their success remains limited to high-resource domains.
Approach: They analyze the coverage and representativeness of specialized-domain QA datasets against large-scale reference datasets.
Outcome: The proposed methods and evaluations highlight the challenges faced by LLMs in low-resource domains.
Seed-Free Synthetic Data Generation Framework for Instruction-Tuning LLMs: A Case Study in Thai (2024.acl-srw)

Copied to clipboard

Challenge: Xue et al., 2024) have demonstrated that large language models can perform at human level across multitudes of tasks and domains.
Approach: They propose a seed-free framework for generating synthetic instruction-tuning data that incorporates fluency, diversity, and cultural context.
Outcome: The proposed framework achieves competitive performance using only 5,000 instructions compared to state-of-the-art Thai LLMs trained on hundreds of thousands of instructions.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations