Challenge: a cross-lingual dataset captures a transnational cultural phenomenon . risky health behaviors (RHB) are often linked to complex mental health conditions .
Approach: They present the first cross-lingual dataset that captures a transnational cultural phenomenon . their dataset of more than 15,000 annotated social media posts forms the core of JiraiBench .
Outcome: The study shows that cultural context can be more influential than linguistic similarity . the study also shows that the Japanese prompts better handle Chinese content .

Similar Papers

WorldValuesBench: A Large-Scale Benchmark Dataset for Multi-Cultural Value Awareness of Language Models (2024.lrec-main)

Copied to clipboard

Challenge: a global dataset for multi-cultural value prediction task is lacking in the computer science community . a multi-culture awareness of LMs is critical to generating safe and personalized responses .
Approach: They present a global multi-cultural value prediction task using a world value survey dataset . they construct more than 20 million examples of the type "(demographic attributes, value question) answer" they show that the task is challenging for strong open and closed-source models .
Outcome: The proposed model can generate a rating response to a value question based on demographic contexts on 11.1%, 25.0%, 72.2%, and 75.0% of the questions.
LiveCLKTBench: Towards Reliable Evaluation of Cross-Lingual Knowledge Transfer in Multilingual LLMs (2026.acl-long)

Copied to clipboard

Challenge: Evaluating cross-lingual knowledge transfer in large language models is challenging, as correct answers in a target language may arise either from genuine transfer or from prior exposure during pre-training.
Approach: They propose a pipeline to isolate and measure cross-lingual knowledge transfer by identifying self-contained, time-sensitive knowledge entities from real-world domains and generating factual questions.
Outcome: The proposed pipeline analyzes multiple LLMs across five languages and shows that cross-lingual transfer is strongly influenced by linguistic distance and often asymmetric across language directions.
Toxicity Red-Teaming: Benchmarking LLM Safety in Singapore’s Low-Resource Languages (2025.emnlp-main)

Copied to clipboard

Challenge: Large language models (LLMs) have transformed natural language processing, but their safety mechanisms remain under-explored in low-resource, multilingual settings.
Approach: They propose a red-teaming approach to probe LLM vulnerabilities in Singapore's diverse linguistic context using a dataset and evaluation framework.
Outcome: The proposed framework systematically probes LLM vulnerabilities in three real-world scenarios including Singlish, Chinese, Malay, and Tamil.
SEA-SafeguardBench: Culturally Grounded Safety Benchmark for Southeast Asian Languages (2026.findings-acl)

Copied to clipboard

Challenge: Existing multilingual safety benchmarks rely on machine-translated English data, which fails to capture nuances in low-resource languages.
Approach: They propose to use a human-verified safety benchmark for Southeast Asian languages to validate their safety and cultural diversity.
Outcome: The proposed model outperforms existing models in general, in-the-wild, and content generation across eight languages and 21,640 samples across three subsets: general, and in- the-wild.
MedRiskEval: Medical Risk Evaluation Benchmark of Language Models, On the Importance of User Perspectives in Healthcare Settings (2026.eacl-industry)

Copied to clipboard

Challenge: Existing risk evaluations focused on general safety benchmarks, resulting in role-dependent vulnerabilities in real-world medical and clinical deployments.
Approach: They propose a patient-oriented dataset called PatientSafetyBench that evaluates a variety of open- and closed-source LLMs.
Outcome: The proposed benchmark examines medical risks from 466 open- and closed-source LLMs across 5 risk categories.
MMAC: A Multilingual, Multimodal Alignment Framework for Cultural Grounding Evaluation (2026.acl-long)

Copied to clipboard

Challenge: Existing models lack cultural alignment across modalities and languages . a new framework to assess cultural awareness across linguistics and languages is needed .
Approach: They propose a framework that integrates tri-modally aligned cultural benchmarks and a five-dimensional evaluation protocol to assess cross-country awareness disparities.
Outcome: The proposed framework assesses cultural awareness disparities across modalities and languages . it is the first dataset aligned at the input level across text, image, and speech .
BenchMAX: A Comprehensive Multilingual Evaluation Suite for Large Language Models (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing multilingual benchmarks focus primarily on language understanding tasks.
Approach: They develop a multi-way multilingual benchmark that measures critical capabilities of large language models across languages.
Outcome: Extensive experiments on BenchMAX reveal uneven utilization of core capabilities across languages, emphasizing the performance gaps that scaling model size alone does not resolve.
ValueBench: Towards Comprehensively Evaluating Value Orientations and Understanding of Large Language Models (2024.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) are transforming diverse fields and gaining increasing influence as human proxies.
Approach: They propose a psychometric evaluation pipeline grounded in realistic human-AI interactions to probe value orientations and novel tasks for evaluating value understanding in an open-ended value space.
Outcome: The proposed evaluation pipeline is grounded in realistic human-AI interactions and performs tasks that approximate expert conclusions in value-related extraction and generation tasks.
IndoSafety: Culturally Grounded Safety for LLMs in Indonesian Languages (2025.emnlp-main)

Copied to clipboard

Challenge: Existing safety standards are often based on direct translations from English, which overlook key aspects of local communication.
Approach: They propose a high-quality, human-verified safety evaluation dataset tailored for the Indonesian context.
Outcome: The proposed dataset covers formal and colloquial Indonesian, along with three major local languages: Javanese, Sundanese, and Minangkabau.
JMedBench: A Benchmark for Evaluating Japanese Biomedical Large Language Models (2025.coling-main)

Copied to clipboard

Challenge: Existing large language models (LLMs) focus on general domains, with fewer advancements in Japanese biomedical LLMs.
Approach: They propose a benchmark for Japanese large language models with eight LLMs across four categories and 20 Japanese biomedical datasets for comparison.
Outcome: The proposed benchmark includes eight LLMs across four categories and 20 Japanese biomedical datasets across five tasks.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations