Challenge: Existing evaluation benchmarks for large language models lack annotations that justify moral classifications and focus on English constrain moral reasoning across diverse cultural settings.
Approach: They propose a multilingual benchmark dataset for evaluating moral reasoning of large language models . it includes 3,000 tweets annotated with binary hate speech labels, moral categories and rationales .
Outcome: The proposed dataset shows a misalignment between LLM outputs and human annotations in moral reasoning tasks.

Similar Papers

Moral Foundations of Large Language Models (2024.emnlp-main)

Copied to clipboard

Challenge: Moral foundations theory (MFT) is a psychological assessment tool that decomposes human moral reasoning into five factors, including care/harm, liberty/oppression, and sanctity/degradation.
Approach: They propose to use moral foundations theory to analyze whether popular LLMs have acquired a bias towards a particular set of moral values.
Outcome: The proposed model can be adversarially selected to exhibit a particular moral foundations and can affect downstream tasks.
Multi3Hate: Multimodal, Multilingual, and Multicultural Hate Speech Detection with Vision–Language Models (2025.naacl-long)

Copied to clipboard

Challenge: a new study shows that cultural background significantly affects multimodal hate speech moderation models . a limited dataset excludes multi-modal forms of hate and excludes non-English-speaking cultures . the lowest pairwise label agreement between the USA and India is due to cultural factors .
Approach: They use a multimodal and multilingual parallel hate speech dataset to examine cultural differences . they find that cultural background significantly affects multimodal hate speech annotation .
Outcome: The proposed dataset shows that cultural background significantly affects multimodal hate speech annotation.
Do Moral Judgment and Reasoning Capability of LLMs Change with Language? A Study using the Multilingual Defining Issues Test (2024.eacl-long)

Copied to clipboard

Challenge: Existing studies have shown that moral judgment depends on the language in which the dilemma is presented.
Approach: They extend the work of beyond English, to 5 new languages (Chinese, Hindi, Russian, Spanish and Swahili) and probe three LLMs that show substantial multilingual text processing and generation abilities.
Outcome: The models show substantial multilingual text processing and generation abilities.
Probing LLMs for hate speech detection: strengths and vulnerabilities (2023.findings-emnlp)

Copied to clipboard

Challenge: Recent efforts to detect hateful or toxic language using large language models have not used explanation, additional context and victim community information in the detection process.
Approach: They use different prompt variations, input information and victim community information to evaluate large language models in zero shot setting without adding any in-context examples.
Outcome: The proposed models perform significantly better when included in the pipeline than baseline models.
Adaptable Moral Stances of Large Language Models on Sexist Content: Implications for Society and Gender Discourse (2024.emnlp-main)

Copied to clipboard

Challenge: Using large language models, large language model learning has become more integrated into our daily lives, making it increasingly important to ensure they reflect ethical and equitable values.
Approach: They assess how LLMs can apply moral reasoning to both criticize and defend sexist language by evaluating their models and evaluating the moral foundations cited by them.
Outcome: The models show they can provide comprehensible and contextually relevant text for understanding diverse views on how sexism is perceived.
Are Rules Meant to be Broken? Understanding Multilingual Moral Reasoning as a Computational Pipeline with UniMoral (2025.acl-long)

Copied to clipboard

Challenge: Existing approaches to analyze moral reasoning are discordant and lack cohesion, focusing on isolated aspects of the process.
Approach: They propose a unified dataset that integrates moral dilemmas annotated with labels for action choices, ethical principles, contributing factors, and consequences, and captures diverse socio-cultural contexts.
Outcome: The proposed dataset integrates moral dilemmas annotated with labels for action choices, ethical principles, contributing factors, and consequences, along with annotators’ moral and cultural profiles.
HateBRXplain: A Benchmark Dataset with Human-Annotated Rationales for Explainable Hate Speech Detection in Brazilian Portuguese (2025.coling-main)

Copied to clipboard

Challenge: Hate speech detection systems have been developed to inhibit offensive and hateful language from being published or spread on the Web and social media.
Approach: They propose to use a Portuguese dataset to provide rationales for hate speech detection with text span annotations.
Outcome: The proposed models outperform the baselines in Portuguese and showed that they provide plausible explanations when compared to human annotations.
Speaking Multiple Languages Affects the Moral Bias of Language Models (2023.findings-acl)

Copied to clipboard

Challenge: Pre-trained multilingual language models are often better on English than other languages . however, they are trained on varying amounts of data for each language .
Approach: They apply the MORALDIRECTION framework to multilingual models and analyse their results . they find that PMLMs encode differing moral biases, but these do not correspond to cultural differences or commonalities in human opinions.
Outcome: The proposed model captures moral norms from English and imposes them on other languages.
Probing Narrative Morals: A New Character-Focused MFT Framework for Use with Large Language Models (2025.emnlp-main)

Copied to clipboard

Challenge: Existing methods to categorize moral foundations in storytelling are limited.
Approach: They propose a character-centric method to quantify moral foundations in storytelling using large language models and a novel Moral Foundations Character Action Questionnaire to validate their approach against human annotations.
Outcome: The proposed method validates against human annotations and then applies to 2,697 folktales from 55 countries.
AfriHate: A Multilingual Collection of Hate Speech and Abusive Language Datasets for African Languages (2025.naacl-long)

Copied to clipboard

Challenge: Hate speech and abusive language are global phenomena that need sociocultural background knowledge to be understood, identified, and moderated.
Approach: They propose to use a multilingual dataset to collect hate speech and abusive language in 15 African languages to help improve model performance.
Outcome: The proposed datasets are based on tweets annotated by native speakers familiar with the regional culture and show that they perform well in low-resource settings.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations