MFTCXplain: A Multilingual Benchmark Dataset for Evaluating the Moral Reasoning of LLMs through Multi-hop Hate Speech Explanation (2025.findings-emnlp)
Copied to clipboard
Jackson Trager, Francielle Vargas, Diego Alves, Matteo Guida, Mikel K. Ngueajio, Ameeta Agrawal, Yalda Daryani, Farzan Karimi Malekabadi, Flor Miriam Plaza-del-Arco
| Challenge: | Existing evaluation benchmarks for large language models lack annotations that justify moral classifications and focus on English constrain moral reasoning across diverse cultural settings. |
| Approach: | They propose a multilingual benchmark dataset for evaluating moral reasoning of large language models . it includes 3,000 tweets annotated with binary hate speech labels, moral categories and rationales . |
| Outcome: | The proposed dataset shows a misalignment between LLM outputs and human annotations in moral reasoning tasks. |
Similar Papers
Moral Foundations of Large Language Models (2024.emnlp-main)
Copied to clipboard
| Challenge: | Moral foundations theory (MFT) is a psychological assessment tool that decomposes human moral reasoning into five factors, including care/harm, liberty/oppression, and sanctity/degradation. |
| Approach: | They propose to use moral foundations theory to analyze whether popular LLMs have acquired a bias towards a particular set of moral values. |
| Outcome: | The proposed model can be adversarially selected to exhibit a particular moral foundations and can affect downstream tasks. |
Multi3Hate: Multimodal, Multilingual, and Multicultural Hate Speech Detection with Vision–Language Models (2025.naacl-long)
Copied to clipboard
| Challenge: | a new study shows that cultural background significantly affects multimodal hate speech moderation models . a limited dataset excludes multi-modal forms of hate and excludes non-English-speaking cultures . the lowest pairwise label agreement between the USA and India is due to cultural factors . |
| Approach: | They use a multimodal and multilingual parallel hate speech dataset to examine cultural differences . they find that cultural background significantly affects multimodal hate speech annotation . |
| Outcome: | The proposed dataset shows that cultural background significantly affects multimodal hate speech annotation. |
Do Moral Judgment and Reasoning Capability of LLMs Change with Language? A Study using the Multilingual Defining Issues Test (2024.eacl-long)
Copied to clipboard
| Challenge: | Existing studies have shown that moral judgment depends on the language in which the dilemma is presented. |
| Approach: | They extend the work of beyond English, to 5 new languages (Chinese, Hindi, Russian, Spanish and Swahili) and probe three LLMs that show substantial multilingual text processing and generation abilities. |
| Outcome: | The models show substantial multilingual text processing and generation abilities. |
Probing LLMs for hate speech detection: strengths and vulnerabilities (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Recent efforts to detect hateful or toxic language using large language models have not used explanation, additional context and victim community information in the detection process. |
| Approach: | They use different prompt variations, input information and victim community information to evaluate large language models in zero shot setting without adding any in-context examples. |
| Outcome: | The proposed models perform significantly better when included in the pipeline than baseline models. |
Adaptable Moral Stances of Large Language Models on Sexist Content: Implications for Society and Gender Discourse (2024.emnlp-main)
Copied to clipboard
| Challenge: | Using large language models, large language model learning has become more integrated into our daily lives, making it increasingly important to ensure they reflect ethical and equitable values. |
| Approach: | They assess how LLMs can apply moral reasoning to both criticize and defend sexist language by evaluating their models and evaluating the moral foundations cited by them. |
| Outcome: | The models show they can provide comprehensible and contextually relevant text for understanding diverse views on how sexism is perceived. |
Are Rules Meant to be Broken? Understanding Multilingual Moral Reasoning as a Computational Pipeline with UniMoral (2025.acl-long)
Copied to clipboard
| Challenge: | Existing approaches to analyze moral reasoning are discordant and lack cohesion, focusing on isolated aspects of the process. |
| Approach: | They propose a unified dataset that integrates moral dilemmas annotated with labels for action choices, ethical principles, contributing factors, and consequences, and captures diverse socio-cultural contexts. |
| Outcome: | The proposed dataset integrates moral dilemmas annotated with labels for action choices, ethical principles, contributing factors, and consequences, along with annotators’ moral and cultural profiles. |
HateBRXplain: A Benchmark Dataset with Human-Annotated Rationales for Explainable Hate Speech Detection in Brazilian Portuguese (2025.coling-main)
Copied to clipboard
| Challenge: | Hate speech detection systems have been developed to inhibit offensive and hateful language from being published or spread on the Web and social media. |
| Approach: | They propose to use a Portuguese dataset to provide rationales for hate speech detection with text span annotations. |
| Outcome: | The proposed models outperform the baselines in Portuguese and showed that they provide plausible explanations when compared to human annotations. |
Speaking Multiple Languages Affects the Moral Bias of Language Models (2023.findings-acl)
Copied to clipboard
Katharina Haemmerl, Bjoern Deiseroth, Patrick Schramowski, Jindřich Libovický, Constantin Rothkopf, Alexander Fraser, Kristian Kersting
| Challenge: | Pre-trained multilingual language models are often better on English than other languages . however, they are trained on varying amounts of data for each language . |
| Approach: | They apply the MORALDIRECTION framework to multilingual models and analyse their results . they find that PMLMs encode differing moral biases, but these do not correspond to cultural differences or commonalities in human opinions. |
| Outcome: | The proposed model captures moral norms from English and imposes them on other languages. |
Probing Narrative Morals: A New Character-Focused MFT Framework for Use with Large Language Models (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing methods to categorize moral foundations in storytelling are limited. |
| Approach: | They propose a character-centric method to quantify moral foundations in storytelling using large language models and a novel Moral Foundations Character Action Questionnaire to validate their approach against human annotations. |
| Outcome: | The proposed method validates against human annotations and then applies to 2,697 folktales from 55 countries. |
AfriHate: A Multilingual Collection of Hate Speech and Abusive Language Datasets for African Languages (2025.naacl-long)
Copied to clipboard
Shamsuddeen Hassan Muhammad, Idris Abdulmumin, Abinew Ali Ayele, David Ifeoluwa Adelani, Ibrahim Said Ahmad, Saminu Mohammad Aliyu, Paul Röttger, Abigail Oppong, Andiswa Bukula, Chiamaka Ijeoma Chukwuneke, Ebrahim Chekol Jibril, Elyas Abdi Ismail, Esubalew Alemneh, Hagos Tesfahun Gebremichael, Lukman Jibril Aliyu, Meriem Beloucif, Oumaima Hourrane, Rooweither Mabuya, Salomey Osei, Samuel Rutunda, Tadesse Destaw Belay, Tadesse Kebede Guge, Tesfa Tegegne Asfaw, Lilian Diana Awuor Wanzare, Nelson Odhiambo Onyango, Seid Muhie Yimam, Nedjma Ousidhoum
| Challenge: | Hate speech and abusive language are global phenomena that need sociocultural background knowledge to be understood, identified, and moderated. |
| Approach: | They propose to use a multilingual dataset to collect hate speech and abusive language in 15 African languages to help improve model performance. |
| Outcome: | The proposed datasets are based on tweets annotated by native speakers familiar with the regional culture and show that they perform well in low-resource settings. |