| Challenge: | Moral foundations theory (MFT) is a psychological assessment tool that decomposes human moral reasoning into five factors, including care/harm, liberty/oppression, and sanctity/degradation. |
| Approach: | They propose to use moral foundations theory to analyze whether popular LLMs have acquired a bias towards a particular set of moral values. |
| Outcome: | The proposed model can be adversarially selected to exhibit a particular moral foundations and can affect downstream tasks. |
Similar Papers
Probing Narrative Morals: A New Character-Focused MFT Framework for Use with Large Language Models (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing methods to categorize moral foundations in storytelling are limited. |
| Approach: | They propose a character-centric method to quantify moral foundations in storytelling using large language models and a novel Moral Foundations Character Action Questionnaire to validate their approach against human annotations. |
| Outcome: | The proposed method validates against human annotations and then applies to 2,697 folktales from 55 countries. |
MFTCXplain: A Multilingual Benchmark Dataset for Evaluating the Moral Reasoning of LLMs through Multi-hop Hate Speech Explanation (2025.findings-emnlp)
Copied to clipboard
Jackson Trager, Francielle Vargas, Diego Alves, Matteo Guida, Mikel K. Ngueajio, Ameeta Agrawal, Yalda Daryani, Farzan Karimi Malekabadi, Flor Miriam Plaza-del-Arco
| Challenge: | Existing evaluation benchmarks for large language models lack annotations that justify moral classifications and focus on English constrain moral reasoning across diverse cultural settings. |
| Approach: | They propose a multilingual benchmark dataset for evaluating moral reasoning of large language models . it includes 3,000 tweets annotated with binary hate speech labels, moral categories and rationales . |
| Outcome: | The proposed dataset shows a misalignment between LLM outputs and human annotations in moral reasoning tasks. |
Exploring LLMs’ Ability to Spontaneously and Conditionally Modify Moral Expressions through Text Manipulation (2025.acl-long)
Copied to clipboard
| Challenge: | Existing studies on moral-related tasks based on large language models have not been conducted. |
| Approach: | They analyze behavior of Large Language Models (LLMs) among open and uncensored models and use human-annotated datasets to analyze moral-related data. |
| Outcome: | The results show that large language models can alter moral dimensions through text manipulation tasks and moral-related conditioning prompts. |
Moral Mimicry: Large Language Models Produce Moral Rationalizations Tailored to Political Identity (2023.acl-srw)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have demonstrated impressive capabilities in generating fluent text, as well as tendencies to reproduce undesirable social biases. |
| Approach: | They propose that LLMs reproduce moral biases associated with political groups in the United States, an instance of a broader capability termed moral mimicry. |
| Outcome: | The LLMs generated by the models reproduce moral biases associated with political groups in the United States, and this is an instance of a broader capability termed moral mimicry. |
Speaking Multiple Languages Affects the Moral Bias of Language Models (2023.findings-acl)
Copied to clipboard
Katharina Haemmerl, Bjoern Deiseroth, Patrick Schramowski, Jindřich Libovický, Constantin Rothkopf, Alexander Fraser, Kristian Kersting
| Challenge: | Pre-trained multilingual language models are often better on English than other languages . however, they are trained on varying amounts of data for each language . |
| Approach: | They apply the MORALDIRECTION framework to multilingual models and analyse their results . they find that PMLMs encode differing moral biases, but these do not correspond to cultural differences or commonalities in human opinions. |
| Outcome: | The proposed model captures moral norms from English and imposes them on other languages. |
Adaptable Moral Stances of Large Language Models on Sexist Content: Implications for Society and Gender Discourse (2024.emnlp-main)
Copied to clipboard
| Challenge: | Using large language models, large language model learning has become more integrated into our daily lives, making it increasingly important to ensure they reflect ethical and equitable values. |
| Approach: | They assess how LLMs can apply moral reasoning to both criticize and defend sexist language by evaluating their models and evaluating the moral foundations cited by them. |
| Outcome: | The models show they can provide comprehensible and contextually relevant text for understanding diverse views on how sexism is perceived. |
Do Morals Guide How LLMs Think? The Role of Ethical Perspectives in General Problem Solving (2026.acl-long)
Copied to clipboard
| Challenge: | Experimental results show that different moral perspectives lead to changes in the model’s decision-making during general reasoning, reflected in both responses and internal representations. |
| Approach: | They define distinct moral stages based on Kohlberg’s theory of moral development and design prompts to elicit model responses aligned with each condition. |
| Outcome: | The proposed model responses are validated using the Defining Issues Test, a human evaluation tool. |
Evaluating Moral Beliefs across LLMs through a Pluralistic Framework (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Proper moral beliefs are fundamental for language models, yet assessing these beliefs poses a significant challenge. |
| Approach: | They propose a framework to evaluate the moral beliefs of four large language models . they use a dataset containing 472 moral choice scenarios in Chinese . |
| Outcome: | The proposed framework evaluates the moral beliefs of four large language models. |
Moral Framing in Politics (MFiP): A new resource and models for moral framing (2025.emnlp-main)
Copied to clipboard
| Challenge: | Recent studies have focused on detecting moral values in political communication, trying to identify moral frames used by political actors or parties to convey their messages. |
| Approach: | They propose to code German parliamentary debates to identify moral framing and to detect subtle differences in politicians’ moral framming. |
| Outcome: | The proposed model distinguishes between different types of moral frames and includes narrative roles, together with the moral foundations for each frame. |
Comparing Moral Values in Western English-speaking societies and LLMs with Word Associations (2025.acl-long)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are trained on extensive corpora to learn linguistic patterns, contextual nuances, and implicit elements of human values. |
| Approach: | They propose to use word associations as low-level underlying representations to obtain a more robust picture of LLMs’ moral reasoning. |
| Outcome: | The proposed method reveals detailed but systematic differences between LLMs and human associations. |