Moral Foundations of Large Language Models (2024.emnlp-main)

Copied to clipboard

Challenge: Moral foundations theory (MFT) is a psychological assessment tool that decomposes human moral reasoning into five factors, including care/harm, liberty/oppression, and sanctity/degradation.
Approach: They propose to use moral foundations theory to analyze whether popular LLMs have acquired a bias towards a particular set of moral values.
Outcome: The proposed model can be adversarially selected to exhibit a particular moral foundations and can affect downstream tasks.

Similar Papers

Probing Narrative Morals: A New Character-Focused MFT Framework for Use with Large Language Models (2025.emnlp-main)

Copied to clipboard

Challenge: Existing methods to categorize moral foundations in storytelling are limited.
Approach: They propose a character-centric method to quantify moral foundations in storytelling using large language models and a novel Moral Foundations Character Action Questionnaire to validate their approach against human annotations.
Outcome: The proposed method validates against human annotations and then applies to 2,697 folktales from 55 countries.
MFTCXplain: A Multilingual Benchmark Dataset for Evaluating the Moral Reasoning of LLMs through Multi-hop Hate Speech Explanation (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing evaluation benchmarks for large language models lack annotations that justify moral classifications and focus on English constrain moral reasoning across diverse cultural settings.
Approach: They propose a multilingual benchmark dataset for evaluating moral reasoning of large language models . it includes 3,000 tweets annotated with binary hate speech labels, moral categories and rationales .
Outcome: The proposed dataset shows a misalignment between LLM outputs and human annotations in moral reasoning tasks.
Exploring LLMs’ Ability to Spontaneously and Conditionally Modify Moral Expressions through Text Manipulation (2025.acl-long)

Copied to clipboard

Challenge: Existing studies on moral-related tasks based on large language models have not been conducted.
Approach: They analyze behavior of Large Language Models (LLMs) among open and uncensored models and use human-annotated datasets to analyze moral-related data.
Outcome: The results show that large language models can alter moral dimensions through text manipulation tasks and moral-related conditioning prompts.
Moral Mimicry: Large Language Models Produce Moral Rationalizations Tailored to Political Identity (2023.acl-srw)

Copied to clipboard

Challenge: Large Language Models (LLMs) have demonstrated impressive capabilities in generating fluent text, as well as tendencies to reproduce undesirable social biases.
Approach: They propose that LLMs reproduce moral biases associated with political groups in the United States, an instance of a broader capability termed moral mimicry.
Outcome: The LLMs generated by the models reproduce moral biases associated with political groups in the United States, and this is an instance of a broader capability termed moral mimicry.
Speaking Multiple Languages Affects the Moral Bias of Language Models (2023.findings-acl)

Copied to clipboard

Challenge: Pre-trained multilingual language models are often better on English than other languages . however, they are trained on varying amounts of data for each language .
Approach: They apply the MORALDIRECTION framework to multilingual models and analyse their results . they find that PMLMs encode differing moral biases, but these do not correspond to cultural differences or commonalities in human opinions.
Outcome: The proposed model captures moral norms from English and imposes them on other languages.
Adaptable Moral Stances of Large Language Models on Sexist Content: Implications for Society and Gender Discourse (2024.emnlp-main)

Copied to clipboard

Challenge: Using large language models, large language model learning has become more integrated into our daily lives, making it increasingly important to ensure they reflect ethical and equitable values.
Approach: They assess how LLMs can apply moral reasoning to both criticize and defend sexist language by evaluating their models and evaluating the moral foundations cited by them.
Outcome: The models show they can provide comprehensible and contextually relevant text for understanding diverse views on how sexism is perceived.
Do Morals Guide How LLMs Think? The Role of Ethical Perspectives in General Problem Solving (2026.acl-long)

Copied to clipboard

Challenge: Experimental results show that different moral perspectives lead to changes in the model’s decision-making during general reasoning, reflected in both responses and internal representations.
Approach: They define distinct moral stages based on Kohlberg’s theory of moral development and design prompts to elicit model responses aligned with each condition.
Outcome: The proposed model responses are validated using the Defining Issues Test, a human evaluation tool.
Evaluating Moral Beliefs across LLMs through a Pluralistic Framework (2024.findings-emnlp)

Copied to clipboard

Challenge: Proper moral beliefs are fundamental for language models, yet assessing these beliefs poses a significant challenge.
Approach: They propose a framework to evaluate the moral beliefs of four large language models . they use a dataset containing 472 moral choice scenarios in Chinese .
Outcome: The proposed framework evaluates the moral beliefs of four large language models.
Moral Framing in Politics (MFiP): A new resource and models for moral framing (2025.emnlp-main)

Copied to clipboard

Challenge: Recent studies have focused on detecting moral values in political communication, trying to identify moral frames used by political actors or parties to convey their messages.
Approach: They propose to code German parliamentary debates to identify moral framing and to detect subtle differences in politicians’ moral framming.
Outcome: The proposed model distinguishes between different types of moral frames and includes narrative roles, together with the moral foundations for each frame.
Comparing Moral Values in Western English-speaking societies and LLMs with Word Associations (2025.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) are trained on extensive corpora to learn linguistic patterns, contextual nuances, and implicit elements of human values.
Approach: They propose to use word associations as low-level underlying representations to obtain a more robust picture of LLMs’ moral reasoning.
Outcome: The proposed method reveals detailed but systematic differences between LLMs and human associations.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations