Challenge: Existing methods to categorize moral foundations in storytelling are limited.
Approach: They propose a character-centric method to quantify moral foundations in storytelling using large language models and a novel Moral Foundations Character Action Questionnaire to validate their approach against human annotations.
Outcome: The proposed method validates against human annotations and then applies to 2,697 folktales from 55 countries.

Similar Papers

Moral Foundations of Large Language Models (2024.emnlp-main)

Copied to clipboard

Challenge: Moral foundations theory (MFT) is a psychological assessment tool that decomposes human moral reasoning into five factors, including care/harm, liberty/oppression, and sanctity/degradation.
Approach: They propose to use moral foundations theory to analyze whether popular LLMs have acquired a bias towards a particular set of moral values.
Outcome: The proposed model can be adversarially selected to exhibit a particular moral foundations and can affect downstream tasks.
Story Morals: Surfacing value-driven narrative schemas using large language models (2024.emnlp-main)

Copied to clipboard

Challenge: Using large language models, we extract and validate story morals across a diverse set of narrative genres.
Approach: They propose a task of narrative schema labelling based on the concept of "story morals" they use large language models to extract and validate story morals across a diverse set of genres .
Outcome: The proposed method extracts and validates story morals across folktales, novels, movies and TV, personal stories from social media and the news using automated metrics and human assessments.
Tales of Morality: Comparing Human- and LLM-Generated Moral Stories from Visual Cues (2025.findings-emnlp)

Copied to clipboard

Challenge: a recent study has found that stories are central to how humans communicate moral values .
Approach: They compare human- and LLM-generated moral narratives based on images annotated by humans for moral content . authors propose a framework for evaluating moral storytelling in vision-language models .
Outcome: The proposed model compared human- and LLM-generated narratives on images . human stories reflect a balanced distribution of moral foundations and coherent narrative arcs, but LLMs emphasize Care foundation and lack emotional resolution.
MFTCXplain: A Multilingual Benchmark Dataset for Evaluating the Moral Reasoning of LLMs through Multi-hop Hate Speech Explanation (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing evaluation benchmarks for large language models lack annotations that justify moral classifications and focus on English constrain moral reasoning across diverse cultural settings.
Approach: They propose a multilingual benchmark dataset for evaluating moral reasoning of large language models . it includes 3,000 tweets annotated with binary hate speech labels, moral categories and rationales .
Outcome: The proposed dataset shows a misalignment between LLM outputs and human annotations in moral reasoning tasks.
Moral Framing in Politics (MFiP): A new resource and models for moral framing (2025.emnlp-main)

Copied to clipboard

Challenge: Recent studies have focused on detecting moral values in political communication, trying to identify moral frames used by political actors or parties to convey their messages.
Approach: They propose to code German parliamentary debates to identify moral framing and to detect subtle differences in politicians’ moral framming.
Outcome: The proposed model distinguishes between different types of moral frames and includes narrative roles, together with the moral foundations for each frame.
Exploring LLMs’ Ability to Spontaneously and Conditionally Modify Moral Expressions through Text Manipulation (2025.acl-long)

Copied to clipboard

Challenge: Existing studies on moral-related tasks based on large language models have not been conducted.
Approach: They analyze behavior of Large Language Models (LLMs) among open and uncensored models and use human-annotated datasets to analyze moral-related data.
Outcome: The results show that large language models can alter moral dimensions through text manipulation tasks and moral-related conditioning prompts.
Are Rules Meant to be Broken? Understanding Multilingual Moral Reasoning as a Computational Pipeline with UniMoral (2025.acl-long)

Copied to clipboard

Challenge: Existing approaches to analyze moral reasoning are discordant and lack cohesion, focusing on isolated aspects of the process.
Approach: They propose a unified dataset that integrates moral dilemmas annotated with labels for action choices, ethical principles, contributing factors, and consequences, and captures diverse socio-cultural contexts.
Outcome: The proposed dataset integrates moral dilemmas annotated with labels for action choices, ethical principles, contributing factors, and consequences, along with annotators’ moral and cultural profiles.
Adaptable Moral Stances of Large Language Models on Sexist Content: Implications for Society and Gender Discourse (2024.emnlp-main)

Copied to clipboard

Challenge: Using large language models, large language model learning has become more integrated into our daily lives, making it increasingly important to ensure they reflect ethical and equitable values.
Approach: They assess how LLMs can apply moral reasoning to both criticize and defend sexist language by evaluating their models and evaluating the moral foundations cited by them.
Outcome: The models show they can provide comprehensible and contextually relevant text for understanding diverse views on how sexism is perceived.
Evaluating Moral Beliefs across LLMs through a Pluralistic Framework (2024.findings-emnlp)

Copied to clipboard

Challenge: Proper moral beliefs are fundamental for language models, yet assessing these beliefs poses a significant challenge.
Approach: They propose a framework to evaluate the moral beliefs of four large language models . they use a dataset containing 472 moral choice scenarios in Chinese .
Outcome: The proposed framework evaluates the moral beliefs of four large language models.
CharMoral: A Character Morality Dataset for Morally Dynamic Character Analysis in Long-Form Narratives (2025.coling-main)

Copied to clipboard

Challenge: Existing studies on character analysis focus on character identification, social network analysis, and the exploration of characters' personas or personalities.
Approach: They propose a four-stage framework to automatically classify actions as moral or immoral based on context.
Outcome: The proposed framework is effective in moral reasoning tasks in multiple genres.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations