Challenge: a recent study has found that stories are central to how humans communicate moral values .
Approach: They compare human- and LLM-generated moral narratives based on images annotated by humans for moral content . authors propose a framework for evaluating moral storytelling in vision-language models .
Outcome: The proposed model compared human- and LLM-generated narratives on images . human stories reflect a balanced distribution of moral foundations and coherent narrative arcs, but LLMs emphasize Care foundation and lack emotional resolution.

Similar Papers

Story Morals: Surfacing value-driven narrative schemas using large language models (2024.emnlp-main)

Copied to clipboard

Challenge: Using large language models, we extract and validate story morals across a diverse set of narrative genres.
Approach: They propose a task of narrative schema labelling based on the concept of "story morals" they use large language models to extract and validate story morals across a diverse set of genres .
Outcome: The proposed method extracts and validates story morals across folktales, novels, movies and TV, personal stories from social media and the news using automated metrics and human assessments.
Comparing Moral Values in Western English-speaking societies and LLMs with Word Associations (2025.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) are trained on extensive corpora to learn linguistic patterns, contextual nuances, and implicit elements of human values.
Approach: They propose to use word associations as low-level underlying representations to obtain a more robust picture of LLMs’ moral reasoning.
Outcome: The proposed method reveals detailed but systematic differences between LLMs and human associations.
Exploring LLMs’ Ability to Spontaneously and Conditionally Modify Moral Expressions through Text Manipulation (2025.acl-long)

Copied to clipboard

Challenge: Existing studies on moral-related tasks based on large language models have not been conducted.
Approach: They analyze behavior of Large Language Models (LLMs) among open and uncensored models and use human-annotated datasets to analyze moral-related data.
Outcome: The results show that large language models can alter moral dimensions through text manipulation tasks and moral-related conditioning prompts.
Probing Narrative Morals: A New Character-Focused MFT Framework for Use with Large Language Models (2025.emnlp-main)

Copied to clipboard

Challenge: Existing methods to categorize moral foundations in storytelling are limited.
Approach: They propose a character-centric method to quantify moral foundations in storytelling using large language models and a novel Moral Foundations Character Action Questionnaire to validate their approach against human annotations.
Outcome: The proposed method validates against human annotations and then applies to 2,697 folktales from 55 countries.
Biased Tales: Cultural and Topic Bias in Generating Children’s Stories (2025.emnlp-main)

Copied to clipboard

Challenge: Personalized stories are often preferred because they reflect a child's interests, experiences, and developmental needs.
Approach: They analyze a dataset to examine how biases influence protagonists’ attributes and story elements in LLM-generated stories.
Outcome: The proposed dataset shows that gender stereotypes influence protagonist attributes and story elements in LLM-generated stories.
The Pluralistic Moral Gap: Understanding Moral Judgment and Value Differences between Humans and Large Language Models (2026.eacl-long)

Copied to clipboard

Challenge: Existing studies have shown that Large Language Models (LLMs) are not fully aligned with human moral judgments.
Approach: They propose a dataset of 1,618 real-world moral dilemmas paired with a distribution of human moral judgments consisting of a binary evaluation and a free-text rationale to examine how closely LLMs align with human moral judgements.
Outcome: The proposed model reproduces human judgments only under high consensus; alignment deteriorates sharply when human disagreement increases.
HISTOIRESMORALES: A French Dataset for Assessing Moral Alignment (2025.naacl-long)

Copied to clipboard

Challenge: HistoiresMorales is a dataset based on moralStories in French . it is based upon annotations of moral values within the dataset .
Approach: They propose a dataset in French that aims to align language models with moral values . they use annotations to ensure their alignment with French norms .
Outcome: The proposed dataset guarantees grammatical accuracy and adaptation to the French cultural context.
From Surveys to Narratives: Rethinking Cultural Value Adaptation in LLMs (2025.emnlp-main)

Copied to clipboard

Challenge: Adapting cultural values in Large Language Models presents significant challenges due to biases and data limitations.
Approach: They propose to augment World Values Survey (WVS) data with encyclopedic and scenario-based cultural narratives from Wikipedia and NormAd to address these limitations.
Outcome: The proposed approach enhances cultural distinctiveness and improves classification performance across cultures.
Moral Foundations of Large Language Models (2024.emnlp-main)

Copied to clipboard

Challenge: Moral foundations theory (MFT) is a psychological assessment tool that decomposes human moral reasoning into five factors, including care/harm, liberty/oppression, and sanctity/degradation.
Approach: They propose to use moral foundations theory to analyze whether popular LLMs have acquired a bias towards a particular set of moral values.
Outcome: The proposed model can be adversarially selected to exhibit a particular moral foundations and can affect downstream tasks.
Adaptable Moral Stances of Large Language Models on Sexist Content: Implications for Society and Gender Discourse (2024.emnlp-main)

Copied to clipboard

Challenge: Using large language models, large language model learning has become more integrated into our daily lives, making it increasingly important to ensure they reflect ethical and equitable values.
Approach: They assess how LLMs can apply moral reasoning to both criticize and defend sexist language by evaluating their models and evaluating the moral foundations cited by them.
Outcome: The models show they can provide comprehensible and contextually relevant text for understanding diverse views on how sexism is perceived.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations