Challenge: Large Language Models (LLMs) are trained on extensive corpora to learn linguistic patterns, contextual nuances, and implicit elements of human values.
Approach: They propose to use word associations as low-level underlying representations to obtain a more robust picture of LLMs’ moral reasoning.
Outcome: The proposed method reveals detailed but systematic differences between LLMs and human associations.

Similar Papers

Tales of Morality: Comparing Human- and LLM-Generated Moral Stories from Visual Cues (2025.findings-emnlp)

Copied to clipboard

Challenge: a recent study has found that stories are central to how humans communicate moral values .
Approach: They compare human- and LLM-generated moral narratives based on images annotated by humans for moral content . authors propose a framework for evaluating moral storytelling in vision-language models .
Outcome: The proposed model compared human- and LLM-generated narratives on images . human stories reflect a balanced distribution of moral foundations and coherent narrative arcs, but LLMs emphasize Care foundation and lack emotional resolution.
Knowledge of cultural moral norms in large language models (2023.acl-long)

Copied to clipboard

Challenge: Existing studies do not examine moral variation in a diverse cultural setting.
Approach: They investigate whether monolingual English language models capture moral variation across cultures . they use data from the World Values Survey and PEW global surveys .
Outcome: The proposed models predict moral norms worse than the English models reported previously . the models improve inference across countries at the expense of an accurate estimate .
Structured Moral Reasoning in Language Models: A Value-Grounded Evaluation Framework (2025.emnlp-main)

Copied to clipboard

Challenge: Large language models (LLMs) are increasingly deployed in domains requiring moral understanding, yet their reasoning often remains shallow and misaligned with human reasoning.
Approach: They propose a value-grounded framework for evaluating and distilling structured moral reasoning in large language models.
Outcome: The proposed framework evaluates 12 open-source models across four moral datasets.
Moral Foundations of Large Language Models (2024.emnlp-main)

Copied to clipboard

Challenge: Moral foundations theory (MFT) is a psychological assessment tool that decomposes human moral reasoning into five factors, including care/harm, liberty/oppression, and sanctity/degradation.
Approach: They propose to use moral foundations theory to analyze whether popular LLMs have acquired a bias towards a particular set of moral values.
Outcome: The proposed model can be adversarially selected to exhibit a particular moral foundations and can affect downstream tasks.
From Word to World: Evaluate and Mitigate Culture Bias in LLMs via Word Association Test (2025.emnlp-main)

Copied to clipboard

Challenge: Multilingual and cross-cultural WAT reveal how culture modulates perceptual and interactive patterns.
Approach: They propose to embed cultural-specific semantic associations directly within large language models (LLMs) to address cultural preference.
Outcome: The proposed model significantly improves cross-cultural alignment, capturing diverse semantic associations.
Exploring Multilingual Concepts of Human Values in Large Language Models: Is Value Alignment Consistent, Transferable and Controllable across Languages? (2024.findings-emnlp)

Copied to clipboard

Challenge: Prior research has revealed that certain abstract concepts are linearly represented as directions in the representation space of LLMs, predominantly centered around English.
Approach: They extend previous research that shows certain abstract concepts are linearly represented as directions in LLMs, predominantly centered around English.
Outcome: The proposed model can be used to align LLMs with human values, and it can generate toxic, untruthful, biased, and even illegal content.
Exploring LLMs’ Ability to Spontaneously and Conditionally Modify Moral Expressions through Text Manipulation (2025.acl-long)

Copied to clipboard

Challenge: Existing studies on moral-related tasks based on large language models have not been conducted.
Approach: They analyze behavior of Large Language Models (LLMs) among open and uncensored models and use human-annotated datasets to analyze moral-related data.
Outcome: The results show that large language models can alter moral dimensions through text manipulation tasks and moral-related conditioning prompts.
HISTOIRESMORALES: A French Dataset for Assessing Moral Alignment (2025.naacl-long)

Copied to clipboard

Challenge: HistoiresMorales is a dataset based on moralStories in French . it is based upon annotations of moral values within the dataset .
Approach: They propose a dataset in French that aims to align language models with moral values . they use annotations to ensure their alignment with French norms .
Outcome: The proposed dataset guarantees grammatical accuracy and adaptation to the French cultural context.
The Pluralistic Moral Gap: Understanding Moral Judgment and Value Differences between Humans and Large Language Models (2026.eacl-long)

Copied to clipboard

Challenge: Existing studies have shown that Large Language Models (LLMs) are not fully aligned with human moral judgments.
Approach: They propose a dataset of 1,618 real-world moral dilemmas paired with a distribution of human moral judgments consisting of a binary evaluation and a free-text rationale to examine how closely LLMs align with human moral judgements.
Outcome: The proposed model reproduces human judgments only under high consensus; alignment deteriorates sharply when human disagreement increases.
Moral Mimicry: Large Language Models Produce Moral Rationalizations Tailored to Political Identity (2023.acl-srw)

Copied to clipboard

Challenge: Large Language Models (LLMs) have demonstrated impressive capabilities in generating fluent text, as well as tendencies to reproduce undesirable social biases.
Approach: They propose that LLMs reproduce moral biases associated with political groups in the United States, an instance of a broader capability termed moral mimicry.
Outcome: The LLMs generated by the models reproduce moral biases associated with political groups in the United States, and this is an instance of a broader capability termed moral mimicry.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations