Challenge: Recent studies on transformer-based language models have shown that there seems to be a 'moral dimension' to LMs, as they show high accuracy in related downstream tasks such as moral reasoning and action classification.
Approach: They propose a mechanism based on deontic logic to allow for a flexible adaptation of individual norms by de-biasing training data sets and a task-reduction to textual entailment.
Outcome: The proposed mechanism de-biases training data sets and reduces tasks to textual entailment.

Similar Papers

Moral Stories: Situated Reasoning about Norms, Intents, Actions, and their Consequences (2021.emnlp-main)

Copied to clipboard

Challenge: aaron carroll: in social settings, human behavior is governed by unspoken rules of conduct rooted in societal norms . carroll and colleagues examine whether language generation models can serve as behavioral priors if they are not . they say we examine whether they can generate descriptions of actions that accomplish predefined goals .
Approach: They propose to combine multiple expert models to improve quality of generated actions, consequences, and norms.
Outcome: The proposed models significantly improve the quality of generated actions, consequences, and norms compared to baselines.
Measuring Social Norms of Large Language Models (2024.findings-naacl)

Copied to clipboard

Challenge: Existing datasets that evaluate a general understanding of social science are inadequate to understand social norms.
Approach: They propose a multi-agent framework to improve large language models’ ability to understand social norms by comparing them to elementary students.
Outcome: The proposed framework improves large language models to be on par with humans.
Toward Inclusive Language Models: Sparsity-Driven Calibration for Systematic and Interpretable Mitigation of Social Biases in LLMs (2025.findings-emnlp)

Copied to clipboard

Challenge: a new method to mitigate stereotypical bias in large language models is needed . inherent biases from training on vast Internet datasets can amplify harmful stereotypes .
Approach: They propose a method to identify stereotypical bias in decoder-only transformer models . they apply a localization mechanism that correlates internal activations with a new Context Influence score .
Outcome: The proposed method reduces stereotypical biases on BBQ, StereoSet, and CrowS-Pairs while improving reasoning performance on MMLU by 10%.
How Inclusively do LMs Perceive Social and Moral Norms? (2025.findings-naacl)

Copied to clipboard

Challenge: Language models (LMs) are used in decision-making systems and as interactive assistants.
Approach: They propose to prompt 11 LMs on rules-of-thumb and compare their outputs with 100 human annotators.
Outcome: The proposed model is compared with 100 human annotators to find out if they are inclusive of diverse human values.
NormAd: A Framework for Measuring the Cultural Adaptability of Large Language Models (2025.naacl-long)

Copied to clipboard

Challenge: Large language models (LLMs) are widely used and engage millions of users from diverse contexts and cultures.
Approach: They propose an evaluation framework to assess LLMs’ cultural adaptability by measuring their ability to judge social acceptability across varying levels of cultural norm specificity.
Outcome: The proposed model shows stronger adaptability to English-centric cultures over those from the Global South.
NormDial: A Comparable Bilingual Synthetic Dialog Dataset for Modeling Social Norm Adherence and Violation (2023.emnlp-main)

Copied to clipboard

Challenge: Social norms fundamentally shape interpersonal communication.
Approach: They propose a human-in-the-loop pipeline to synthesize a bilingual dyadic dialogue dataset with turn-by-turn annotations of social norms for Chinese and American cultures.
Outcome: The proposed dataset is high-quality through human evaluation and compares with existing models.
It’s Morphin’ Time! Combating Linguistic Discrimination with Inflectional Perturbations (2020.acl-main)

Copied to clipboard

Challenge: Existing work on societal bias in NLP focuses on race and gender . linguistic background is a unique attribute that has been largely ignored in the field .
Approach: They examine linguistic background to craft plausible adversarial examples that expose biases in popular NLP models.
Outcome: The proposed model improves robustness without sacrificing performance on clean data.
Aligning to Social Norms and Values in Interactive Narratives (2022.naacl-main)

Copied to clipboard

Challenge: Social value alignment is the ability to create agents that act in alignment with socially beneficial norms and values in interactive narratives or text-based games.
Approach: They introduce a game-value ALignment agent that uses social commonsense to restrict its action space to actions that are aligned with socially beneficial values.
Outcome: The proposed agent improves state-of-the-art task performance by 4% while reducing the frequency of socially harmful behaviors by 25% compared to strong contemporary value alignment approaches.
Fairness Evaluation and Inference Level Mitigation in LLMs (2026.findings-acl)

Copied to clipboard

Challenge: Large language models display undesirable behaviors embedded in their internal representations, undermining fairness, inconsistency drift, and the propagation of unwanted patterns during extended dialogues.
Approach: They propose a pruning-based framework that detects context-aware neuron activations and applies adaptive masking to modulate their influence during generation.
Outcome: The proposed framework detects context-aware neuron activations and applies adaptive masking to modulate their influence during generation.
He is very intelligent, she is very beautiful? On Mitigating Social Biases in Language Modelling and Generation (2021.findings-acl)

Copied to clipboard

Challenge: Existing studies have focused on mitigating social biases in context-free representations, with recent shift to contextual ones.
Approach: They propose an approach to mitigate social biases in a large pre-trained contextual language model . they propose lexical co-occurrence-based bias penalization in the decoder units .
Outcome: The proposed approach reduces biases in fill-in-the-blank sentences and summarizes . it also reduces the biased representations in the frameworks, the authors show .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations