Challenge: Large language models (LLMs) are increasingly used to model and augment collective decision-making.
Approach: They propose a framework for assessing collective alignment using the Lost at Sea social psychology task.
Outcome: The proposed framework compares LLMs with human-AI alignment on the Lost at Sea social psychology task.

Similar Papers

Do LLMs Align Human Values Regarding Social Biases? Judging and Explaining Social Biases with LLMs (2025.findings-emnlp)

Copied to clipboard

Challenge: Large language models can lead to undesired consequences when misaligned with human values . previous studies have shown misalignment of LLMs with human value using expert-designed or agent-based emulated bias scenarios .
Approach: They investigate whether large language models (LLMs) are misaligned with human values . they find no significant differences in understanding of HVSB between LLMs .
Outcome: The results show that large language models do not have lower misalignment rates and attack success rates . the study also shows that smaller language models have the ability to explain HVSB .
Aligning Black-box Language Models with Human Judgments (2025.findings-naacl)

Copied to clipboard

Challenge: Large language models (LLMs) are increasingly used as automated judges to evaluate recommendation systems, search engines, and other subjective tasks.
Approach: They propose a framework to align LLM judgments with individual human evaluators or their aggregated judgments without retraining or fine-tuning the LLM.
Outcome: The proposed framework achieves 142% improvement in agreement across 29 tasks and exceeds inter-human agreement on four out of six tasks.
Human Alignment: How Much Do We Adapt to LLMs? (2025.acl-short)

Copied to clipboard

Challenge: Large Language Models (LLMs) are becoming a common part of our lives, yet few studies have examined how they influence our behavior.
Approach: They propose a cooperative language game in which players aim to converge on a word and play a game in a group.
Outcome: The proposed game shows that humans notice and adapt to differences regardless of whether they are aware they are interacting with an LLM.
Aligning Language Models to User Opinions (2023.findings-emnlp)

Copied to clipboard

Challenge: Personality is a defining feature of human beings, shaped by a complex interplay of demographic characteristics, moral principles, and social experiences.
Approach: They use public opinion surveys to model past user opinions in addition to user demographics and ideology to achieve up to 7 points accuracy gains in predicting public opinions from survey questions.
Outcome: The proposed model achieves 7 points accuracy gains in predicting public opinions from public opinion surveys across a broad set of topics.
An Empirical Study of Group Conformity in Multi-Agent Systems (2025.findings-acl)

Copied to clipboard

Challenge: Recent advances in Large Language Models (LLMs) have enabled multi-agent systems that simulate real-world interactions with near-human reasoning.
Approach: They analyze how LLM agents shape public opinion through debates on five contentious topics by simulating over 2,500 debates.
Outcome: The proposed models show that LLM agents adopt specific stances over time and align with numerically dominant groups or more intelligent agents, exerting a greater influence.
Reasons to Reject? Aligning Language Models with Judgments (2024.findings-acl)

Copied to clipboard

Challenge: a new framework for aligning large language models with judgments is proposed to help with alignment . a framework that allows for fine-grained inappropriate content detection and correction based on judgments . large language model alignment is critical for making artificial intelligence a reliable ally for humanity .
Approach: They propose a framework that allows for fine-grained inappropriate content detection and correction based on judgments.
Outcome: The proposed framework beats the 175B DaVinci003 and improves on AlpacaEval using judgments.
Persona-Assigned Large Language Models Exhibit Human-Like Motivated Reasoning (2026.findings-acl)

Copied to clipboard

Challenge: Prior studies have reported that large language models (LLMs) are also susceptible to human-like cognitive biases, but the extent to which LLMs selectively reason toward identity-congruent conclusions remains unexplored.
Approach: They investigate whether assigning 8 personas across 4 political and socio-demographic attributes induces motivated reasoning in LLMs.
Outcome: The proposed model is assigned 8 personas across 4 political and socio-demographic attributes and shows that they have 9% reduced veracity discernment compared to models without persona.
SocialGaze: Improving the Integration of Human Social Norms in Large Language Models (2024.findings-emnlp)

Copied to clipboard

Challenge: Increasingly, large language models (LLMs) are able to understand and rationalize socially acceptable behaviors, but they are often misaligned with human consensus.
Approach: They propose a multi-step prompting framework that verbalizes a social situation from multiple perspectives before forming a judgment.
Outcome: The proposed framework improves the alignment with human judgments by up to 11 F1 points with the GPT-3.5 model.
Can Large Language Models Capture Dissenting Human Voices? (2023.emnlp-main)

Copied to clipboard

Challenge: Large language models (LLMs) have shown impressive achievements in solving a broad range of tasks.
Approach: They evaluate the performance and alignment of large language models with humans using Monte Carlo Estimation and Log Probability Estimationic methods to estimate the multinomial distribution.
Outcome: The proposed models fail to capture human disagreement distribution and inference and human alignment performance plunge even further on data samples with high disagreement levels raising concerns about their natural language understanding ability and representativeness to a larger human population.
Bias in the Mirror : Are LLMs opinions robust to their own adversarial attacks (2025.acl-long)

Copied to clipboard

Challenge: Existing work on large language models lacks robustness, highlighting the limitations of such models.
Approach: They propose a novel approach where two LLMs engage in self-debate to persuade a neutral version of the model.
Outcome: The proposed approach examines whether large language models are robust during interactions and whether they are susceptible to reinforcing misinformation or shifting to harmful viewpoints.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations