Challenge: Existing safety evaluations focus on refusal-based methods that test whether models avoid responding to inappropriate or violent requests, leaving open questions about how models behave in interactive social settings.
Approach: They propose to use a meta-LLM to construct a closed behavioral taxonomy from a multi-agent simulation to examine adversarial behavior of large language models.
Outcome: The proposed model-based model-driven model-model-based taxonomy shows that the model-led model-learning model exhibits distinct behavioral profiles and influences social stability and competitive success.

Similar Papers

Bayesian Social Deduction with Graph-Informed Language Models (2026.acl-long)

Copied to clipboard

Challenge: Large language models (LLMs) have demonstrated remarkable general-purpose reasoning capabilities across a wide range of tasks.
Approach: They propose a hybrid reasoning framework that externalizes belief inference to a structured probabilistic model while using an LLM for language understanding and interaction.
Outcome: The proposed framework achieves competitive performance with larger models in Agent-Agent play and is the first language agent to defeat human players in a controlled study.
An Empirical Study of Collective Behaviors and Social Dynamics in Large Language Model Agents (2026.eacl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) are increasingly mediating our social, cultural, and political interactions.
Approach: They propose a method that reminds LLM agents to avoid harmful posting . they analyze 7M posts and interactions among 32K LLMs over a year .
Outcome: The proposed method aims to find out whether LLMs influence toxic posting patterns and polarization in their community.
A Group Fairness Lens for Large Language Models (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing methods focusing on a few groups lack a comprehensive categorical perspective to evaluate LLMs’ potential biases and unfairness.
Approach: They propose to evaluate LLM biases from a group fairness lens using a hierarchical schema characterizing diverse social groups.
Outcome: The proposed method mitigates biases in LLMs from a group fairness lens and encapsulates target-attribute combinations across multiple dimensions.
Do LLM Agents Mirror Socio-Cognitive Effects in Power-Asymmetric Conversations? (2026.acl-long)

Copied to clipboard

Challenge: Power differences shape human communication through well-documented socio-cognitive effects . asymmetric relationships or power differentials give rise to well-known socio-computational effects - lianelli, 1976 .
Approach: They simulate multi-turn, power-asymmetric dialogues with personas from diverse professions . they find that LLMs show key socio-cognitive effects of power, albeit with nuances and variability .
Outcome: The results show that large language models exhibit socio-cognitive effects of power . the results are consistent with previous studies on LLMs .
Exploring the Choice Behavior of Large Language Models (2025.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) are increasingly being adopted across various domains where they help to make choices.
Approach: They construct a virtual QA platform that includes three different experimental conditions, with four models from GPT and Llama series participating in repeated experiments.
Outcome: The proposed model includes three experimental conditions and four models from GPT and Llama series.
Evaluating Behavioral Alignment in Conflict Dialogue: A Multi-Dimensional Comparison of LLM Agents and Humans (2025.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) are increasingly used in socially complex, interaction-driven tasks, yet their ability to mirror human behavior in emotionally and strategically complex contexts remains underexplored.
Approach: They examine alignment of personality-prompted Large Language Models in conflict dialogues that incorporate negotiation by simulating a five-factor personality profile.
Outcome: The proposed model achieves the closest alignment with humans in linguistic style and emotional dynamics while Claude-3.7-Sonnet best reflects strategic behavior.
Will LLMs Sink or Swim? Exploring Decision-Making Under Pressure (2024.findings-emnlp)

Copied to clipboard

Challenge: Recent advances in Large Language Models (LLMs) have shown their ability to simulate human-like decision-making, yet the impact of psychological pressures on their decision- making processes remains underexplored.
Approach: They used explicit and implicit pressure prompts to induce specific pressures and tested them on reasoning, psychometric, and game theory tasks.
Outcome: The results show that pressures significantly affect LLMs’ decision-making, varying across tasks and models.
Rethinking Pragmatics in Large Language Models: Towards Open-Ended Evaluation and Preference Tuning (2024.emnlp-main)

Copied to clipboard

Challenge: Existing methods to assess social-pragmatic inference in large language models are inadequacy, and preferential tuning is the best approach.
Approach: They propose to use free-form models' responses as a measure to assess social-pragmatic reasoning and advocate for preference optimization over supervised finetuning (SFT).
Outcome: The proposed model outperforms supervised finetuning (SFT) and offers a near-free launch in pragmatic abilities without compromising general capabilities.
Systematic Biases in LLM Simulations of Debates (2024.emnlp-main)

Copied to clipboard

Challenge: Current research suggests that LLM-based agents become increasingly human-like in their performance, sparking interest in using these AI agents as substitutes for human participants in behavioral studies.
Approach: They propose to use LLMs to simulate political debates on topics that are important aspects of people’s day-to-day lives and decision-making processes.
Outcome: The proposed model can simulate political debates on topics that are important aspects of people’s day-to-day lives and decision-making processes.
Is this the real life? Is this just fantasy? The Misleading Success of Simulating Social Interactions With LLMs (2024.emnlp-main)

Copied to clipboard

Challenge: Recent advances in large language models have enabled richer social simulations . however, the role of information asymmetry in these simulations has been overlooked .
Approach: They develop an evaluation framework to simulate social interactions with LLMs in different settings.
Outcome: The proposed framework performs better in unrealistic, omniscient simulation settings but struggles in those with information asymmetry.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations