Measuring Large Language Models’ Adversarial Behavior in Social Deduction Games (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing safety evaluations focus on refusal-based methods that test whether models avoid responding to inappropriate or violent requests, leaving open questions about how models behave in interactive social settings. |
| Approach: | They propose to use a meta-LLM to construct a closed behavioral taxonomy from a multi-agent simulation to examine adversarial behavior of large language models. |
| Outcome: | The proposed model-based model-driven model-model-based taxonomy shows that the model-led model-learning model exhibits distinct behavioral profiles and influences social stability and competitive success. |
Similar Papers
Bayesian Social Deduction with Graph-Informed Language Models (2026.acl-long)
Copied to clipboard
Shahab Rahimirad, Guven Gergerli, Lucia Romero, Angela Qian, Matthew Lyle Olson, Simon Stepputtis, Joseph Campbell
| Challenge: | Large language models (LLMs) have demonstrated remarkable general-purpose reasoning capabilities across a wide range of tasks. |
| Approach: | They propose a hybrid reasoning framework that externalizes belief inference to a structured probabilistic model while using an LLM for language understanding and interaction. |
| Outcome: | The proposed framework achieves competitive performance with larger models in Agent-Agent play and is the first language agent to defeat human players in a controlled study. |
An Empirical Study of Collective Behaviors and Social Dynamics in Large Language Model Agents (2026.eacl-long)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are increasingly mediating our social, cultural, and political interactions. |
| Approach: | They propose a method that reminds LLM agents to avoid harmful posting . they analyze 7M posts and interactions among 32K LLMs over a year . |
| Outcome: | The proposed method aims to find out whether LLMs influence toxic posting patterns and polarization in their community. |
A Group Fairness Lens for Large Language Models (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Existing methods focusing on a few groups lack a comprehensive categorical perspective to evaluate LLMs’ potential biases and unfairness. |
| Approach: | They propose to evaluate LLM biases from a group fairness lens using a hierarchical schema characterizing diverse social groups. |
| Outcome: | The proposed method mitigates biases in LLMs from a group fairness lens and encapsulates target-attribute combinations across multiple dimensions. |
Do LLM Agents Mirror Socio-Cognitive Effects in Power-Asymmetric Conversations? (2026.acl-long)
Copied to clipboard
| Challenge: | Power differences shape human communication through well-documented socio-cognitive effects . asymmetric relationships or power differentials give rise to well-known socio-computational effects - lianelli, 1976 . |
| Approach: | They simulate multi-turn, power-asymmetric dialogues with personas from diverse professions . they find that LLMs show key socio-cognitive effects of power, albeit with nuances and variability . |
| Outcome: | The results show that large language models exhibit socio-cognitive effects of power . the results are consistent with previous studies on LLMs . |
Exploring the Choice Behavior of Large Language Models (2025.findings-acl)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are increasingly being adopted across various domains where they help to make choices. |
| Approach: | They construct a virtual QA platform that includes three different experimental conditions, with four models from GPT and Llama series participating in repeated experiments. |
| Outcome: | The proposed model includes three experimental conditions and four models from GPT and Llama series. |
Evaluating Behavioral Alignment in Conflict Dialogue: A Multi-Dimensional Comparison of LLM Agents and Humans (2025.emnlp-main)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are increasingly used in socially complex, interaction-driven tasks, yet their ability to mirror human behavior in emotionally and strategically complex contexts remains underexplored. |
| Approach: | They examine alignment of personality-prompted Large Language Models in conflict dialogues that incorporate negotiation by simulating a five-factor personality profile. |
| Outcome: | The proposed model achieves the closest alignment with humans in linguistic style and emotional dynamics while Claude-3.7-Sonnet best reflects strategic behavior. |
Will LLMs Sink or Swim? Exploring Decision-Making Under Pressure (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Recent advances in Large Language Models (LLMs) have shown their ability to simulate human-like decision-making, yet the impact of psychological pressures on their decision- making processes remains underexplored. |
| Approach: | They used explicit and implicit pressure prompts to induce specific pressures and tested them on reasoning, psychometric, and game theory tasks. |
| Outcome: | The results show that pressures significantly affect LLMs’ decision-making, varying across tasks and models. |
Rethinking Pragmatics in Large Language Models: Towards Open-Ended Evaluation and Preference Tuning (2024.emnlp-main)
Copied to clipboard
| Challenge: | Existing methods to assess social-pragmatic inference in large language models are inadequacy, and preferential tuning is the best approach. |
| Approach: | They propose to use free-form models' responses as a measure to assess social-pragmatic reasoning and advocate for preference optimization over supervised finetuning (SFT). |
| Outcome: | The proposed model outperforms supervised finetuning (SFT) and offers a near-free launch in pragmatic abilities without compromising general capabilities. |
Systematic Biases in LLM Simulations of Debates (2024.emnlp-main)
Copied to clipboard
| Challenge: | Current research suggests that LLM-based agents become increasingly human-like in their performance, sparking interest in using these AI agents as substitutes for human participants in behavioral studies. |
| Approach: | They propose to use LLMs to simulate political debates on topics that are important aspects of people’s day-to-day lives and decision-making processes. |
| Outcome: | The proposed model can simulate political debates on topics that are important aspects of people’s day-to-day lives and decision-making processes. |
Is this the real life? Is this just fantasy? The Misleading Success of Simulating Social Interactions With LLMs (2024.emnlp-main)
Copied to clipboard
| Challenge: | Recent advances in large language models have enabled richer social simulations . however, the role of information asymmetry in these simulations has been overlooked . |
| Approach: | They develop an evaluation framework to simulate social interactions with LLMs in different settings. |
| Outcome: | The proposed framework performs better in unrealistic, omniscient simulation settings but struggles in those with information asymmetry. |