Challenge: Game theory provides a framework for studying human behaviors through incentivized games that simulate social situations.
Approach: They used two validated games from the cognitive science literature to study how well several recent open- and closed-source LLMs predict player actions with underlying human motives.
Outcome: The results show that state-of-the-art LLMs can achieve accuracy close to human levels in predicting players’ actions with underlying human motives in SPGs, but failed to recognize statistical patterns in players’ action.

Similar Papers

MotiveBench: How Far Are We From Human-Like Motivational Reasoning in Large Language Models? (2025.findings-acl)

Copied to clipboard

Challenge: Existing benchmarks for large language models lack information asymmetry with real-world situations.
Approach: They propose a benchmark to evaluate the human-like motivational and behavioral reasoning ability of LLMs with detailed, realistic situations.
Outcome: The proposed benchmark compared LLMs with real-world scenarios on seven model families and found that the most advanced models struggle with understanding "love & belonging" needs.
InterIntent: Investigating Social Intelligence of LLMs via Intention Understanding in an Interactive Game Context (2024.emnlp-main)

Copied to clipboard

Challenge: Large language models (LLMs) have demonstrated the potential to mimic human social intelligence, but most studies focus on static self-report or performance-based tests.
Approach: They propose a framework to assess LLMs' ability to understand and manage intentions by mapping their ability to infer the intentions of others in a game setting.
Outcome: The proposed framework assesses LLMs' ability to understand and manage intentions in a game setting.
Investigating Human and LLMs’ Decisions in Unverifiable Environments: A Case Study with GitHub Activity Overview (2026.findings-acl)

Copied to clipboard

Challenge: examining the behaviors of Large Language Models as artificial social actors is underexplored, especially in unverifiable scenarios where conventional benchmarking has little to help improve their abilities.
Approach: They propose a method to collect, compare, and reason about human and LLMs' decisions in an unverifiable scenario and use it to examine their behaviors.
Outcome: The proposed method compared human and LLM decisions in an unverifiable scenario on GitHub and found that proprietary LLMs behave more like humans than open-source LLM systems.
SocialGaze: Improving the Integration of Human Social Norms in Large Language Models (2024.findings-emnlp)

Copied to clipboard

Challenge: Increasingly, large language models (LLMs) are able to understand and rationalize socially acceptable behaviors, but they are often misaligned with human consensus.
Approach: They propose a multi-step prompting framework that verbalizes a social situation from multiple perspectives before forming a judgment.
Outcome: The proposed framework improves the alignment with human judgments by up to 11 F1 points with the GPT-3.5 model.
Noise, Adaptation, and Strategy: Assessing LLM Fidelity in Decision-Making (2025.emnlp-main)

Copied to clipboard

Challenge: Large language models (LLMs) are increasingly used for social science simulations . however, most evaluations focus on task optimality rather than variability and adaptation characteristic of human decision-making.
Approach: They propose a process-oriented evaluation framework with progressive interventions to evaluate two economics tasks using large language models.
Outcome: The proposed evaluation framework targets two economic tasks with progressive interventions.
Rethinking Pragmatics in Large Language Models: Towards Open-Ended Evaluation and Preference Tuning (2024.emnlp-main)

Copied to clipboard

Challenge: Existing methods to assess social-pragmatic inference in large language models are inadequacy, and preferential tuning is the best approach.
Approach: They propose to use free-form models' responses as a measure to assess social-pragmatic reasoning and advocate for preference optimization over supervised finetuning (SFT).
Outcome: The proposed model outperforms supervised finetuning (SFT) and offers a near-free launch in pragmatic abilities without compromising general capabilities.
Persona-Assigned Large Language Models Exhibit Human-Like Motivated Reasoning (2026.findings-acl)

Copied to clipboard

Challenge: Prior studies have reported that large language models (LLMs) are also susceptible to human-like cognitive biases, but the extent to which LLMs selectively reason toward identity-congruent conclusions remains unexplored.
Approach: They investigate whether assigning 8 personas across 4 political and socio-demographic attributes induces motivated reasoning in LLMs.
Outcome: The proposed model is assigned 8 personas across 4 political and socio-demographic attributes and shows that they have 9% reduced veracity discernment compared to models without persona.
Comparing Inferential Strategies of Humans and Large Language Models in Deductive Reasoning (2024.acl-long)

Copied to clipboard

Challenge: Recent advances in the domain of large language models (LLMs) have showcased their capability in executing deductive reasoning tasks.
Approach: They examine inferential strategies employed by large language models through a detailed evaluation of their responses to propositional logic problems.
Outcome: The proposed model shows that it displays reasoning patterns similar to humans, including strategies like supposition following or chain construction.
InMind: Evaluating LLMs in Capturing and Applying Individual Human Reasoning Styles (2025.emnlp-main)

Copied to clipboard

Challenge: Recent large language models (LLMs) have demonstrated strong reasoning abilities across complex mathematical and scientific domains.
Approach: They propose a framework to assess whether LLMs can capture and apply personalized reasoning styles in social deduction games.
Outcome: The proposed framework evaluates LLMs on the game Avalon and shows that they can capture and apply individualized reasoning styles.
Decision Biases and Intent-Irony Decoupling in Large Language Models (2026.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) exhibit impressive linguistic fluency, but it remains unclear whether they possess human-like Theory of Mind (ToM) or rely on statistical heuristics . a recent study examined the performance of LLMs against 300 human participants .
Approach: a study establishes a framework for large language models that modulates contextual contrast, linguistic cues, and cognitive mechanisms.
Outcome: a new evaluation framework compares ten state-of-the-art LLMs against 300 human participants . the framework systematically modulates contextual contrast, linguistic cues, and cognitive mechanisms .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations