MotiveBench: How Far Are We From Human-Like Motivational Reasoning in Large Language Models? (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing benchmarks for large language models lack information asymmetry with real-world situations. |
| Approach: | They propose a benchmark to evaluate the human-like motivational and behavioral reasoning ability of LLMs with detailed, realistic situations. |
| Outcome: | The proposed benchmark compared LLMs with real-world scenarios on seven model families and found that the most advanced models struggle with understanding "love & belonging" needs. |
Similar Papers
Towards large language model-based personal agents in the enterprise: Current trends and open problems (2023.findings-emnlp)
Copied to clipboard
Vinod Muthusamy, Yara Rizk, Kiran Kate, Praveen Venkateswaran, Vatche Isahagian, Ashu Gulati, Parijat Dube
| Challenge: | Existing large language models (LLMs) are brittle to input changes and can produce inconsistent results for the same inputs. |
| Approach: | They propose to use large language models to reason about complex goals and orchestrate a set of pluggable tools or APIs to accomplish a goal. |
| Outcome: | The proposed use cases have many open problems in an exciting area of NLP research, such as trust and explainability, consistency and reproducibility, and the need for new metrics and benchmarks. |
LogicBench: Towards Systematic Evaluation of Logical Reasoning Ability of Large Language Models (2024.acl-long)
Copied to clipboard
Mihir Parmar, Nisarg Patel, Neeraj Varshney, Mutsumi Nakamura, Man Luo, Santosh Mashetty, Arindam Mitra, Chitta Baral
| Challenge: | Existing work investigating the logical reasoning ability of large language models has focused only on a couple of inference rules of propositional and first-order logics. |
| Approach: | They propose to use a natural language question-answering dataset to evaluate the logical reasoning ability of large language models. |
| Outcome: | The proposed model performs poorly on a range of natural language questions using chain-of-thought prompting. |
Persona-Assigned Large Language Models Exhibit Human-Like Motivated Reasoning (2026.findings-acl)
Copied to clipboard
| Challenge: | Prior studies have reported that large language models (LLMs) are also susceptible to human-like cognitive biases, but the extent to which LLMs selectively reason toward identity-congruent conclusions remains unexplored. |
| Approach: | They investigate whether assigning 8 personas across 4 political and socio-demographic attributes induces motivated reasoning in LLMs. |
| Outcome: | The proposed model is assigned 8 personas across 4 political and socio-demographic attributes and shows that they have 9% reduced veracity discernment compared to models without persona. |
Large Human Language Models: A Need and the Challenges (2024.naacl-long)
Copied to clipboard
| Challenge: | a growing recognition of the importance of modeling human and social factors into human-centered NLP models . authors advocate for three positions toward creating large human language models based on psychological and behavioral sciences . |
| Approach: | et al. advocate for three positions toward creating large human language models . they argue that LM training should include the human context and recognize that people are more than their group . |
| Outcome: | a new study shows that learning language from linguistic signals alone is not adequate, according to a recent paper . authors advocate for three positions toward creating large human language models . a human-centered model should include the human context, and account for the dynamic nature of the human environment, they say . |
SocioBench: Modeling Human Behavior in Sociological Surveys with Large Language Models (2025.emnlp-main)
Copied to clipboard
| Challenge: | Large language models (LLMs) lack large-scale, systematically constructed benchmarks for evaluating their alignment with real-world social attitudes. |
| Approach: | They propose a benchmark to assess LLMs' alignment with real-world social attitudes . they find LLM models achieve only 30–40% accuracy when simulating individuals . |
| Outcome: | The proposed benchmark shows that LLMs achieve only 30% accuracy when simulating individuals in complex survey scenarios. |
Towards Reasoning in Large Language Models: A Survey (2023.findings-acl)
Copied to clipboard
| Challenge: | Reasoning is a fundamental aspect of human intelligence that plays a crucial role in many intellectual activities. |
| Approach: | They propose to improve LLMs' ability to elicit reasoning by providing exemplars or prompts to model reasoning. |
| Outcome: | This paper provides a comprehensive overview of the state of knowledge on reasoning in large language models. |
HER: Human-like Reasoning and Reinforcement Learning for LLM Role-playing (2026.findings-acl)
Copied to clipboard
Chengyu Du, Xintao Wang, Aili Chen, Weiyuan Li, Rui Xu, Junteng Liu, Zishan Huang, Rong Tian, Zijun Sun, Yuhao Li, Liheng Feng, Deming Ding, Pengyu Zhao, Yanghua Xiao
| Challenge: | Existing models for LLM role-playing lack high-quality datasets with explicit reasoning traces and reliable reward signals aligned with human preferences. |
| Approach: | They propose a unified framework for cognitive-level persona simulation that strictly distinguishes characters’ first-person thinking processes from LLMs’ third-person reasoning. |
| Outcome: | The proposed framework outperforms the Qwen3-32B baseline model and achieves a 30.26% and 14.97% performance on the minimax benchmarks. |
ValueBench: Towards Comprehensively Evaluating Value Orientations and Understanding of Large Language Models (2024.acl-long)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are transforming diverse fields and gaining increasing influence as human proxies. |
| Approach: | They propose a psychometric evaluation pipeline grounded in realistic human-AI interactions to probe value orientations and novel tasks for evaluating value understanding in an open-ended value space. |
| Outcome: | The proposed evaluation pipeline is grounded in realistic human-AI interactions and performs tasks that approximate expert conclusions in value-related extraction and generation tasks. |
Can Language Models Recognize Convincing Arguments? (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Existing studies have found that large language models can generate persuasive content without engaging in human experimentation. |
| Approach: | They extend a dataset with debates, votes, and user traits to measure LLMs' ability to distinguish between strong and weak arguments, predict stances based on beliefs and demographic characteristics, and determine appeal of argument to individual based upon their traits. |
| Outcome: | The proposed tasks outperform human predictions in detecting convincing arguments in debates, votes, and user traits. |
RiddleBench: A New Generative Reasoning Benchmark for LLMs (2026.findings-eacl)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) show remarkable capabilities, but complex reasoning skills require deeper investigation. |
| Approach: | They propose a benchmark of 1,737 puzzles to test reasoning beyond simple pattern matching. |
| Outcome: | The proposed model performs poorly when faced with reordered constraints or irrelevant information. |