Challenge: Existing methods for evaluating curiosity-like behaviors in large language models lack curiosity-inspired features.
Approach: They propose a psychology-inspired framework to evaluate curiosity in large language models . they adapt the Five-Dimensional Curiosity scale Revised (5DCR) to LLMs .
Outcome: The proposed framework evaluates curiosity in large language models using questionnaires and behavioral studies.

Similar Papers

Modeling, Evaluating, and Embodying Personality in LLMs: A Survey (2025.findings-emnlp)

Copied to clipboard

Challenge: This survey provides a comprehensive overview of the LLM-driven personality scenario.
Approach: This survey provides a comprehensive overview of the LLM-driven personality scenario.
Outcome: The proposed taxonomy analyzes the limitations of existing methods and identifies key research gaps.
MotiveBench: How Far Are We From Human-Like Motivational Reasoning in Large Language Models? (2025.findings-acl)

Copied to clipboard

Challenge: Existing benchmarks for large language models lack information asymmetry with real-world situations.
Approach: They propose a benchmark to evaluate the human-like motivational and behavioral reasoning ability of LLMs with detailed, realistic situations.
Outcome: The proposed benchmark compared LLMs with real-world scenarios on seven model families and found that the most advanced models struggle with understanding "love & belonging" needs.
Exploring the Choice Behavior of Large Language Models (2025.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) are increasingly being adopted across various domains where they help to make choices.
Approach: They construct a virtual QA platform that includes three different experimental conditions, with four models from GPT and Llama series participating in repeated experiments.
Outcome: The proposed model includes three experimental conditions and four models from GPT and Llama series.
A Comprehensive Survey on Learning from Rewards for Large Language Models: Reward Models and Learning Strategies (2025.findings-emnlp)

Copied to clipboard

Challenge: Recent developments in Large Language Models have shifted from pre-training to post-training and test-time scaling.
Approach: They present a comprehensive overview of learning from rewards from the perspective of reward models and learning strategies across training, inference, and post-inference stages.
Outcome: The proposed paradigm enables the transition from passive learning from static data to active learning from dynamic feedback.
A Comprehensive Survey of Process Reward Models: Data Generation, Model Construction, and Usage (2026.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) have advanced reasoning ability, yet conventional alignment remains dominated by outcome reward models that judge only final answers.
Approach: They summarize applications across math, code, text, multimodal reasoning, robotics, and agents . goal is to clarify design spaces, reveal open challenges, and guide future research toward fine-grained, robust reasoning alignment.
Outcome: The proposed model enables finer credit assignment, richer diagnostics, and improved robustness.
Curiosity-Driven Reinforcement Learning from Human Feedback (2025.acl-long)

Copied to clipboard

Challenge: Reinforcement learning from human feedback (RLHF) has proven effective in aligning large language models with human preferences, but often at the cost of reduced output diversity.
Approach: They propose a framework that incorporates intrinsic rewards for novel states alongside traditional sparse extrinsic rewards to optimize both output diversity and alignment quality.
Outcome: The proposed framework achieves significant gains in diversity on multiple diversity-oriented metrics while maintaining alignment with human preferences comparable to standard RLHF.
ReEfBench: Quantifying the Reasoning Efficiency of LLMs (2026.acl-long)

Copied to clipboard

Challenge: Existing methods for Chain-of-Thought evaluations do not distinguish between genuine reasoning and mere verbosity.
Approach: They propose a framework for the non-intrusive, comprehensive process-centric evaluation of reasoning grounded in First-Order Logic.
Outcome: The proposed framework identifies four distinct behavioral prototypes and diagnoses the failure modes.
When the LM misunderstood the human chuckled: Analyzing garden path effects in humans and language models (2025.acl-long)

Copied to clipboard

Challenge: Modern Large Language Models (LLMs) have shown human-like abilities in many language tasks, sparking interest in comparing LLMs’ and humans’ language processing.
Approach: They propose to answer two questions: 1. What makes garden-path sentences hard for humans? 2. Do the same reasons make garden- path sentences hard?
Outcome: The proposed models show that humans struggle with specific syntactic complexities, with some models showing high correlation with human comprehension.
A Monte-Carlo Sampling Framework For Reliable Evaluation of Large Language Models Using Behavioral Analysis (2025.findings-emnlp)

Copied to clipboard

Challenge: Current approaches to evaluation of large language models ignore high entropy of LLM responses.
Approach: They propose a Monte-Carlo evaluation framework for evaluating large language models . they test multiple LLMs to see if they are susceptible to cognitive biases .
Outcome: The proposed framework shows that LLMs are more human-like and less rational . it also shows that larger LLM models are more susceptible to cognitive biases .
BIG5-CHAT: Shaping LLM Personalities Through Training on Human-Grounded Data (2025.acl-long)

Copied to clipboard

Challenge: Existing methods for embedding human personality traits into LLMs are limited by realism and validity issues.
Approach: They propose to use a large-scale dataset to embed human personality traits into LLMs . they use supervised fine-tuning and direct preference optimization to train LLM models .
Outcome: The proposed methods outperform prompting on personality assessments and IPIP-NEO, and show higher conscientiousness, agreeableness, lower extraversion, and lower neuroticism on reasoning tasks.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations