Challenge: Large Language Models (LLMs) exhibit impressive linguistic fluency, but it remains unclear whether they possess human-like Theory of Mind (ToM) or rely on statistical heuristics . a recent study examined the performance of LLMs against 300 human participants .
Approach: a study establishes a framework for large language models that modulates contextual contrast, linguistic cues, and cognitive mechanisms.
Outcome: a new evaluation framework compares ten state-of-the-art LLMs against 300 human participants . the framework systematically modulates contextual contrast, linguistic cues, and cognitive mechanisms .

Similar Papers

Cognitive Bias in Decision-Making with LLMs (2024.findings-emnlp)

Copied to clipboard

Challenge: Large language models inherit societal biases against protected groups and can be subject to functionally resembling cognitive bias.
Approach: They propose a framework to uncover, evaluate, and mitigate cognitive bias in large language models by using a dataset containing 13,465 prompts to evaluate LLM decisions on different cognitive biases.
Outcome: The proposed framework uncovers, evaluates, and mitigates cognitive bias in large language models, particularly in high-stakes decision-making tasks.
Systematic Biases in LLM Simulations of Debates (2024.emnlp-main)

Copied to clipboard

Challenge: Current research suggests that LLM-based agents become increasingly human-like in their performance, sparking interest in using these AI agents as substitutes for human participants in behavioral studies.
Approach: They propose to use LLMs to simulate political debates on topics that are important aspects of people’s day-to-day lives and decision-making processes.
Outcome: The proposed model can simulate political debates on topics that are important aspects of people’s day-to-day lives and decision-making processes.
Benchmarking Cognitive Biases in Large Language Models as Evaluators (2024.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) have been shown to be effective as automatic evaluators with simple prompting and in-context learning.
Approach: They assemble 16 Large Language Models and evaluate their outputs by preference ranking . they introduce a cognitive bias benchmark to measure six different cognitive biases in LLM evaluation outputs.
Outcome: The proposed model is biased on the CoBBLer benchmark, indicating that machine preferences are misaligned with humans.
I’m sure you’re a real scholar yourself: Exploring Ironic Content Generation by Large Language Models (2024.findings-emnlp)

Copied to clipboard

Challenge: Moreover, irony is highly subjective and can depend on various factors, such as social, cultural, or generational aspects.
Approach: They propose to fine-tune two large language models to generate ironic and non-ironic content and analyze their outputs from a linguistic perspective.
Outcome: The proposed models generate ironic and non-ironic responses to a given social media post and analyze their outputs from a linguistic perspective.
Do LLMs Align Human Values Regarding Social Biases? Judging and Explaining Social Biases with LLMs (2025.findings-emnlp)

Copied to clipboard

Challenge: Large language models can lead to undesired consequences when misaligned with human values . previous studies have shown misalignment of LLMs with human value using expert-designed or agent-based emulated bias scenarios .
Approach: They investigate whether large language models (LLMs) are misaligned with human values . they find no significant differences in understanding of HVSB between LLMs .
Outcome: The results show that large language models do not have lower misalignment rates and attack success rates . the study also shows that smaller language models have the ability to explain HVSB .
LLMs in Sarcasm Detection? It’s elementary! (Or is it?) (2026.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) are often cited for their sophisticated pragmatic reasoning, but they collapse to random guessing on organic human speech.
Approach: They propose that LLMs have near-human competence in sarcasm detection . authors propose that this proficiency may be deceptive .
Outcome: The proposed model performance on synthetic leaderboards is a statistical mirage of competence.
A Monte-Carlo Sampling Framework For Reliable Evaluation of Large Language Models Using Behavioral Analysis (2025.findings-emnlp)

Copied to clipboard

Challenge: Current approaches to evaluation of large language models ignore high entropy of LLM responses.
Approach: They propose a Monte-Carlo evaluation framework for evaluating large language models . they test multiple LLMs to see if they are susceptible to cognitive biases .
Outcome: The proposed framework shows that LLMs are more human-like and less rational . it also shows that larger LLM models are more susceptible to cognitive biases .
A Survey in Automatic Irony Processing: Linguistic, Cognitive, and Multi-X Perspectives (2022.coling-1)

Copied to clipboard

Challenge: figurative language research has focused on sarcasm and irony, but there is still a gap in the field.
Approach: They propose to review computational irony, cognitive science, and neural models of irony processing . they aim to encourage a balanced and equal research environment in figurative languages .
Outcome: The proposed multi-X irony processing perspectives will provide an overview of computational irony, insights from linguisic theory and cognitive science, and interactions with downstream NLP tasks.
Language Model Council: Democratically Benchmarking Foundation Models on Highly Subjective Tasks (2025.naacl-long)

Copied to clipboard

Challenge: Existing evaluations of Large Language Models (LLMs) rely on a single large model to score outputs from other LLMs, but this is prone to intra-model bias and many tasks may be too subjective for a one model to judge fairly.
Approach: They propose a language model council where a group of LLMs collaborate to create tests, respond to them, and evaluate each other’s responses to produce a ranking in a democratic fashion.
Outcome: The proposed model produces rankings that are more separable and robust than any individual LLM judge.
Bias in the Mirror : Are LLMs opinions robust to their own adversarial attacks (2025.acl-long)

Copied to clipboard

Challenge: Existing work on large language models lacks robustness, highlighting the limitations of such models.
Approach: They propose a novel approach where two LLMs engage in self-debate to persuade a neutral version of the model.
Outcome: The proposed approach examines whether large language models are robust during interactions and whether they are susceptible to reinforcing misinformation or shifting to harmful viewpoints.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations