Evaluating Reasoning Models for Queries with Presuppositions (2026.findings-acl)

Copied to clipboard

Challenge: Prior work notes that large language models fail to challenge erroneous assumptions and can reinforce users’ misinformed opinions.
Approach: They construct queries with varying degrees of presuppositions spanning health, science, and general knowledge and evaluate several widely-deployed models.
Outcome: The proposed models achieve higher accuracy but fail to challenge a large fraction of false presuppositions.

Similar Papers

Evaluating Large Language Models for Health-related Queries with Presuppositions (2024.findings-acl)

Copied to clipboard

Challenge: a large number of health-related queries require factually accurate answers . however, the lack of accurate answers may cause real-world harm .
Approach: They evaluate the factual accuracy and consistency of large language models using a dataset consisting of health-related queries with varying degrees of presuppositions.
Outcome: The proposed model responses agree with 23-32% of existing false claims and 49-55% with novel fabricated claims.
Towards Reasoning in Large Language Models: A Survey (2023.findings-acl)

Copied to clipboard

Challenge: Reasoning is a fundamental aspect of human intelligence that plays a crucial role in many intellectual activities.
Approach: They propose to improve LLMs' ability to elicit reasoning by providing exemplars or prompts to model reasoning.
Outcome: This paper provides a comprehensive overview of the state of knowledge on reasoning in large language models.
Factuality of Large Language Models: A Survey (2024.emnlp-main)

Copied to clipboard

Challenge: Large language models (LLMs) are factually incorrect, which limits their applicability in real-world scenarios.
Approach: They analyze existing work to identify major challenges and their associated causes . they propose to evaluate LLMs using a variety of measures to mitigate factual errors .
Outcome: The proposed methods are based on a variety of datasets and proposed strategies to mitigate factual errors.
Measuring Bias and Agreement in Large Language Model Presupposition Judgments (2025.findings-acl)

Copied to clipboard

Challenge: Identifying linguistic bias in text requires the identification of explicit statements and presuppositions . large language models can be used to detect subtle forms of bias with no clear lexical signals .
Approach: They propose to prompt large language models to evaluate presuppositions across texts . they find that LLMs may inadvertently reflect societal biases when identifying presuposed content .
Outcome: The proposed model can be used to detect linguistic biases in text, but its accuracy is unclear . linguistic factors associated with human-model alignment suggest biase influenced by gender and ideology.
Comparing Inferential Strategies of Humans and Large Language Models in Deductive Reasoning (2024.acl-long)

Copied to clipboard

Challenge: Recent advances in the domain of large language models (LLMs) have showcased their capability in executing deductive reasoning tasks.
Approach: They examine inferential strategies employed by large language models through a detailed evaluation of their responses to propositional logic problems.
Outcome: The proposed model shows that it displays reasoning patterns similar to humans, including strategies like supposition following or chain construction.
A Peek into Token Bias: Large Language Models Are Not Yet Genuine Reasoners (2024.emnlp-main)

Copied to clipboard

Challenge: a new hypothesis-testing framework is developed to assess whether large language models possess genuine reasoning abilities or primarily depend on token bias.
Approach: They propose a framework to assess whether large language models have genuine reasoning abilities or primarily depend on token bias.
Outcome: The proposed framework outlines a list of hypotheses where token biases are readily identifiable . the results suggest that most LLMs still struggle with logical reasoning .
A Psycholinguistic Evaluation of Language Models’ Sensitivity to Argument Roles (2024.findings-emnlp)

Copied to clipboard

Challenge: a systematic evaluation of large language models' sensitivity to argument roles is presented . a recent study shows that argument roles have a delayed impact on verb prediction in human sentence processing.
Approach: They propose to replicate psycholinguistic studies on human argument role processing . they find that language models are able to distinguish verbs that appear in plausible and implausible contexts .
Outcome: The proposed models are able to distinguish verbs that appear in plausible and implausible contexts, but none captures the same selective patterns that human comprehenders exhibit during real-time verb prediction.
Current Advances in LLM Reasoning (2026.acl-tutorials)

Copied to clipboard

Challenge: This tutorial examines comprehensive evaluation strategies to assess the reasoning abilities of large language models (LLMs) advanced inference time methods and post-training methods that aim to make LLMs think more like humans are discussed in this tutorial.
Approach: This tutorial explores comprehensive evaluation strategies to assess the reasoning abilities of large language models (LLMs) and discusses two types of methods to improve models’ reasoning: advanced inference time methods, structured and self-improvement inference methods, and post-training methods, such as RLHF, DPO, and GRPO.
Outcome: This tutorial examines evaluation strategies to assess the reasoning abilities of large language models and discusses two types of methods to improve models’ reasoning.
Do Language Models Have Semantics? On the Five Standard Positions (2025.acl-long)

Copied to clipboard

Challenge: Large language models (LLMs) are trained to solve the so-called cloze task . solving clozing tasks is essentially a memorization task, says a recent study .
Approach: They propose to use five positions to determine whether large language models exhibit semantic understanding . large language model is trained to solve the so-called cloze task .
Outcome: The proposed theory is based on a pairwise comparison of five positions on semantic understanding in large language models and chatbots.
Large Language Models for Mathematical Reasoning: Progresses and Challenges (2024.eacl-srw)

Copied to clipboard

Challenge: a survey examines the landscape of mathematical problem-solving techniques . large language models have proven to be potent assets in unraveling nuances of mathematical reasoning .
Approach: They examine the evolution of Large Language Models (LLMs) for solving mathematical problems . they examine the spectrum of LLM-oriented techniques proposed for solving math problems - and their challenges .
Outcome: The survey examines the spectrum of proposed LLM-oriented techniques in solving math problems.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations