Evaluating Reasoning Models for Queries with Presuppositions (2026.findings-acl)
Copied to clipboard
| Challenge: | Prior work notes that large language models fail to challenge erroneous assumptions and can reinforce users’ misinformed opinions. |
| Approach: | They construct queries with varying degrees of presuppositions spanning health, science, and general knowledge and evaluate several widely-deployed models. |
| Outcome: | The proposed models achieve higher accuracy but fail to challenge a large fraction of false presuppositions. |
Similar Papers
Evaluating Large Language Models for Health-related Queries with Presuppositions (2024.findings-acl)
Copied to clipboard
| Challenge: | a large number of health-related queries require factually accurate answers . however, the lack of accurate answers may cause real-world harm . |
| Approach: | They evaluate the factual accuracy and consistency of large language models using a dataset consisting of health-related queries with varying degrees of presuppositions. |
| Outcome: | The proposed model responses agree with 23-32% of existing false claims and 49-55% with novel fabricated claims. |
Towards Reasoning in Large Language Models: A Survey (2023.findings-acl)
Copied to clipboard
| Challenge: | Reasoning is a fundamental aspect of human intelligence that plays a crucial role in many intellectual activities. |
| Approach: | They propose to improve LLMs' ability to elicit reasoning by providing exemplars or prompts to model reasoning. |
| Outcome: | This paper provides a comprehensive overview of the state of knowledge on reasoning in large language models. |
Factuality of Large Language Models: A Survey (2024.emnlp-main)
Copied to clipboard
Yuxia Wang, Minghan Wang, Muhammad Arslan Manzoor, Fei Liu, Georgi Georgiev, Rocktim Das, Preslav Nakov
| Challenge: | Large language models (LLMs) are factually incorrect, which limits their applicability in real-world scenarios. |
| Approach: | They analyze existing work to identify major challenges and their associated causes . they propose to evaluate LLMs using a variety of measures to mitigate factual errors . |
| Outcome: | The proposed methods are based on a variety of datasets and proposed strategies to mitigate factual errors. |
Measuring Bias and Agreement in Large Language Model Presupposition Judgments (2025.findings-acl)
Copied to clipboard
| Challenge: | Identifying linguistic bias in text requires the identification of explicit statements and presuppositions . large language models can be used to detect subtle forms of bias with no clear lexical signals . |
| Approach: | They propose to prompt large language models to evaluate presuppositions across texts . they find that LLMs may inadvertently reflect societal biases when identifying presuposed content . |
| Outcome: | The proposed model can be used to detect linguistic biases in text, but its accuracy is unclear . linguistic factors associated with human-model alignment suggest biase influenced by gender and ideology. |
Comparing Inferential Strategies of Humans and Large Language Models in Deductive Reasoning (2024.acl-long)
Copied to clipboard
| Challenge: | Recent advances in the domain of large language models (LLMs) have showcased their capability in executing deductive reasoning tasks. |
| Approach: | They examine inferential strategies employed by large language models through a detailed evaluation of their responses to propositional logic problems. |
| Outcome: | The proposed model shows that it displays reasoning patterns similar to humans, including strategies like supposition following or chain construction. |
A Peek into Token Bias: Large Language Models Are Not Yet Genuine Reasoners (2024.emnlp-main)
Copied to clipboard
Bowen Jiang, Yangxinyu Xie, Zhuoqun Hao, Xiaomeng Wang, Tanwi Mallick, Weijie Su, Camillo Taylor, Dan Roth
| Challenge: | a new hypothesis-testing framework is developed to assess whether large language models possess genuine reasoning abilities or primarily depend on token bias. |
| Approach: | They propose a framework to assess whether large language models have genuine reasoning abilities or primarily depend on token bias. |
| Outcome: | The proposed framework outlines a list of hypotheses where token biases are readily identifiable . the results suggest that most LLMs still struggle with logical reasoning . |
A Psycholinguistic Evaluation of Language Models’ Sensitivity to Argument Roles (2024.findings-emnlp)
Copied to clipboard
| Challenge: | a systematic evaluation of large language models' sensitivity to argument roles is presented . a recent study shows that argument roles have a delayed impact on verb prediction in human sentence processing. |
| Approach: | They propose to replicate psycholinguistic studies on human argument role processing . they find that language models are able to distinguish verbs that appear in plausible and implausible contexts . |
| Outcome: | The proposed models are able to distinguish verbs that appear in plausible and implausible contexts, but none captures the same selective patterns that human comprehenders exhibit during real-time verb prediction. |
Current Advances in LLM Reasoning (2026.acl-tutorials)
Copied to clipboard
| Challenge: | This tutorial examines comprehensive evaluation strategies to assess the reasoning abilities of large language models (LLMs) advanced inference time methods and post-training methods that aim to make LLMs think more like humans are discussed in this tutorial. |
| Approach: | This tutorial explores comprehensive evaluation strategies to assess the reasoning abilities of large language models (LLMs) and discusses two types of methods to improve models’ reasoning: advanced inference time methods, structured and self-improvement inference methods, and post-training methods, such as RLHF, DPO, and GRPO. |
| Outcome: | This tutorial examines evaluation strategies to assess the reasoning abilities of large language models and discusses two types of methods to improve models’ reasoning. |
Do Language Models Have Semantics? On the Five Standard Positions (2025.acl-long)
Copied to clipboard
| Challenge: | Large language models (LLMs) are trained to solve the so-called cloze task . solving clozing tasks is essentially a memorization task, says a recent study . |
| Approach: | They propose to use five positions to determine whether large language models exhibit semantic understanding . large language model is trained to solve the so-called cloze task . |
| Outcome: | The proposed theory is based on a pairwise comparison of five positions on semantic understanding in large language models and chatbots. |
Large Language Models for Mathematical Reasoning: Progresses and Challenges (2024.eacl-srw)
Copied to clipboard
| Challenge: | a survey examines the landscape of mathematical problem-solving techniques . large language models have proven to be potent assets in unraveling nuances of mathematical reasoning . |
| Approach: | They examine the evolution of Large Language Models (LLMs) for solving mathematical problems . they examine the spectrum of LLM-oriented techniques proposed for solving math problems - and their challenges . |
| Outcome: | The survey examines the spectrum of proposed LLM-oriented techniques in solving math problems. |