Challenge: Existing models of large language reasoning exhibit reasoning biases similar to humans, a study shows . syllogistic reasoning is one of the basic forms of deductive reasoning .
Approach: They propose to use a syllogism dataset to evaluate models' reasoning abilities . they propose to ask LLMs to translate slogismatic slurs into abstract logical expressions .
Outcome: The proposed method shows that models exhibit reasoning biases similar to humans, and that there is room for improvement in reasoning problems where premises and hypotheses are neither entailment nor contradiction.

Similar Papers

A Systematic Analysis of Large Language Models as Soft Reasoners: The Case of Syllogistic Inferences (2024.emnlp-main)

Copied to clipboard

Challenge: syllogistic reasoning is a deductive reasoning skill that is crucial in everyday problem-solving and decision-making experiences.
Approach: They propose to study the reasoning abilities of Large Language Models (LLMs) they propose to use supervised fine-tuning and chain-of-thought reasoning to investigate their results.
Outcome: The proposed models exhibit reasoning biases, avoid answering that no conclusion follows, align with human difficulties, and struggle with multi-step reasoning.
A Systematic Comparison of Syllogistic Reasoning in Humans and Language Models (2024.naacl-long)

Copied to clipboard

Challenge: Psychologists have documented several ways in which humans’ inferences deviate from the rules of logic.
Approach: They focus on syllogisms, which are inferences from two simple premises, and show that larger models are more logical than smaller ones.
Outcome: The results show that language models often mimic human biases, but overcome them in some cases.
A Peek into Token Bias: Large Language Models Are Not Yet Genuine Reasoners (2024.emnlp-main)

Copied to clipboard

Challenge: a new hypothesis-testing framework is developed to assess whether large language models possess genuine reasoning abilities or primarily depend on token bias.
Approach: They propose a framework to assess whether large language models have genuine reasoning abilities or primarily depend on token bias.
Outcome: The proposed framework outlines a list of hypotheses where token biases are readily identifiable . the results suggest that most LLMs still struggle with logical reasoning .
Logical forms complement probability in understanding language model (and human) performance (2025.acl-long)

Copied to clipboard

Challenge: Existing studies on LLMs have shown that they perform well on logical reasoning problems, but there is still a lack of fine-grained understanding of the logical forms.
Approach: They propose a dataset of hypothetical and disjunctive syllogisms in propositional and modal logic and use it as the testbed for understanding LLM performance.
Outcome: The proposed model performs well on proposi-tional and modal logics, but does it exhibit preferences for certain argument forms?
Large Language Models Are Partially Primed in Pronoun Interpretation (2023.findings-acl)

Copied to clipboard

Challenge: Existing studies suggest large language models acquire rich linguistic representations, but little is known about whether they adapt to linguistic biases in a human-like way.
Approach: They examine whether large language models display human-like referential biases using stimuli and procedures from real psycholinguistic experiments.
Outcome: The proposed models display human-like referential biases when exposed to referential patterns in the local context.
Can Activation Steering Generalize Across Languages? A Study on Syllogistic Reasoning in Language Models (2026.eacl-long)

Copied to clipboard

Challenge: Prior work has focused on activation steering for Large Language Models (LLMs) this technique can be used to improve reasoning accuracy and transferability across languages.
Approach: They propose to use activation steering to steer models towards a cross-lingual reasoning space.
Outcome: The proposed techniques generalise well to multilingual datasets while minimizing language modelling performance.
SylloBio-NLI: Evaluating Large Language Models on Biomedical Syllogistic Reasoning (2025.naacl-long)

Copied to clipboard

Challenge: Existing models are far from achieving the robustness and consistency required for safe biomedical NLI applications.
Approach: They propose a framework that leverages external ontologies to instantiate diverse syllogistic arguments for biomedical NLI by identifying valid conclusions and extracting supporting evidence.
Outcome: The proposed framework evaluates large language models on identifying valid conclusions and extracting supporting evidence across 28 syllogistic schemes instantiated with human genome pathways.
Assessing Step-by-Step Reasoning against Lexical Negation: A Case Study on Syllogism (2023.emnlp-main)

Copied to clipboard

Challenge: Large language models (LLMs) take advantage of step-by-step reasoning instructions . negation is a core linguistic phenomenon that is difficult to process .
Approach: They examine the step-by-step reasoning ability of large language models with a focus on negation . negation is a core linguistic phenomenon that is difficult to process .
Outcome: The proposed models perform better when using chain-of-thought prompting . the results highlight unique limitations in each LLM family .
Persona-Assigned Large Language Models Exhibit Human-Like Motivated Reasoning (2026.findings-acl)

Copied to clipboard

Challenge: Prior studies have reported that large language models (LLMs) are also susceptible to human-like cognitive biases, but the extent to which LLMs selectively reason toward identity-congruent conclusions remains unexplored.
Approach: They investigate whether assigning 8 personas across 4 political and socio-demographic attributes induces motivated reasoning in LLMs.
Outcome: The proposed model is assigned 8 personas across 4 political and socio-demographic attributes and shows that they have 9% reduced veracity discernment compared to models without persona.
Towards the Roots of the Negation Problem: A Multilingual NLI Dataset and Model Scaling Analysis (2025.findings-emnlp)

Copied to clipboard

Challenge: Negations are key to determining sentence meaning, making them essential for logical reasoning.
Approach: They construct and publish two new textual entailment datasets in four languages with paired examples differing in negation.
Outcome: The results show that increasing the model size may improve the models’ ability to handle negations.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations