Challenge: Existing studies on Large Language Models (LLMs) have not investigated the impact of question types on LLM performance.
Approach: They evaluate the performance of five Large Language Models on reasoning tasks . they use quantitative reasoning tasks and deductive reasoning tasks to evaluate the models .
Outcome: The results show that Reasoning accuracy does not correlate with final selection accuracy.

Similar Papers

Evaluating the Deductive Competence of Large Language Models (2024.naacl-long)

Copied to clipboard

Challenge: Existing large language models have limited abilities to solve deductive reasoning problems . performance differences between conditions do not improve overall performance .
Approach: They investigate whether several large language models can solve a deductive reasoning problem in their conventional form.
Outcome: The proposed models can solve a classic type of deductive reasoning problem in their conventional form.
Can Multiple-choice Questions Really Be Useful in Detecting the Abilities of LLMs? (2024.lrec-main)

Copied to clipboard

Challenge: Multiple-choice questions (MCQs) are widely used in the evaluation of large language models (LLMs) however, there are concerns about whether MCQ can truly measure LLM’s capabilities.
Approach: They propose to use multiple choice questions to evaluate large language models (LLMs) to assess their capabilities.
Outcome: The proposed methods show that MCQs are less reliable than LFGQs in terms of expected calibration error.
What Factors Influence LLMs’ Judgments? A Case Study on Question Answering (2024.lrec-main)

Copied to clipboard

Challenge: Existing studies indicate that Large Language Models perform at a level comparable to humans with advantages of speed and cost-effectiveness in different fields.
Approach: They propose to introduce four unexplored factors and a new dimension of question difficulty to provide a more comprehensive understanding of LLMs’ judgments across varying question intricacies.
Outcome: The proposed dimensions of question difficulty and answer quantity provide valuable insights into optimizing LLMs’ performance as judges.
Do Large Language Models excel in Complex Logical Reasoning with Formal Language? (2025.emnlp-main)

Copied to clipboard

Challenge: Existing studies on LLMs have focused on formal language, but evaluations of their performance are limited.
Approach: They propose to use a formal language to evaluate LLMs across logical reasoning problems using formal languages.
Outcome: The proposed model outperforms Instruct models in three dimensions, taxonomy of tasks, and format of trajectories, and achieves the best generalization performance across other languages.
Comparing Inferential Strategies of Humans and Large Language Models in Deductive Reasoning (2024.acl-long)

Copied to clipboard

Challenge: Recent advances in the domain of large language models (LLMs) have showcased their capability in executing deductive reasoning tasks.
Approach: They examine inferential strategies employed by large language models through a detailed evaluation of their responses to propositional logic problems.
Outcome: The proposed model shows that it displays reasoning patterns similar to humans, including strategies like supposition following or chain construction.
Are Large Language Models Consistent over Value-laden Questions? (2024.findings-emnlp)

Copied to clipboard

Challenge: Large language models (LLMs) appear to bias survey answers toward certain values . however, some argue that LLMs are inconsistent to simulate particular values - a recent study .
Approach: They define value consistency as similarity of answers across paraphrases, related questions and multilingual translations of a question to English, Chinese, German, and Japanese.
Outcome: The proposed model is consistent across paraphrases, use-cases, translations, and within a topic.
Towards Reasoning in Large Language Models: A Survey (2023.findings-acl)

Copied to clipboard

Challenge: Reasoning is a fundamental aspect of human intelligence that plays a crucial role in many intellectual activities.
Approach: They propose to improve LLMs' ability to elicit reasoning by providing exemplars or prompts to model reasoning.
Outcome: This paper provides a comprehensive overview of the state of knowledge on reasoning in large language models.
Puzzle Solving using Reasoning of Large Language Models: A Survey (2024.emnlp-main)

Copied to clipboard

Challenge: Recent advances in Large Language Models (LLMs) have demonstrated their logical reasoning abilities across various domains.
Approach: They propose to divide puzzles into rule-based and rule-less categories and critically assess LLMs' performance through various methodologies.
Outcome: The proposed models have demonstrated capabilities in deductive reasoning and inductive reasoning, but they face limitations in inductive thinking.
An Empirical Study of LLM Reasoning Ability Under Strict Output Length Constraint (2025.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) are a powerful tool for test-time scaling, but they are often used under time constraints.
Approach: They propose to use LLMs to make models think before answering questions . they also use self-correction and best-of-N decoding to encourage deeper thinking .
Outcome: The proposed models are able to achieve higher inference accuracy with extra inference computation under time constraints.
Can LLMs Reason Like Doctors? Exploring the Limits of Large Language Models in Complex Medical Reasoning (2026.findings-eacl)

Copied to clipboard

Challenge: Large language models (LLMs) have shown remarkable progress in reasoning across multiple domains, but it remains unclear whether their abilities reflect genuine reasoning or sophisticated pattern matching.
Approach: They conduct one of the largest evaluations to date, assessing 77 LLMs . they select three medical question answering (QA) benchmarks targeting reasoning processes .
Outcome: The results highlight the need to improve specific reasoning strategies to better reflect medical decision-making.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations