Challenge: Existing methods focus on whether the reasoning chain leads to the correct conclusion, but this view may confound reasoning quality with other spurious shortcuts to predict the answer.
Approach: They propose a framework that evaluates reasoning chains via two key properties: (1) correctness, i.e., each step makes a valid inference based on information contained within the step, preceding steps, and input context, and (2) informativeness, respectively.
Outcome: The proposed framework evaluates reasoning chains via two key properties: (1) correctness, i.e., each step makes a valid inference based on information contained within the step, preceding steps, and input context, and (2) informativeness, which is helpful towards deriving the generated answer.

Similar Papers

A Chain-of-Thought Is as Strong as Its Weakest Link: A Benchmark for Verifiers of Reasoning Chains (2024.acl-long)

Copied to clipboard

Challenge: Recent literature discusses automatic methods to evaluate reasoning to improve their correctness, but no fine-grained step-level datasets are available to enable thorough evaluation of such verification methods.
Approach: They propose to benchmark automatic verifiers of complex Chain-of-Thought reasoning in open-domain question-answering settings using a dataset that includes comprehensive labels for relevance, attribution to evidence passages, and logical correctness of each reasoning step.
Outcome: The proposed dataset shows that verifiers struggle at verifying reasoning chains, particularly verifying logical correctness and detecting contradictions.
Evaluating Step-by-step Reasoning Traces: A Survey (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing evaluation practices are inconsistent, resulting in fragmented progress across evaluator design and benchmark development.
Approach: a survey provides a comprehensive overview of step-by-step reasoning evaluation . existing evaluation practices are inconsistent, resulting in fragmented progress .
Outcome: The proposed evaluation criteria are based on four top-level categories . the results are presented in a systematic review of the literature.
Step-by-Step Reasoning to Solve Grid Puzzles: Where do LLMs Falter? (2024.emnlp-main)

Copied to clipboard

Challenge: Existing studies evaluate only the final predicted answer of a puzzle, without providing any finer metrics to evaluate them.
Approach: They propose to use a grid-based evaluation dataset to evaluate LLMs' reasoning abilities and a new error taxonomy to evaluate their reasoning chains.
Outcome: The proposed model outperforms existing prompting methods on a wide range of natural language understanding tasks previously thought to be exclusive to humans.
Verifying the Steps of Deductive Reasoning Chains (2025.findings-acl)

Copied to clipboard

Challenge: Large Language Models have been shown to improve the reasoning capabilities of the models.
Approach: They propose to automate verification of individual reasoning steps in a logical deductive Chain-of-Thought.
Outcome: The proposed method can detect unsound reasoning steps fairly well, but under-performs symbolic methods.
Stepwise Informativeness Search for Improving LLM Reasoning (2025.emnlp-main)

Copied to clipboard

Challenge: Recent advances in large language models have improved multistep reasoning but they lose focus over the middle of long contexts.
Approach: They propose a tree search framework that proactively identifies underutilized steps and minimizing redundant information between steps.
Outcome: The proposed framework generates more accurate and concise rationales with reduced errors and redundancy.
CheckEval: A reliable LLM-as-a-Judge framework for evaluating text generation using checklists (2025.emnlp-main)

Copied to clipboard

Challenge: Existing evaluation protocols for text generation suffer from rating inconsistencies . lexical overlap-based metrics align poorly with human judgments .
Approach: They propose a checklist-based evaluation framework that improves rating reliability via decomposed binary questions.
Outcome: The proposed framework improves rating reliability by decomposing binary questions . it improves agreement across evaluator models by 0.45 and reduces score variance . human evaluation remains the gold standard, but it #, Equal contribution.
LogicBench: Towards Systematic Evaluation of Logical Reasoning Ability of Large Language Models (2024.acl-long)

Copied to clipboard

Challenge: Existing work investigating the logical reasoning ability of large language models has focused only on a couple of inference rules of propositional and first-order logics.
Approach: They propose to use a natural language question-answering dataset to evaluate the logical reasoning ability of large language models.
Outcome: The proposed model performs poorly on a range of natural language questions using chain-of-thought prompting.
CiteEval: Principle-Driven Citation Evaluation for Source Attribution (2025.acl-long)

Copied to clipboard

Challenge: Current evaluation frameworks rely on NLI to assess binary or ternary support from cited sources, which is suboptimal for citation evaluation.
Approach: They propose a citation evaluation framework based on fine-grained citation ratings within a broad context and construct a multi-domain benchmark with high-quality human annotations.
Outcome: The proposed framework provides a high-quality human annotation benchmark and a suite of model-based metrics that exhibit strong correlation with human judgments.
FACT-E: Causality-Inspired Evaluation for Trustworthy Chain-of-Thought Reasoning (2026.findings-acl)

Copied to clipboard

Challenge: Existing models generate explanations that appear coherent while containing unfaithful intermediate steps.
Approach: They propose a causality-inspired framework for evaluating CoT quality using controlled perturbations as an instrumental signal to separate genuine step-to-step dependence from bias-driven artifacts.
Outcome: Experiments on GSM8K, MATH, and CommonsenseQA show that FACT-E improves reasoning-trajectory selection and yields stronger in-context learning exemplars.
How Interpretable are Reasoning Explanations from Prompting Large Language Models? (2024.findings-naacl)

Copied to clipboard

Challenge: Prompt Engineering has garnered significant attention for enhancing the performance of large language models across a multitude of tasks.
Approach: They propose a simple prompting technique that yields more than 70% improvement in interpretability.
Outcome: The proposed method improves interpretability by 70% across multiple dimensions.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations