Papers with Reasoning

300 papers
MoNaCo: More Natural and Complex Questions for Reasoning Across Dozens of Documents (2026.tacl-1)

Copied to clipboard

Challenge: Automated agents powered by large language models are becoming more ingrained into how people seek information . but evaluation benchmarks for LLMs rarely feature natural questions that are time-consuming . a new benchmark, MoNaCo, aims to address this gap by eliciting and manually answering time-wasting questions .
Approach: They propose a benchmark of 1,315 natural and time-consuming questions that require dozens of intermediate steps to solve.
Outcome: MoNaCo benchmarks achieve at least 61.2% F1 in real-world time-consuming questions hampered by low recall and hallucinations . Frontier LLMs evaluated on MoN achieving at least 61% F1, harmed by low memory and halluzinations.
CHATREPORT: Democratizing Sustainability Disclosure Analysis through LLM-based Tools (2023.emnlp-demo)

Copied to clipboard

Challenge: a lack of transparency in sustainability reporting is a key challenge due to the sheer volume and complexity of sustainability reports . only a few entities worldwide have the resources to analyze these reports at scale . a novel LLM-based system to automate the analysis of corporate sustainability reports is needed .
Approach: They propose a novel LLM-based system to automate the analysis of corporate sustainability reports.
Outcome: The proposed system automates the analysis of corporate sustainability reports.
Logic Haystacks: Probing LLMs’ Long-Context Logical Reasoning (Without Easily Identifiable Unrelated Padding) (2026.eacl-short)

Copied to clipboard

Challenge: Recent large language models claim long context windows, but evaluations often involve simple retrieval tasks or synthetic tasks padded with irrelevant text.
Approach: They use grammars to generate simplified English with logical representations to create long input text while controlling its semantics.
Outcome: The proposed model performs better with realistic distractors than with standard models.
How LLMs React to Industrial Spatio-Temporal Data? Assessing Hallucination with a Novel Traffic Incident Benchmark Dataset (2025.naacl-industry)

Copied to clipboard

Challenge: Large language models (LLMs) have a high potential to digitize and enhance the health & public services industry.
Approach: They propose to use a cross-lingual benchmark dataset to assess the robustness of state-of-the-art LLMs in the spatio vs temporal domain for traffic incident classification.
Outcome: The proposed model performs well in the spatio-temporal domain and in the non-English context.
Current Advances in LLM Reasoning (2026.acl-tutorials)

Copied to clipboard

Challenge: This tutorial examines comprehensive evaluation strategies to assess the reasoning abilities of large language models (LLMs) advanced inference time methods and post-training methods that aim to make LLMs think more like humans are discussed in this tutorial.
Approach: This tutorial explores comprehensive evaluation strategies to assess the reasoning abilities of large language models (LLMs) and discusses two types of methods to improve models’ reasoning: advanced inference time methods, structured and self-improvement inference methods, and post-training methods, such as RLHF, DPO, and GRPO.
Outcome: This tutorial examines evaluation strategies to assess the reasoning abilities of large language models and discusses two types of methods to improve models’ reasoning.
Simple yet Effective Bridge Reasoning for Open-Domain Multi-Hop Question Answering (D19-58)

Copied to clipboard

Challenge: Existing work on open-domain multi-hop question answering relies on off-the-shelf information retrieval techniques to retrieve answer passages.
Approach: They propose a new subproblem for open-domain multi-hop question answering . they aim to recognize the anchor from a set of start passages with a reading comprehension model .
Outcome: The proposed method significantly improves the baseline method on the open-domain hotpotQA benchmark.
SymPyBench: A Dynamic Benchmark for Scientific Reasoning with Executable Python Code (2026.eacl-industry)

Copied to clipboard

Challenge: Existing benchmarks do not capture the complexity of structured, step-by-step reasoning essential in physics and related domains.
Approach: They propose a large-scale synthetic benchmark of 15K university-level physics problems . they use structured, step-by-step reasoning and executable Python code to produce the ground-truth solution.
Outcome: The proposed model is based on a set of 15K university-level physics problems with three question types.
A Multilingual Reading Comprehension System for more than 100 Languages (2020.coling-demos)

Copied to clipboard

Challenge: Recent advances in open domain question answering (QA) have focused on machine reading comprehension (MRC)
Approach: They propose a multilingual machine reading comprehension (MRC) demo which can answer questions in over 100 languages.
Outcome: The proposed system can answer questions in over 100 languages and integrates with IBM Watson's machine translation widget to improve language accessibility.
CIKQA: Learning Commonsense Inference with a Unified Knowledge-in-the-loop QA Paradigm (2023.findings-eacl)

Copied to clipboard

Challenge: Existing commonsense reasoning datasets target different knowledge types, modalities, and formats, but how to help machines acquire and infer over commonsensical knowledge is still unclear.
Approach: They propose a commonsense reasoning benchmark to motivate commonsensing progress from two perspectives: (1) Evaluating whether models can distinguish knowledge quality by predicting if the knowledge is enough to answer the question or not.
Outcome: The proposed model outperforms existing models in evaluating their generalization capabilities across tasks while demonstrating that distinguishing knowledge quality remains challenging for current models.
Book QA: Stories of Challenges and Opportunities (D19-58)

Copied to clipboard

Challenge: Existing approaches to answer questions based on the full text of books are limited by their unique characteristics.
Approach: They propose a system for answering questions based on the full text of books . they use a memory network to reason and predict an answer, and a novel question generator to improve generalization.
Outcome: The proposed system improves on the recently published NarrativeQA corpus on Who questions . it shows that the proposed system is highly challenging and needs more research .
Benchmarking LLM’s Capability in Reasoning over Conflicting Web References (2026.acl-long)

Copied to clipboard

Challenge: Large language models (LLMs) integrated with retrieval-augmented generation (RAG) are a dominant framework for building intelligent assistants.
Approach: They propose a benchmark to evaluate LLMs' reasoning capability over real-world conflicting documents retrieved from the web.
Outcome: The proposed benchmark evaluates LLMs' reasoning capability over real-world conflicting documents retrieved from the web.
“President Vows to Cut <Taxes> Hair”: Dataset and Analysis of Creative Text Editing for Humorous Headlines (N19-1)

Copied to clipboard

Challenge: Existing datasets address specific humor templates, such as funny one-liners and filling in Mad Libs R.
Approach: They introduce a dataset for research in computational humor that uses crowdsourced editing techniques to create funny headlines.
Outcome: The new dataset supports classic theories of humor, including incongruity, superiority, setup/punchline.
Commonsense inference in human-robot communication (D19-60)

Copied to clipboard

Challenge: a gap exists in natural language understanding of commands between humans and machines.
Approach: They propose a method for commonsense inference to transform high-level commands into action commands for robotic systems to execute.
Outcome: The proposed method allows to build a knowledge base that consists of a large set of commonsense inferences.
Implicit Temporal Reasoning for Evidence-Based Fact-Checking (2023.findings-eacl)

Copied to clipboard

Challenge: Temporal reasoning is implicit since models learn from data how to leverage temporal information.
Approach: They propose to ground claims and associated evidence on shared timelines using publication dates and time expressions extracted from their text.
Outcome: The proposed model outperforms existing models that explicitly model temporal relations between evidence and the document by up to 9% Micro F1 and 15% Macro F1 on the MultiFC dataset.
Comprehensive Supersense Disambiguation of English Prepositions and Possessives (P18-1)

Copied to clipboard

Challenge: Frequent prepositions like for are maddeningly polysemous, their interpretation depends especially on the object of the preposition.
Approach: They propose a new annotation scheme, corpus, and task for the disambiguation of prepositions and possessives in English.
Outcome: The proposed annotations are comprehensive with respect to types and tokens of these markers and use broadly applicable supersense classes rather than fine-grained dictionary definitions.
Error Classification of Large Language Models on Math Word Problems: A Dynamically Adaptive Framework (2025.findings-emnlp)

Copied to clipboard

Challenge: Current error classification methods rely on static and predefined categories to capture error patterns.
Approach: They propose a framework for automated dynamic error classification in mathematical reasoning that incorporates common error patterns as explicit guidance.
Outcome: The proposed framework reduces human bias and fine-grained analysis of error patterns.
Verifying the Steps of Deductive Reasoning Chains (2025.findings-acl)

Copied to clipboard

Challenge: Large Language Models have been shown to improve the reasoning capabilities of the models.
Approach: They propose to automate verification of individual reasoning steps in a logical deductive Chain-of-Thought.
Outcome: The proposed method can detect unsound reasoning steps fairly well, but under-performs symbolic methods.
COGEN: Abductive Commonsense Language Generation (2023.acl-short)

Copied to clipboard

Challenge: Existing training methods for NLP models to perform on two main tasks are needed to introduce these capabilities into the field of reasoning.
Approach: They propose a model that integrates commonsense reasoning with contextual filtering to improve the inference.
Outcome: The proposed model outperforms existing models and sets new state-of-the-art in regards to alphaNLI and alphaNGG tasks.
Stimulating Creativity with FunLines: A Case Study of Humor Generation in Headlines (2020.acl-demos)

Copied to clipboard

Challenge: FunLines is an online game that allows players to generate and rate funny news headlines . it is difficult to generate data that depends on human creativity, and measuring creativity often requires more effort.
Approach: They propose a game where players edit news headlines to make them funny and rate the funniness of headlines edited by others.
Outcome: The proposed game outperforms other crowdsourcing approaches in generating humor datasets.
Disentangling the Effects of Unlearning in Measuring Parametric Faithfulness of Chain-of-Thought (2026.acl-srw)

Copied to clipboard

Challenge: Chain-of-Thought (CoT) has been debated as a model's faithfulness to internal reasoning process.
Approach: They propose to use unlearning to measure parametric faithfulness of models by adjusting for unintended artifacts of unlearning.
Outcome: The proposed metric accounts for the unintended artifacts of unlearning and shows that it is non-negligible.
Pragmatic Reasoning Unlocks Quantifier Semantics for Foundation Models (2023.emnlp-main)

Copied to clipboard

Challenge: Generalized quantifiers are used to indicate the proportions predicates satisfy (e.g., some apples are red).
Approach: They propose a framework to model quantifier semantics for textbased foundation models by combining natural language inference and the Rational Speech Acts framework.
Outcome: The proposed framework shows a 20% improvement over a literal listener baseline in predicting percentage scopes for quantifier comprehension even with no training.
Rethinking Loss Functions for Fact Verification (2024.eacl-short)

Copied to clipboard

Challenge: Existing objective functions for fact verification fail to capture heterogeneity among verdict classes . cross-entropy loss treats all misclassification types uniformly, which is problematic .
Approach: They propose two task-specific objective functions that capture the heterogeneity among verdict classes . they use a dictionary-based objective function to classify Wikipedia sentences into three verdict classes.
Outcome: The proposed objectives outperform the standard cross-entropy loss objective . the proposed objectives are combined with simple class weighting to overcome imbalance .
Common Sense or Ableism? Rethinking Commonsense Reasoning Through the Lens of Disability (2026.eacl-short)

Copied to clipboard

Challenge: a recent study finds that commonsense reasoning is not always universal and can leave disabled people behind . a case study of disabled people with long COVID shows that common sense is not universal .
Approach: They investigate how datasets and models deal with disability in commonsense reasoning . they use annotations from disabled and non-disabled persons for ableism .
Outcome: The proposed datasets have low sensitivity to human-detected ableism but still detect 5 to 25% of entries as ableist.
LLM DEBATE OPPONENT : Counter-argument Generation focusing on Implicit and Critical Premises (2025.naacl-srw)

Copied to clipboard

Challenge: Recent advances in Large Language Models (LLMs) show promise in automating counter-argument generation.
Approach: They compare multi-step and one-step generation methods for counter-arguments across 100 debate topics.
Outcome: The proposed model outperforms multi-step and one-step pipelines for counter-arguments across 100 debate topics.
ET5: A Novel End-to-end Framework for Conversational Machine Reading Comprehension (2022.coling-1)

Copied to clipboard

Challenge: Existing methods require three steps to understand text, but span extraction and question rephrasing steps are not fully exploited.
Approach: They propose a framework for conversational machine reading comprehension based on shared parameter mechanism . experimental results show the proposed framework achieves new state-of-the-art results on the ShARC leaderboard .
Outcome: The proposed framework achieves state-of-the-art on the ShARC leaderboard with the BLEU-4 score of 55.2.
SHARP: Unlocking Interactive Hallucination via Stance Transfer in Role-Playing LLMs (2025.findings-acl)

Copied to clipboard

Challenge: Existing studies on social interactions neglect hallucination while struggling with poor generalizability and implicit character fidelity judgments.
Approach: They propose a generalizable and explicit paradigm for uncovering interactive patterns of Large Language Models across diverse worldviews by defining interactive hallucination through stance transfer and SHARP, a benchmark built by extracting relations from commonsense knowledge graphs.
Outcome: The proposed paradigm is generalizable and explicit and demonstrates its effectiveness and stability.
TARo: Token-level Adaptive Routing for LLM Test-time Alignment (2026.findings-acl)

Copied to clipboard

Challenge: Large language models (LLMs) exhibit strong reasoning capabilities but typically require expensive post-training to reach high performance.
Approach: They propose to use token-level Adaptive Routing to steer frozen LLMs toward structured reasoning entirely at inference time.
Outcome: Extensive experiments show that TARo significantly improves reasoning performance by up to +22.4% over base model and +8.4% .
Evidence-Driven Reasoning for Industrial Maintenance Using Heterogeneous Data (2026.acl-industry)

Copied to clipboard

Challenge: Existing maintenance systems do not support conditional reasoning, argues a new study . large language models (LLMs) offer flexible reasoning, but naively applying generative models introduces risks, he says .
Approach: They propose a maintenance language-based reasoning framework that constrains reasoning through deterministic evidence construction and structured failure knowledge.
Outcome: The proposed framework produces evidence-grounded explanations and advisory actions under heterogeneous data, a study shows . it constrains reasoning through deterministic evidence construction and structured failure knowledge, and applies a rule-based verification loop to suppress unsupported conclusions.
Benchmarking LLM Faithfulness in RAG with Evolving Leaderboards (2025.emnlp-industry)

Copied to clipboard

Challenge: Large language models (LLMs) excel in various tasks, but often produce hallucinations . retrieved contexts, misrepresent information, or generate outright contradictions .
Approach: They propose a framework that measures hallucination faithfulness of large language models . they introduce a leaderboard that leverages diverse human-annotated hallucinian examples .
Outcome: The proposed framework improves hallucination evaluations by leveraging human-annotated examples.
OVM, Outcome-supervised Value Models for Planning in Mathematical Reasoning (2024.findings-naacl)

Copied to clipboard

Challenge: Large language models (LLMs) struggle with maintaining accuracy throughout multiple reasoning steps, especially in mathematical reasoning where an error in earlier steps can propagate to subsequent ones and ultimately leading to an incorrect answer.
Approach: They propose an Outcome-supervised Value Model (OVM) that employs outcome supervision for training a value model, which prioritizes steps that lead to accurate conclusions.
Outcome: The proposed model performs better on two multi-step reasoning datasets, GSM8K and Game of 24.
A Corpus for Modeling User and Language Effects in Argumentation on Online Debating (P19-1)

Copied to clipboard

Challenge: Existing argumentation datasets have allowed only limited assessment of "user" traits because information on background of users is generally unavailable.
Approach: They present a dataset of 78,376 debates generated over a 10-year period along with surprisingly comprehensive participant profiles.
Outcome: The proposed dataset includes 78,376 debates generated over a 10-year period along with comprehensive participant profiles.
How Well Do Multi-hop Reading Comprehension Models Understand Date Information? (2022.aacl-short)

Copied to clipboard

Challenge: Existing multi-hop reading comprehension datasets have reasoning shortcuts that can be used to answer comparison questions without performing multi- hop reasoning.
Approach: They propose a dataset with three probing tasks in addition to the main question . they then evaluate the model's ability to understand date information .
Outcome: The proposed model performs well in date comparison and number subtraction tasks.
LOGICAL-COMMONSENSEQA: A Benchmark for Logical Commonsense Reasoning (2026.acl-short)

Copied to clipboard

Challenge: LOGICAL-COMMONSENSEQA benchmarks evaluate commonsense reasoning as logical composition over pairs of atomic statements . commonsensible reasoning is central to human cognition and a long-standing challenge in artificial intelligence and natural language understanding.
Approach: They propose a benchmark that reframes commonsense reasoning as logical composition over pairs of atomic statements using plausibility-level operators.
Outcome: LOGICAL-COMMONSENSEQA exposes fundamental reasoning limitations and provides a framework for advancing compositional commonsense reasoning.
Arithmetic Reasoning with LLM: Prolog Generation & Permutation (2024.naacl-short)

Copied to clipboard

Challenge: Existing work has shown that large language models can generate arithmetic and commonsense reasoning, but they are not native to mathematical operations and symbolic manipulations.
Approach: They propose to use large language models to generate Prolog programs to solve math problems using a code interpreter to generate arithmetic and symbolic formulas.
Outcome: The proposed model outperforms CoT generation in the GSM8K benchmark across three LLMs.
Guided Knowledge Generation with Language Models for Commonsense Reasoning (2024.findings-emnlp)

Copied to clipboard

Challenge: Large Language Models (LLMs) have achieved notable success in commonsense reasoning tasks, benefiting from extensive world knowledge acquired through extensive pretraining.
Approach: They propose a method to generate knowledge explanations and to automatically assign labels based on the probability of correct answers.
Outcome: The proposed method outperforms baselines on four widely-used commonsense reasoning benchmarks and shows that it can generate high quality knowledge leading to correct answers.
AM4DSP: Argumentation Mining in Structured Decentralized Discussion Platforms for Deliberative Democracy (2025.emnlp-demos)

Copied to clipboard

Challenge: Argument mining is the automated process of identification and extraction of argumentative structures in natural language.
Approach: They propose to use argument mining to extract arguments from online discussions in the context of deliberative democracy.
Outcome: The proposed system enables the extraction and analysis of arguments from online discussions in the context of deliberative democracy.
J4R: Learning to Judge with Equivalent Initial State Group Relative Policy Optimization (2026.acl-long)

Copied to clipboard

Challenge: Large language models (LLMs) are increasingly being used for reasoning intensive tasks.
Approach: They propose an algorithm that trains judges to be robust to positional biases . they also propose a benchmark that evaluates judges in diverse reasoning settings .
Outcome: The proposed algorithm outperforms GPT-4o and the next best small judge by 6.7% and 9% on ReasoningJudgeBench and JudgeBench.
Think Wider, Detect Sharper: Reinforced Reference Coverage for Document-Level Self-Contradiction Detection (2025.emnlp-main)

Copied to clipboard

Challenge: Recent approaches to document-level contradiction detection (DSCD) only gain marginal improvement and often introduce inconsistencies across repeated responses.
Approach: They propose a method that combines supervised fine-tuning and reinforcement learning to enhance document-level contradiction detection (DSCD) they propose to use a task-specific reward function to expand the model’s reasoning scope, boosting both accuracy and consistency.
Outcome: The proposed method significantly boosts Llama 3.1-8B-Instruct’s accuracy from 38.5% to 51.1%, and consistency from 59.6% to76.2%.
MapAgent: A Hierarchical Agent for Geospatial Reasoning with Dynamic Map Tool Integration (2026.findings-eacl)

Copied to clipboard

Challenge: Existing frameworks for large language models are tailored to domains such as mathematics, coding, or web automation.
Approach: They propose a hierarchical multi-agent plug-and-play framework with customized toolsets and agentic scaffolds for map-integrated geospatial reasoning.
Outcome: The proposed framework decouples planning from execution and reduces cognitive load on users.
Controlled Hallucinations: Learning to Generate Faithfully from Noisy Data (2020.findings-emnlp)

Copied to clipboard

Challenge: Neural text generation (data- or text-to-text) demonstrates remarkable performance when training data is abundant which for many applications is not the case.
Approach: They propose a technique to treat hallucinations as a controllable aspect of the generated text without dismissing any input and without modifying the model architecture.
Outcome: The proposed technique can be used on a WikiBio dataset and in a human evaluation.
Make Your Decision Convincing! A Unified Two-Stage Framework: Self-Attribution and Decision-Making (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing frameworks for explaining black-box model behavior are unreliable . large-scale pre-trained models often rely on superficial clues for predictions .
Approach: They propose a unified two-stage framework that uses subsequences from the input text as a rationale to generate model decision.
Outcome: The proposed framework achieves competitive results on five reasoning datasets and in semi-supervised scenarios.
COM2SENSE: A Commonsense Reasoning Benchmark with Complementary Sentences (2021.findings-acl)

Copied to clipboard

Challenge: Recent advances in pretrained language models have shown promising results on commonsense reasoning benchmark datasets.
Approach: They propose a commonsense reasoning benchmark dataset with 4k sentence pairs . they propose 'gamified' model-in-the-loop setup to incentivize challenging samples .
Outcome: The proposed benchmarks show that the proposed model achieves 71% standard accuracy and 51% pairwise accuracy, well below human performance.
Invoke Interfaces Only When Needed: Adaptive Invocation for Large Language Models in Question Answering (2025.findings-emnlp)

Copied to clipboard

Challenge: a new metric is developed to pinpoint the moment of invocation when hallucinations arise in small LMs.
Approach: They propose a metric that measures hallucinations during the generation process of small LMs.
Outcome: The proposed metric outperforms baselines in hallucination detection across multiple QA datasets.
DANLI: Deliberative Agent for Following Natural Language Instructions (2022.emnlp-main)

Copied to clipboard

Challenge: Recent work on embodied AI agents that can perform tasks by following human language instructions is limited by reactive methods, which are insufficient for long-horizon complex tasks.
Approach: They propose a neuro-symbolic deliberative agent that, while following language instructions, proactively applies reasoning and planning based on its neural and symbolic representations acquired from past experience.
Outcome: The proposed agent achieves greater than 70% improvement over reactive baselines on the challenging TEACh benchmark.
Enhancing Event Causality Identification with Counterfactual Reasoning (2023.acl-short)

Copied to clipboard

Challenge: Existing methods for event causality identification (ECI) focus on mining potential causal signals, but causal signals are ambiguous, which may lead to the context-keywords bias and the event-pairs bias.
Approach: They propose a method that explicitly estimates the influence of context keywords and event pairs in training to eliminate biases in inference.
Outcome: The proposed method eliminates biases in inference on two datasets.
Legal Mathematical Reasoning with LLMs: Procedural Alignment through Two-Stage Reinforcement Learning (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing legal mathematical reasoning models lack structured numerical reasoning . existing models perform poorly on LexNum, while LexPam improves both mathematical accuracy and legal coherence.
Approach: They propose a legal mathematical reasoning benchmark LexNum and LexPam to address this problem . LexPam is a two-stage reinforcement learning framework for efficient legal reasoning training.
Outcome: The proposed framework improves mathematical accuracy and legal coherence . it also improves legal cohesion and generalizes effectively across tasks and domains.
One Missing Piece for Open-Source Reasoning Models: A Dataset to Mitigate Cold-Starting Short CoT LLMs in RL (2025.acl-industry)

Copied to clipboard

Challenge: Existing large reasoning models are limited by their closed nature and high API costs and safety issues.
Approach: They propose to build a long CoT dataset with existing short CoT LLMs that are not trained for inference-time scaling.
Outcome: The proposed model achieves quality comparable to—or slightly below—R1 and is able to think longer and provide control over the thought budget to better manage the overthinking problem.
IIRC: A Dataset of Incomplete Information Reading Comprehension Questions (2020.emnlp-main)

Copied to clipboard

Challenge: Existing reading comprehension tasks focus on questions for which the contexts provide all the information required to answer them, thus not evaluating a system’s performance at identifying a potential lack of sufficient information and locating sources for that information.
Approach: They propose to use a dataset with 13K questions over paragraphs from English Wikipedia that provide only partial information to answer them, with the missing information occurring in one or more linked documents.
Outcome: The proposed model achieves 31.1% F1 on the reading comprehension task, while estimated human performance is 88.4%.
Garbage In, Reasoning Out? Why Benchmark Scores are Unreliable and What to Do About It (2026.findings-eacl)

Copied to clipboard

Challenge: Using social reasoning benchmarks, we uncover pervasive flaws in both benchmark items and evaluation methodology.
Approach: They audit three widely used social reasoning benchmarks and identify flaws in their design and evaluation methodology.
Outcome: The results challenge the validity of current benchmark-based claims about social reasoning in large language models.
CriticBench: Benchmarking LLMs for Critique-Correct Reasoning (2024.findings-acl)

Copied to clipboard

Challenge: CriticBench is a benchmark designed to assess LLMs’ abilities to critique and refine their reasoning across a variety of tasks.
Approach: They propose a benchmark to assess LLMs' ability to critique and correct reasoning across a variety of tasks.
Outcome: The proposed benchmark examines the performance of 17 large language models in generation, critique, and correction reasoning.
From Surrogacy to Adoption; From Bitcoin to Cryptocurrency: Debate Topic Expansion (P19-1)

Copied to clipboard

Challenge: Recent advances in argumentation mining have left much of the relevant argumentative content out of reach.
Approach: They propose a task of Debate Topic Expansion to find related topics for a given debate topic, along with an annotated dataset for the task.
Outcome: The proposed algorithms differ from well-studied lexical-semantic relations and show they work well in argumentation mining.
Explainable Hallucination through Natural Language Inference Mapping (2025.findings-acl)

Copied to clipboard

Challenge: Large language models (LLMs) often generate hallucinated content, making it crucial to identify and quantify inconsistencies in their outputs.
Approach: They propose a framework that maps entailment and contradiction relations between inputs and outputs using a natural language inference model.
Outcome: The proposed framework outperforms state-of-the-art methods by five percentage points while providing clear, interpretable explanations.
Best-of-L: Cross-Lingual Reward Modeling for Mathematical Reasoning (2026.findings-eacl)

Copied to clipboard

Challenge: Recent studies have focused on improving reasoning ability in English models, with multilingual models receiving comparatively little attention.
Approach: They propose a framework that ranks candidate reasoning traces across languages rather than within a single language.
Outcome: The proposed framework improves accuracy by up to 10 points in English compared to using reward modeling within a single language.
Chaining Event Spans for Temporal Relation Grounding (2024.eacl-long)

Copied to clipboard

Challenge: Existing approaches to understanding temporal relations between events have relied on answer overlaps as a proxy label to distinguish similar and dissimilar questions.
Approach: They propose a timeline reasoning network that elicits proper reasoning behaviors through a module for predicting time spans of events.
Outcome: The proposed approach outperforms existing methods by resolving spurious overlaps using the predicted timeline.
Exploring End-to-End Differentiable Natural Logic Modeling (2020.coling-main)

Copied to clipboard

Challenge: Existing approaches to integrate natural logic with neural networks are brittle and prone to fail in the presence of noise and uncertainty.
Approach: They propose to integrate natural logic with neural networks to create differentiable models that integrate natural reasoning with subsymbolic vector representations and neural components.
Outcome: The proposed model can model monotonicity-based reasoning, compared to baseline models without inductive bias.
TurnaboutLLM: A Deductive Reasoning Benchmark from Detective Games (2025.emnlp-main)

Copied to clipboard

Challenge: evaluating large language models' reasoning abilities via detective stories is often infeasible due to the large answer space and diverse reasoning types presented by its questions.
Approach: They propose a framework and dataset for evaluating the deductive reasoning abilities of Large Language Models (LLMs) by leveraging the interactive gameplay of detective games Ace Attorney and Danganronpa.
Outcome: The proposed framework and dataset are based on the detective games Ace Attorney and Danganronpa and show that they are more efficient than current strategies for enhancing deductive reasoning.
AMO-Bench: Large Language Models Still Struggle in High School Math Competitions (2026.findings-acl)

Copied to clipboard

Challenge: Existing benchmarks for mathematical reasoning are becoming less effective due to performance saturation.
Approach: They propose to use a mathematical reasoning benchmark with Olympiad difficulty to evaluate top-tier LLMs.
Outcome: The proposed benchmarks are cross-validated by experts to meet IMO difficulty standards and entirely original problems to prevent performance leakages from data memorization.
When Facts Change: Temporal Knowledge Conflict Resolution in LLMs (2026.findings-acl)

Copied to clipboard

Challenge: Large language models are increasingly used in retrieval-augmented generation systems to reconcile knowledge conflicts between parametric memory and contextual inputs.
Approach: They propose to use mutability to resolve temporal misalignment in large language models to compare stable and recently updated facts from Wikidata to determine if mutable models can serve as a mediating signal in this process.
Outcome: The proposed model can produce reasoning for facts that actually changed but rarely for stable ones, whereas smaller models rarely detect conflict, while larger models detect it but fail to act on mutability judgments.
RoleCDE: Benchmarking and Mitigating Role–Alignment Trade-offs in Role-Playing Agents (2026.findings-acl)

Copied to clipboard

Challenge: Existing benchmarks for role-playing agents only evaluate surface-level fidelity and provide limited insight into decision making under role–alignment value conflicts.
Approach: They propose a benchmark to evaluate RPAs under role–alignment value conflicts . they use 8k diverse role profiles and 240k dilemma instances to evaluate role-aware decision making .
Outcome: The proposed benchmark covers 8k diverse role profiles and scenarios and nearly 240k dilemma instances across three difficulty levels and eight role categories.
Assessing Monotonicity Reasoning in Dutch through Natural Language Inference (2023.findings-eacl)

Copied to clipboard

Challenge: a novel dataset for natural language inference (NLI) is used to study monotonicity reasoning in Dutch.
Approach: They investigate monotonicity reasoning in Dutch using a novel dataset . they find that models struggle with downward entailing contexts .
Outcome: The proposed dataset shows that models struggle with downward entailing contexts, and argue that this is due to a poor understanding of negation.
Answering Numerical Reasoning Questions in Table-Text Hybrid Contents with Graph-based Encoder and Tree-based Decoder (2022.coling-1)

Copied to clipboard

Challenge: Existing methods for numerical reasoning are not flexible enough to handle diverse expressions.
Approach: They propose a Relational Graph enhanced Hybrid table-text Numerical reasoning model with Tree decoder which captures relationship between numerical value, table schema, and text information on the encoder side.
Outcome: The proposed model outperforms the baseline model and achieves state-of-the-art results on the publicly available tabletext hybrid QA benchmark.
Beyond "Not Novel Enough": Enriching Scholarly Critique with LLM-Assisted Feedback (2026.eacl-long)

Copied to clipboard

Challenge: Novelty assessment is a central yet understudied aspect of peer review . manuscript submissions double roughly every 15 years, and individual reviewers now complete an average of 14 reviews per year.
Approach: They propose a structured approach for automated novelty evaluation that models expert reviewer behavior through three stages: content extraction, retrieval and synthesis of related work, and structured comparison for evidence-based assessment.
Outcome: The proposed approach outperforms existing LLM-based baselines on 182 ICLR 2025 submissions with human-annotated reviewer novelty assessments.
GLGR: Question-aware Global-to-Local Graph Reasoning for Multi-party Dialogue Reading Comprehension (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing approaches for multi-hop reasoning are lacking for local graph reasoning . existing approaches neglect local semantic structures in utterances .
Approach: They propose a question-aware global-to-local graph reasoning approach that expands the canonical Interlocutor-Utterance graph by introducing a query node.
Outcome: The proposed approach outperforms existing methods on Molweni and FriendsQA.
SymBa: Symbolic Backward Chaining for Structured Natural Language Reasoning (2025.naacl-long)

Copied to clipboard

Challenge: Among different methods for structured reasoning, we focus on backward chaining, where the goal is recursively decomposed into subgoals by searching and applying rules.
Approach: They propose a backward chaining system that integrates a symbolic solver and an LLM to improve the performance of LLM-based reasoning.
Outcome: The proposed system improves deductive, relational, and arithmetic reasoning benchmarks compared to baselines.
Modeling Hierarchical Reasoning Chains by Linking Discourse Units and Key Phrases for Reading Comprehension (2022.coling-1)

Copied to clipboard

Challenge: Existing methods of logical reasoning focus on entity-aware information but ignore hierarchical relations that may even have mutual effects.
Approach: They propose a holistic graph network that deals with context at both discourse-level and word-level as the basis for logical reasoning.
Outcome: The proposed method improves on logical reasoning QA datasets and natural language inference datasets.
MemoBrain: Executive Memory as an Agentic Brain for Reasoning (2026.findings-acl)

Copied to clipboard

Challenge: Large language models (LLMs) are inherently long-horizon, causing reasoning traces and tool artifacts to accumulate and strain the working context of large language models.
Approach: They propose a model that constructs a dependency-aware memory over reasoning steps and captures salient intermediate states and their logical relations.
Outcome: The proposed model prunes invalid steps, folds completed sub-trajectories, and preserves a compact, high-salience reasoning backbone under a fixed context budget.
Machine-Aided Annotation for Fine-Grained Proposition Types in Argumentation (2020.lrec-1)

Copied to clipboard

Challenge: a corpus of 2016 debates and commentary contains 4,648 argumentative propositions annotated with fine-grained proposition types.
Approach: They propose a machine learning-human workflow for annotating for four complex proposition types . they demonstrate with preliminary analysis of rhetorical strategies and structure in presidential debates .
Outcome: The proposed method can be used by technical researchers seeking more nuanced representations of argument . it can also be used to analyze rhetorical strategies and structure in presidential debates .
HalluDetect: Detecting, Mitigating, and Benchmarking Hallucinations in Conversational Systems in the Legal Domain (2025.emnlp-industry)

Copied to clipboard

Challenge: Large Language Models (LLMs) are widely used in industry but still produce hallucinations, limiting their reliability in critical applications.
Approach: They propose to reduce hallucinations in consumer grievance chatbots by reducing their token accuracy by 0.4159 per turn.
Outcome: The proposed system achieves an F1 score of 68.92% outperforming baseline detectors by 22.47% while maintaining the highest token accuracy.
The Discussion Tracker Corpus of Collaborative Argumentation (2020.lrec-1)

Copied to clipboard

Challenge: The Discussion Tracker corpus is an annotated dataset of transcripts of spoken, multi-party argumentation transcribed from 985 minutes of audio .
Approach: They analyze 29 multi-party arguments transcribed from 985 minutes of audio . they provide descriptive statistics and code for predicting each dimension separately.
Outcome: The Discussion Tracker corpus was collected in high school English classes and annotated for argument moves, specificity, specificities and collaboration dimensions.
MathMist: A Parallel Multilingual Benchmark Dataset for Mathematical Problem Solving and Reasoning (2026.findings-eacl)

Copied to clipboard

Challenge: Existing benchmarks primarily focus on English or a narrow subset of high-resource languages, leaving significant gaps in assessing multilingual and cross-lingual mathematical reasoning.
Approach: They propose a parallel multilingual benchmark for mathematical problem solving and reasoning that encompasses 2,890 parallel Bangla-English gold standard artifacts.
Outcome: The proposed model encompasses 2,890 parallel Bangla-English gold standard artifacts, totaling 30K aligned question–answer pairs across thirteen languages, representing high-, medium-, and low-resource linguistic settings.
Translating a Math Word Problem to a Expression Tree (D18-1)

Copied to clipboard

Challenge: Sequence-to-sequence (SEQ2SEQ) models have been successfully applied to automatic math word problem solving.
Approach: They propose an equation normalization method to normalize duplicated equations and propose an ensemble model to combine their advantages.
Outcome: The proposed model outperforms the previous state-of-the-art models on the math word problem solving.
Rethinking Sample Polarity in Reinforcement Learning with Verifiable Rewards (2026.acl-long)

Copied to clipboard

Challenge: Large reasoning models are typically trained using reinforcement learning with verifiable reward (RLVR) positive and negative self-generated rollouts are used to update the model's policy . positive samples sharpen existing correct reasoning patterns, while negative samples encourage exploration of new reasoning paths.
Approach: They propose a method that allocates advantage signals to key tokens across different polarities.
Outcome: The proposed method improves the ability of large reasoning models to learn from their own generated rollouts.
Exploiting Numerical-Contextual Knowledge to Improve Numerical Reasoning in Question Answering (2022.findings-naacl)

Copied to clipboard

Challenge: Existing numerical reasoning models overly rely on parametric knowledge at inference time . previous studies show that understanding numbers in text improves numerical reasoning accuracy .
Approach: They propose a numerical reasoning model that leverages parametric knowledge to alleviate this over-reliance on parametric information.
Outcome: The proposed model improves numerical reasoning accuracy and performance in DROP.
Zero-knowledge LLM hallucination detection and mitigation through fine-grained cross-model consistency (2025.emnlp-industry)

Copied to clipboard

Challenge: Existing methods for hallucination management fail to integrate both detection and mitigation without external knowledge sources.
Approach: They propose a black-box framework that leverages fine-grained cross-model consistency to detect and mitigate hallucinations in LLM outputs without external knowledge sources.
Outcome: The proposed framework improves hallucination detection scores by 6-39% on a FELM dataset . it achieves 9 percentage points improvement in answer accuracy on the GPQA-diamond dataset compared to existing approaches .
Fake News Detection using Deep Markov Random Fields (N19-1)

Copied to clipboard

Challenge: Existing deep-learning-based methods ignore the correlations among news articles and only consider each article individually.
Approach: They propose a graph-theoretic method that inherits the power of deep learning while utilizing the correlations among the articles.
Outcome: The proposed model improves on state-of-the-art models on well-known datasets.
Confidence-Aware Reasoning: Optimizing Self-Guided Thinking Trajectories in Large Reasoning Models (2025.emnlp-industry)

Copied to clipboard

Challenge: Chain-of-thought enables large reasoning models to reason through multi-step problems but often leads to unnecessarily long or redundant reasoning traces, a phenomenon known as overthinking.
Approach: They propose an inference-time framework that selectively prunes low-utility reasoning blocks and halts early when sufficient confidence has been achieved.
Outcome: The proposed framework improves answer accuracy by up to +13.3% while reducing average reasoning length by 40%–50%.
Beyond Surface-Level Pattern Trap: LLM Agents for Faster and Smarter Cross-Architecture Code Migration (2026.findings-acl)

Copied to clipboard

Challenge: cross-architecture code migration is a resource-intensive and errorprone task.
Approach: a framework for cross-architecture code migration is proposed to decouple implementation details through functional mining and code refactoring.
Outcome: a new framework improves performance and correctness over state-of-the-art frameworks on OpenCV migration tasks.
Ask, Assess, and Refine: Rectifying Factual Consistency and Hallucination in LLMs with Metric-Guided Feedback Learning (2024.eacl-long)

Copied to clipboard

Challenge: Recent advances in Large Language Models (LLMs) have heralded unprecedented capabilities in information seeking and text generation, but challenges remain regarding citation errors and generating information not present in the evidence (hallucination).
Approach: They propose a framework to assess citation errors and hallucination using an explicit evaluation paradigm to formulate actionable natural language feedback.
Outcome: The proposed approach improves correctness, fluency, and citation quality and reduces hallucinations in the results.
Commonsense Subgraph for Inductive Relation Reasoning with Meta-learning (2025.coling-main)

Copied to clipboard

Challenge: Existing subgraph-based models focus on predicting missing relations in knowledge graphs . a new meta-learning model extracts concepts from entities to construct commonsense subgraphs based on semantic information .
Approach: They propose a commonsense subgraph meta-learning model that extracts concepts from entities to construct commonsensible subgraphs.
Outcome: The proposed model outperforms existing models in inductive reasoning tasks and in few-shot scenarios.
MathGenie: Generating Synthetic Data with Question Back-translation for Enhancing Mathematical Reasoning of LLMs (2024.acl-long)

Copied to clipboard

Challenge: Existing models have demonstrated outstanding capabilities in mathematical reasoning, but there is a performance gap between open-source models and closed-source ones.
Approach: They propose a method for generating diverse and reliable math problems by leveraging the ground-truth solutions of the seed data.
Outcome: The proposed model outperforms open-source models across five representative mathematical reasoning datasets.
Knowledge Verification to Nip Hallucination in the Bud (2024.emnlp-main)

Copied to clipboard

Challenge: Recent studies have shown that large language models generate responses that sound plausible but contradict factual knowledge, a phenomenon known as hallucination.
Approach: They propose a novel approach to align large language models to evaluate knowledge boundaries based on external knowledge to reduce hallucinations .
Outcome: The proposed approach reduces hallucinations across six benchmarks using foundation LLMs of varying backbones and scales.
QUITE: Quantifying Uncertainty in Natural Language Text in Bayesian Reasoning Scenarios (2024.emnlp-main)

Copied to clipboard

Challenge: Existing probabilistic reasoning datasets require the model to only rank textual alternatives or use limited set of templates.
Approach: They propose a question-answering dataset that uses probabilistic rules to express degrees of certainty.
Outcome: The proposed model outperforms existing models on all reasoning types . it is available on Github and is expected to be used in clinical documentation .
Two Heads Are Better Than One: Dual-Model Verbal Reflection at Inference-Time (2025.emnlp-main)

Copied to clipboard

Challenge: Existing LLMs struggle to reliably detect subtle reasoning errors in ASAS tasks.
Approach: They propose a dual-model framework with a dedicated Critic model trained for effective reflection that generates precise verbal feedback.
Outcome: The proposed framework outperforms existing ASAS benchmarks and provides valuable insights into the performance of the proposed framework.
Generating Commonsense Reasoning Questions with Controllable Complexity through Multi-step Structural Composition (2025.coling-main)

Copied to clipboard

Challenge: Existing work mainly learns to map text into questions, lacking a mechanism to control results with desired complexity.
Approach: They propose a novel controllable framework to generate QGs with desired complexity using contextual and commonsense clues from text.
Outcome: The proposed framework can generate complex questions with desired complexity levels.
GLAF: Global-to-Local Aggregation and Fission Network for Semantic Level Fact Verification (2022.coling-1)

Copied to clipboard

Challenge: Existing fact verification models lack fine-grained reasoning over key entities . GLAF uses local fission reasoning to capture latent logical relations between clues .
Approach: They propose a global-to-local fission and fissional network to capture latent logical relations hidden in multiple evidence clues.
Outcome: The proposed network achieves state-of-the-art on a FEVER dataset with a 77.62% FEVER score.
A Probabilistic Framework for LLM Hallucination Detection via Belief Tree Propagation (2025.naacl-long)

Copied to clipboard

Challenge: Current large language models (LLMs) produce factually incorrect statements .
Approach: They propose a probabilistic framework for LLM hallucination detection that generates a belief tree by expanding a statement into logically related claims and reasoning globally about the relationships between these claims.
Outcome: The proposed method improves on multiple hallucination detection benchmarks by 3%-9% over state-of-the-art models.
LLMs Are Not Intelligent Thinkers: Introducing Mathematical Topic Tree Benchmark for Comprehensive Evaluation of LLMs (2025.naacl-long)

Copied to clipboard

Challenge: Large language models (LLMs) have impressive capabilities in mathematical reasoning, but their effectiveness is limited to specific mathematical topics.
Approach: They propose to use the MaTT benchmark to assess large language models' accuracy in multiple-choice scenarios.
Outcome: The proposed model achieved 54% accuracy in a multiple-choice scenario, while the Chain-of-Thought prompting did not improve.
Specializing Pre-trained Language Models for Better Relational Reasoning via Network Pruning (2022.findings-naacl)

Copied to clipboard

Challenge: Pretrained masked language models inherit a considerable amount of relational knowledge from the source corpora.
Approach: They propose to specialize pretrained masked language models into relational models from the perspective of network pruning.
Outcome: The proposed model can represent grounded commonsense relations at non-trivial sparsity while being generalizable . the proposed model improves on a wealth of NLP tasks, but we know little about how much knowledge it imparts .
IRMA: the 335-million-word Italian coRpus for studying MisinformAtion (2023.eacl-main)

Copied to clipboard

Challenge: IRMA contains over 600,000 Italian news articles collected from 56 websites classified as ‘untrustworthy’ by professional fact-checkers.
Approach: They present IRMA, a corpus of italian news articles collected from 56 websites classified as ‘untrustworthy’ by professional fact-checkers.
Outcome: The IRMA corpus contains over 600,000 Italian news articles collected from 56 websites classified as ‘untrustworthy’ by professional fact-checkers.
A Dog Is Passing Over The Jet? A Text-Generation Dataset for Korean Commonsense Reasoning and Evaluation (2022.findings-naacl)

Copied to clipboard

Challenge: Korean pretrained language models struggle to generate short sentences with a given condition based on compositionality and commonsense reasoning.
Approach: They propose a Korean text-generation dataset for Korean generative commonsense reasoning and language model evaluation using a semi-automatic dataset construction approach.
Outcome: The proposed dataset is available at http://aihub.or.kr/opendata/korea-university.
SAGE: Steerable Agentic Data Generation for Deep Search with Execution Feedback (2026.findings-eacl)

Copied to clipboard

Challenge: High-quality, complex question-answer pairs are pivotal for training and evaluating capable deep search agents.
Approach: They propose a pipeline that generates high-quality, difficulty-controlled deep search question-answer pairs for a given corpus and a target difficulty level.
Outcome: The proposed pipeline generates high-quality, difficulty-controlled deep search question-answer pairs for a given corpus and a target difficulty level.
A Comprehensive Literary Chinese Reading Comprehension Dataset with an Evidence Curation Based Solution (2025.emnlp-main)

Copied to clipboard

Challenge: Low-resource language understanding is challenging for large language models (LLMs).
Approach: They propose a CompRehensive lIterary Chinese readIng comprehenSion procedure with a large dataset for CRISIS.
Outcome: The proposed procedure has the largest dataset and substantiates the effectiveness of the proposed procedure with a 7 percent hike in accuracy compared with the baseline.
Uncovering Implicit Inferences for Improved Relational Argument Mining (2023.eacl-main)

Copied to clipboard

Challenge: Argument mining attempts to extract arguments and their structure from unstructured texts.
Approach: They propose a generative neuro-symbolic approach to finding inference chains that connect argument pairs by using the Commonsense Transformer.
Outcome: The proposed approach outperforms the state-of-the-art by 2-5% in F1 score on three datasets.
Efficient Tool Use with Chain-of-Abstraction Reasoning (2025.coling-main)

Copied to clipboard

Challenge: Recent large language models have made progress at interpreting and executing instructions.
Approach: They propose a method to decouple general reasoning from specialized knowledge . they propose to use abstract reasoning chains and domain tools to reify each chain .
Outcome: The proposed method outperforms baseline methods on QA and mathematical reasoning domains.
Convolutional Interaction Network for Natural Language Inference (D18-1)

Copied to clipboard

Challenge: Attention-based neural models have achieved great success in natural language inference (NLI).
Approach: They propose a general model to capture the interaction between two sentences, which can be an alternative to the attention mechanism for NLI.
Outcome: The proposed model can capture complex interactions on three large datasets.
Automating Android Build Repair: Bridging the Reasoning-Execution Gap in LLM Agents with Domain-Specific Tools (2026.eacl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) have been shown to be useful for building applications, but their use for fixing Android build errors remains underexplored.
Approach: They propose a large-level language model agent with domain-specific tools for inspecting and manipulating the Gradle build environment.
Outcome: The proposed agent outperforms a state-of-the-art coding agent that relies on a general-purpose shell significantly on 184 build errors.
Word Surprisal Correlates with Sentential Contradiction in LLMs (2026.eacl-long)

Copied to clipboard

Challenge: Existing models are primarily optimized for task-specific performance, lacking well-defined objectives or linguistic grounding.
Approach: They propose a token-to-word decoding algorithm that extends theoretically grounded probability estimation to open-vocabulary settings.
Outcome: The proposed algorithm can localize sentence-level inconsistency at the word level, establishing a quantitative link between lexical uncertainty and sentential semantics.
Modeling Naive Psychology of Characters in Simple Commonsense Stories (P18-1)

Copied to clipboard

Challenge: Understanding a narrative requires reasoning about the causal links between the events in the story and the mental states of the characters, even when those relationships are not explicitly stated.
Approach: They propose a new annotation framework to explain naive psychology of story characters as fully-specified chains of mental states with respect to motivations and emotional reactions.
Outcome: The proposed framework provides a baseline performance on several new tasks suggesting avenues for future research.
Explicit Utilization of General Knowledge in Machine Reading Comprehension (P19-1)

Copied to clipboard

Challenge: Existing MRC models are unable to integrate general knowledge with human knowledge.
Approach: They propose a data enrichment method which uses WordNet to extract inter-word semantic connections as general knowledge from each given passage-question pair.
Outcome: The proposed model outperforms state-of-the-art models and is robust to noise.
Omni-Chart-600K: A Comprehensive Dataset of Chart Types for Chart Understanding (2025.findings-naacl)

Copied to clipboard

Challenge: Existing chart-related training methods lack capabilities in information extraction, mathematical reasoning, and understanding of multiple chart types.
Approach: They propose a two-stage training strategy and method for jointly training a vision encoder tailored for multi-type charts to address the deficiencies in chart types and limited scope of chart tasks in existing datasets.
Outcome: The proposed dataset includes 21 diverse chart types and tasks, including data retrieval and mathematical reasoning.
Enhancing Arguments Recognition for Financial Mathematical Reasoning over Hybrid Data (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing methods for question answering on textual data are difficult to train and pose a misrecognition problem.
Approach: They propose an approach to train a reasoning program generator to improve argument recognition by aggregating arguments and loss argument set.
Outcome: The proposed method improves the probabilities of proper arguments in a reasoning program generation so that arguments comprising the ground truth have higher weights.
Improving Numerical Reasoning Skills in the Modular Approach for Complex Question Answering on Text (2021.findings-emnlp)

Copied to clipboard

Challenge: Neural Module Networks (NMNs) is an end-to-end differentiable model in the programmer-interpreter paradigm.
Approach: They propose to make the interpreter question-aware and capture the relationship between entities and numbers in both questions and paragraphs.
Outcome: The proposed models outperform the original models on the DROP dataset and are interpertable by nature.
A Vietnamese Dataset for Evaluating Machine Reading Comprehension (2020.coling-main)

Copied to clipboard

Challenge: despite the lack of benchmark datasets for Vietnamese, there are few studies on machine reading comprehension (MRC) . MRC is an essential core for a range of natural language processing applications such as search engines and intelligent agents.
Approach: They propose to use Vietnamese Question Answering Dataset to evaluate machine reading comprehension in Vietnamese . they use over 23,000 human-generated question-answer pairs based on 5,109 Vietnamese articles .
Outcome: The proposed dataset includes over 23,000 human-generated question-answer pairs based on 5,109 passages of 174 Vietnamese articles from Wikipedia.
Persuasion at Play: Understanding Misinformation Dynamics in Demographic-Aware Human-LLM Interactions (2026.eacl-long)

Copied to clipboard

Challenge: Existing challenges in misinformation exposure and susceptibility vary across demographics.
Approach: They propose a framework that investigates the bidirectional persuasion dynamics between LLMs and humans when exposed to misinformation.
Outcome: The proposed framework analyzes the spread of misinformation under persuasion among demographic-oriented LLM agents.
Entropy-Gated Branching for Efficient Test-Time Reasoning (2026.eacl-long)

Copied to clipboard

Challenge: Empirical results show that branching at low uncertainty points can improve reasoning capabilities of large language models . however, these methods require substantially more computational resources, causing errors in high-stakes domains .
Approach: They propose an inference technique that selectively expands prediction sequences at points of high uncertainty.
Outcome: Empirical results show that the proposed method improves accuracy by 22.6% over standard inference while operating 31%-75% faster across math benchmarks.
DeFine: Decision-Making with Analogical Reasoning over Factor Profiles (2025.findings-acl)

Copied to clipboard

Challenge: Large language models are ideal for decision-making, but they can be difficult to process when they are verbose and include repetition, hedging, and vagueness.
Approach: They propose a framework that constructs probabilistic factor profiles from complex scenarios and integrates them with analogical reasoning to guide LLMs in making decisions in new situations.
Outcome: The proposed framework separates the tasks of quantifying uncertainty and incorporating it into LLM decision-making.
Missci: Reconstructing Fallacies in Misrepresented Science (2024.acl-long)

Copied to clipboard

Challenge: False or misleading narratives spread rapidly on social networks, posing challenges for non-experts in discerning credible information.
Approach: They propose a model for fallacious reasoning that focuses on implicit fallacies between relevant content and the inaccurate claim and requires models to verbalize the fallacious thinking in addition to classifying it.
Outcome: The proposed model focuses on implicit fallacies between relevant content and the inaccurate claim and requires models to verbalize the fallacious reasoning in addition to classifying it.
NUT-RC: Noisy User-generated Text-oriented Reading Comprehension (2020.coling-main)

Copied to clipboard

Challenge: Existing RC models focus on extractive or generative, but ignore integration of them.
Approach: They propose a noisy user-generated text-oriented RC model that integrates extractive and generative RC models by a multi-task learning mechanism and an answer selection module.
Outcome: The proposed model outperforms state-of-the-art models on Twitter.
Cosmos QA: Machine Reading Comprehension with Contextual Commonsense Reasoning (D19-1)

Copied to clipboard

Challenge: Existing reading comprehension datasets focus on factual and literal understanding of context paragraphs, but our dataset focuses on reading between the lines over a diverse collection of everyday narratives.
Approach: They propose a large-scale dataset that requires commonsense-based reading comprehension, formulated as multiple-choice questions.
Outcome: The proposed architecture improves over the baselines of existing reading comprehension datasets and shows a significant gap between machine (68.4%) and human performance (94%).
When Reasoning Meets Information Aggregation: A Case Study with Sports Narratives (2024.emnlp-main)

Copied to clipboard

Challenge: Using sports data, an LLM can analyze sports narratives to infer points from actions, identify related entities, attribute points accurately to players and teams, and draw conclusions.
Approach: They propose a method to synthesize NBA basketball game narratives using real NBA basketball data and propose 'SportsGen' they find that most models fail to accurately aggregate basketball scores due to frequent scoring patterns and open-source models suffer from significant score hallucinations.
Outcome: The proposed method can evaluate LLMs’ reasoning capabilities under complex scenarios with varying narrative lengths and density of information.
Cross-Cultural Transfer of Commonsense Reasoning in LLMs: Evidence from the Arab World (2025.findings-emnlp)

Copied to clipboard

Challenge: a recent study examined the potential for cross-cultural transfer of commonsense reasoning . merely 12 culture-specific examples from one country can improve performance in others by 10% on average .
Approach: They evaluate cross-cultural transfer of commonsense reasoning within the arab world . they use in-context learning and demonstration-based reinforcement to evaluate alignment methods .
Outcome: The proposed model can improve performance in cultures with cultural similarities in the Arab world by 10% on average.
Robust Machine Reading Comprehension by Learning Soft labels (2020.coling-main)

Copied to clipboard

Challenge: Neural models have achieved great success on the task of machine reading comprehension, which are typically trained on hard labels.
Approach: They propose a robust training method for machine reading comprehension models to address label sparseness problem by using three strategies to train models on soft labels.
Outcome: The proposed method improves the baseline model performance and achieves state-of-the-art performance on NewsQA and QUOREF.
Topological Sort for Sentence Ordering (2020.acl-main)

Copied to clipboard

Challenge: Recent work on sentence ordering task has framed it as a sequence prediction problem.
Approach: They propose a new constraint solving problem and propose 'human evaluation' they propose to capture coherence in documents by arranging sentences in the correct order .
Outcome: The proposed technique captures coherence in documents better than previous approaches.
Chapter Ordering in Novels (2022.emnlp-main)

Copied to clipboard

Challenge: a major challenge in research on long-form narrative texts is the cost of annotation . authors propose a new task that reconstructs the original order of chapters in novels without the need for human annotation.
Approach: They propose a task that reconstructs the original order of chapters in novels given a random permutation of the text.
Outcome: The proposed task yields a Spearman correlation of 0.59 on the novel and challenging task, substantially above baseline.
Discourse-Aware Semantic Self-Attention for Narrative Reading Comprehension (D19-1)

Copied to clipboard

Challenge: Existing work on discourse-aware self-attention models for reading comprehension uses annotations .
Approach: They propose to use linguistic annotations as a basis for a Discourse-Aware Semantic Self-Attention encoder for reading comprehension on narrative texts.
Outcome: The proposed model improves reading comprehension performance on narrative texts up to +3.4 Rouge-L . it also improves inter- and cross-sentential discourse relations, sentence-internal semantic role relations, and long-distance coreference relations.
Multi-hop Reading Comprehension across Multiple Documents by Reasoning over Heterogeneous Graphs (P19-1)

Copied to clipboard

Challenge: Existing models to tackle multi-hop reading comprehension (RC) are focusing on a single document or paragraph, but they lack the ability to do reasoning across multiple documents.
Approach: They propose a heterogeneous document-entity graph with different types of nodes and edges to solve multi-hop RC problem.
Outcome: The proposed model can do reasoning over the proposed graph with nodes representation initialized with co-attention and self-attention based context encoders.
CLICKER: Cross-Lingual Knowledge Editing via In-Context Learning with Adaptive Stepwise Reasoning (2026.findings-eacl)

Copied to clipboard

Challenge: Existing knowledge editing methods are static and fail to propagate edits across languages.
Approach: They propose a KE method that dynamically retrieves only knowledge relevant to a given query and edits it to maintain cross-lingual consistency.
Outcome: The proposed method outperforms static KE methods on a multilingual dataset with semantically similar but irrelevant prompts.
DiffuCOMET: Contextual Commonsense Knowledge Diffusion (2024.acl-long)

Copied to clipboard

Challenge: Recent methods for identifying contextually relevant commonsense inferences are weak . knowledge models are trained to verbalize tuples from general commonsens knowledge graphs .
Approach: They develop a series of knowledge models that leverage diffusion to reconstruct semantic connections between narrative contexts and relevant commonsense knowledge.
Outcome: The proposed model improves on two benchmarks, ComFact and WebNLG+, to measure commonsense diversity and contextual relevance.
How Do Humans Write Code? Large Models Do It the Same Way Too (2024.emnlp-main)

Copied to clipboard

Challenge: Program-of-Thought (PoT) replaces natural language-based Chain-ofThough (CoT) but introduces more reasoning errors, such as incorrect formulas or flawed logic, compared to CoT.
Approach: They propose a method that integrates CoT and Program-of-Thought to achieve more accurate reasoning and reinforcement learning.
Outcome: The proposed method achieves an average improvement of 6.5% on the Llama-Base model and 4.3% on the Mistral-Bass model across 8 mathematical calculation datasets.
Semantically-Aligned Equation Generation for Solving and Reasoning Math Word Problems (N19-1)

Copied to clipboard

Challenge: Existing methods to solve math word problems require accurate natural language understanding to bridge texts and math expressions.
Approach: They propose a neural approach to automatically solve math word problems by operating symbols according to their semantic meanings in texts.
Outcome: The proposed model outperforms state-of-the-art models and the best non-retrieval-based models over 10% accuracy in a Math23K dataset.
ObjChangeVR: Object State Change Reasoning from Continuous Egocentric Views in VR Environments (2026.eacl-long)

Copied to clipboard

Challenge: Recent advances in multimodal large language models (MLLMs) have shown remarkable capabilities in scene understanding.
Approach: They propose a framework that combines viewpoint-aware and temporal-based retrieval to identify relevant frames, along with cross-view reasoning that reconciles inconsistent evidence from multiple viewpoints.
Outcome: Extensive experiments show that the proposed framework outperforms baseline approaches across multiple MLLMs.
Tracking the Limits of Knowledge Propagation: How LLMs Fail at Multi-Step Reasoning with Conflicting Knowledge (2026.eacl-long)

Copied to clipboard

Challenge: Existing benchmarks for analyzing the performance of Large Language Models (LLMs) focus on single knowledge updates and fact recall, but do not consider how these updates affect downstream reasoning.
Approach: They propose a benchmark to study how LLMs propagate new knowledge when it conflicts with the model's parametric knowledge.
Outcome: The proposed benchmark compared models with no updated facts to show that the new methods worsen performance and improve reasoning performance.
Commonsense Justification for Action Explanation (D18-1)

Copied to clipboard

Challenge: a recent study examines the commonsense reasoning used by humans to justify an AI prediction.
Approach: They propose an approach that models object relations/attributes of the world as latent variables and jointly learns a performer that predicts actions and an explainer that gathers commonsense evidence to justify the action.
Outcome: The proposed model achieves significantly higher performance in both action prediction and justification.
GRPO-LEAD: A Difficulty-Aware Reinforcement Learning Approach for Concise Mathematical Reasoning in Language Models (2025.emnlp-main)

Copied to clipboard

Challenge: Existing methods for group-relative policy optimization face challenges in reward sparsity, verbosity and inadequate focus on problem difficulty.
Approach: They propose a method to improve group relative policy optimization with length-regularized rewards and explicit penalties for incorrect solutions.
Outcome: The proposed method achieves state-of-the-art performance for 14B-scale models . it improves reasoning accuracy, conciseness, and efficiency .
Can LLMs reason over extended multilingual contexts? Towards long-context evaluation beyond retrieval over haystacks (2026.eacl-long)

Copied to clipboard

Challenge: Existing multilingual long-context benchmarks are myopic and inherently limited, as successful recall alone does not indicate a model’s capacity to reason over extended contexts.
Approach: They propose a new synthetic benchmark for multilingual long-context reasoning that includes bAbI-style tasks that test multi-hop inference, aggregation, and epistemic reasoning.
Outcome: The proposed benchmarks are based on a multilingual long-context model and span seven languages.
Retrieval Augmentation for Commonsense Reasoning: A Unified Approach (2022.emnlp-main)

Copied to clipboard

Challenge: Existing methods for retrieving encyclopedic knowledge lack a large corpus and effective commonsense retriever.
Approach: They propose a framework for retrieval-augmented commonsense reasoning with a large commonsensense corpus and a commonseense retriever.
Outcome: The proposed framework outperforms existing methods on commonsense reasoning tasks.
HonestBait: Forward References for Attractive but Faithful Headline Generation (2023.findings-acl)

Copied to clipboard

Challenge: Current approaches to generating attractive headlines often learn directly from data based on clicks and views . clickbait models fail to reveal how much interest is raised by the writing style and how much is due to the event or topic itself .
Approach: They propose a framework for generating headlines using forward references . they use a dataset containing pairs of fake news and verified news .
Outcome: The proposed framework yields more attractive headlines while maintaining high veracity . the framework is based on a dataset containing fake news with verified news .
Identifying the Human Values behind Arguments (2022.acl-long)

Copied to clipboard

Challenge: et al., 2003) examines human values in natural language arguments . authors provide a dataset of 5270 arguments from four geographical cultures .
Approach: They propose a multi-level taxonomy of human values with 54 values and a dataset of 5270 arguments from four geographical cultures, manually annotated for human values.
Outcome: The proposed model shows that human values are more diverse than previously thought . it shows that people disagree on the best course forward on controversial issues .
CipherBank: Exploring the Boundary of LLM Reasoning Capabilities through Cryptography Challenge (2025.findings-acl)

Copied to clipboard

Challenge: Large language models (LLMs) have demonstrated remarkable capabilities, but their capabilities in cryptographic decryption tasks remain underexplored.
Approach: They propose a benchmark to evaluate the reasoning capabilities of large language models in cryptographic decryption tasks.
Outcome: The proposed benchmark examines the reasoning capabilities of large language models in cryptographic decryption tasks.
It’s All in the Heads: Using Attention Heads as a Baseline for Cross-Lingual Transfer in Commonsense Reasoning (2021.findings-acl)

Copied to clipboard

Challenge: gilbert et al.: commonsense reasoning is a key problem in natural language processing but its capabilities are still unstudied. gilland eetal.: a new approach to commonsensible reasoning is needed to solve the problem.
Approach: They propose a method which trains a linear classifier with weights of multi-head attention as features and a multilingual Winograd Schema corpus to measure cross-lingual generalization ability.
Outcome: The proposed approach performs competitively with recent approaches even when applied to other languages in a zero-shot manner.
LAFaCT: Attribution-based Localization and Focused Sequential Analysis of Fact-Critical Tokens for Hallucination Detection (2026.acl-long)

Copied to clipboard

Challenge: Large Language Models suffer from hallucinations, severely undermining their reliability.
Approach: They propose a framework that localizes fact-critical tokens and performs sequential analysis on their hidden states.
Outcome: The proposed framework localizes fact-critical tokens using Factual Criticality . it then performs a focused sequential analysis on their hidden states .
UTMath: A Benchmark for Math Evaluation with Unit Test (2025.findings-emnlp)

Copied to clipboard

Challenge: Prevailing benchmarks for mathematical reasoning include MATH and AIME . predicated on single-instantiation problems with fixed numbers, these models leave generalization on isomorphic problem variants untested.
Approach: They propose a mathematical reasoning benchmark that quantifies solution accuracy and solution space generality.
Outcome: The proposed model solves 1,053 problems spanning 9 mathematical domains . the best-performing model solved only 32.57% of the problems .
Uncovering Implicit Gender Bias in Narratives through Commonsense Inference (2021.findings-emnlp)

Copied to clipboard

Challenge: Pre-trained language models learn harmful biases from their training corpora and may repeat these biase if used for generation.
Approach: They focus on gender biases associated with the protagonist in model-generated stories and use a commonsense reasoning engine to uncover them.
Outcome: The proposed model-generated stories are based on a commonsense reasoning engine and are able to uncover gender biases in the protagonist's motivations, attributes, mental states, and implications on others.
What Makes a Good Reasoning Chain? Uncovering Structural Patterns in Long Chain-of-Thought Reasoning (2025.emnlp-main)

Copied to clipboard

Challenge: Recent advances in reasoning with large language models have popularized Long Chain-of-Thought (LCoT) a framework that converts sequential LCoTs into hierarchical tree structures enables deeper structural analysis of LLM reasoning.
Approach: They propose a framework that converts sequential LCoTs into hierarchical tree structures and enables deeper structural analysis of LLM reasoning.
Outcome: The proposed framework can be used to analyze LLM reasoning in a variety of tasks and models.
“I’m Not Mad”: Commonsense Implications of Negation and Contradiction (2021.naacl-main)

Copied to clipboard

Challenge: a new commonsense knowledge graph for negated and contradicted events is developed to help humans reason about their underlying causes and effects.
Approach: They propose a new commonsense knowledge graph with 624K if-then rules focusing on negated and contradictory events.
Outcome: The proposed model can be used to analyze negated and contradicted statements in natural language.
Self-Consistency Boosts Calibration for Math Reasoning (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing solutions for math reasoning tasks use semantic parsing or AST decoding, but performance can degrade dramatically even with slight changes to the questions.
Approach: They propose three calibration methods based on self-consistency for math reasoning tasks.
Outcome: The proposed methods bridge model confidence and accuracy better than existing methods based on p(True) or logit.
Cards Against Contamination: TCG-Bench for Difficulty-Scalable Multilingual LLM Reasoning (2026.findings-eacl)

Copied to clipboard

Challenge: Recent studies find 25-50% of evaluation datasets appear in training corpora . contamination hinders the possibility to differentiate memorization and reasoning skills.
Approach: They propose a two-player trading card game that is contaminated by a public engine and hidden card implementations to prevent benchmark saturation.
Outcome: The proposed benchmark is based on a new two-player trading card game similar to Magic: The Gathering.
RAG-Star: Enhancing Deliberative Reasoning with Retrieval Augmented Verification and Refinement (2025.naacl-long)

Copied to clipboard

Challenge: Existing large language models (LLMs) show exceptional problem-solving capabilities but struggle with complex reasoning tasks.
Approach: They propose a novel RAG approach that integrates retrieved information to guide tree-based reasoning process based on LLMs.
Outcome: The proposed approach outperforms existing methods in large language models . iteratively plans intermediate sub-queries and answers based on the LLM itself .
Search from History and Reason for Future: Two-stage Reasoning on Temporal Knowledge Graphs (2021.acl-long)

Copied to clipboard

Challenge: Temporal Knowledge Graphs (TKGs) are used in many different areas of research.
Approach: They propose to use a beam search policy to induce multiple clues from historical facts . they propose to adopt a graph convolution network based sequence method to deduce answers from clues .
Outcome: The proposed model can predict future facts in two stages, Clue Searching and Temporal Reasoning.
Connecting the Dots: A Knowledgeable Path Generator for Commonsense Question Answering (2020.findings-emnlp)

Copied to clipboard

Challenge: Existing QA systems do not have commonsense knowledge or cannot reason with it.
Approach: They propose to augment a general commonsense QA framework with a knowledgeable path generator by extrapolating existing paths from a KG with 'state-of-the-art' language model.
Outcome: The generated paths are interpretable, novel, and relevant to the task.
Taming "Zombie" Agents: A Markov State-Aware Framework for Resilient Multi-Agent Evolution (2026.acl-long)

Copied to clipboard

Challenge: Existing approaches to improve efficiency of multi-agent systems rely on aggressive graph topology evolution . however, such hard pruning overlooks the potential for "zombie" agents to recover and contribute in subsequent discussion rounds.
Approach: They propose a Markov state-aware framework for resilient multi-agent evolution that manages agent collaboration through soft state transitions.
Outcome: The proposed framework outperforms baselines and significantly reduces token consumption through state-aware agent scheduling.
An Audit on the Perspectives and Challenges of Hallucinations in NLP (2024.emnlp-main)

Copied to clipboard

Challenge: 103 peer-reviewed publications on hallucination in large language models (LLMs) are characterized by a lack of agreement with the term ‘hallucination’ in the field of NLP.
Approach: They examine 103 peer-reviewed publications on hallucination in large language models (LLMs) and conduct a survey with 171 practitioners from the field of NLP and AI to capture varying perspectives on halllucination.
Outcome: The findings highlight the need for explicit definitions and frameworks outlining hallucination within NLP and highlight potential challenges.
Verifiable, Debuggable, and Repairable Commonsense Logical Reasoning via LLM-based Theory Resolution (2024.emnlp-main)

Copied to clipboard

Challenge: Recent advances in Large Language Models (LLMs) have led to substantial interest in their application to commonsense reasoning tasks.
Approach: They propose a logical reasoning framework that integrates commonsense knowledge with a verifiable logical framework that mitigates hallucinations and facilitates debugging.
Outcome: The proposed framework improves on three language-based reasoning tasks and improves accuracy and reasoning correctness.
Commonsense Reasoning in Arab Culture (2025.acl-long)

Copied to clipboard

Challenge: Existing studies on commonsense reasoning in Arabic have relied on machine translations that lack cultural depth and introduce anglocentric biases.
Approach: They propose a commonsense reasoning dataset in Arabic that covers 13 Arab countries.
Outcome: The proposed dataset covers 13 countries across the Gulf, Levant, North Africa, and the Nile Valley.
Make Every Penny Count: Difficulty-Adaptive Self-Consistency for Cost-Efficient Reasoning (2025.findings-naacl)

Copied to clipboard

Challenge: Existing decoding strategies for chain-of-thought reasoning do not exploit prior information about question difficulty.
Approach: They propose a decoding strategy called self-consistency to improve reasoning performance by adjusting the number of samples based on the posterior distribution of a set of pre-samples.
Outcome: The proposed method outperforms baseline methods on arithmetic, commonsense and symbolic reasoning tasks while achieving comparable performance.
Identifying inherent disagreement in natural language inference (2021.naacl-main)

Copied to clipboard

Challenge: Natural language inference is the task of determining whether text is entailed, contradicted or unrelated to another piece of text.
Approach: They propose to tease systematic inferences from disagreement items by capturing modes in annotations to simulate uncertainty in the annotation process.
Outcome: The proposed approach performs statistically better than baselines on the CommitmentBank corpus in English.
From Generation to Selection: Findings of Converting Analogical Problem-Solving into Multiple-Choice Questions (2024.findings-emnlp)

Copied to clipboard

Challenge: Abstract and Reasoning Corpus (ARC) is a benchmark designed to evaluate reasoning abilities alone by reducing the amount of prior knowledge and data required to solve the tasks.
Approach: They propose a multiple-choice format suitable for assessing stages like Understand and Apply in Large Language Models (LLMs).
Outcome: The proposed model supports analogical reasoning and evidence analysis, but LLMs use shortcuts in the MC-LARC format.
PricingLogic: Evaluating LLMs Reasoning on Complex Tourism Pricing Tasks (2025.emnlp-main)

Copied to clipboard

Challenge: Recent advances in Large Language Models (LLMs) have demonstrated remarkable capabilities across diverse domains, such as code generation, mathematical problem-solving, and general-purpose human instruction following.
Approach: They propose to use large language models to process questions expressed in natural language to automate tourism-booking prices when multiple, overlapping farerules apply.
Outcome: The proposed model can automate tourism-booking prices when multiple, overlapping farerules apply.
Target Inference in Argument Conclusion Generation (2020.acl-main)

Copied to clipboard

Challenge: Existing approaches focus on generating single claims, but there are limitations.
Approach: They propose to use a triplet neural network to infer a conclusion's target from premises' targets and a neural network for a new target.
Outcome: The proposed approach outperforms baselines on two domains.
Thinking Like a Skeptic: Defeasible Inference in Natural Language (2020.findings-emnlp)

Copied to clipboard

Challenge: Defeasible inference is a mode of reasoning in which an inference may be weakened or overturned in light of new evidence.
Approach: They propose a dataset for defeasible inference in natural language that includes extensions to existing inference datasets.
Outcome: Defeasible NLI extends existing datasets for defeaasibility inference in natural language . generative models can weaken or strengthen inferences up to 68% of the time, it shows .
REAR: Reinforced Reasoning Optimization for Event Argument Extraction with Relation-Aware Support (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing methods for EAE restrict integration of relation-level semantics, thereby overlooking the complementary cues from RE.
Approach: They propose a Relation-aware EAE Reinforced optimization framework that integrates relation-level cues from RE into the Large Language Model (LLM)
Outcome: The proposed framework surpasses existing decoder-only methods on the ACE-E, ACE+ and ERE benchmarks.
Deceptive Semantic Shortcuts on Reasoning Chains: How Far Can Models Go without Hallucination? (2024.naacl-long)

Copied to clipboard

Challenge: Existing large language models (LLMs) suffer from hallucinations and unfaithful reasoning due to keyword/entity biases.
Approach: They propose a new probing method and benchmark to quantify this phenomenon by using a keyword/entity biases-based probing technique called EUREQA.
Outcome: The proposed method achieves 62% accuracy on multi-hop and complex QA benchmarks.
On A Scale From 1 to 5: Quantifying Hallucination in Faithfulness Evaluation (2025.findings-naacl)

Copied to clipboard

Challenge: Hallucination is a popular topic in natural language generation (NLG).
Approach: They propose to use large language models to evaluate faithfulness of guided NLGs by a rubric template and large language inference models to score the generation on quantifiable scales.
Outcome: The proposed system can provide accurate judgement and explain whether a source and generation are factually consistent.
Penetrative AI: Making LLMs Comprehend the Physical World (2024.findings-acl)

Copied to clipboard

Challenge: Recent advances in Large Language Models (LLMs) have demonstrated remarkable capabilities across a range of tasks.
Approach: They explore how LLMs can be extended to interact with and reason about the physical world through IoT sensors and actuators, a concept that they call "Penetrative AI".
Outcome: The proposed approach extends LLMs' capabilities to interact with and reason about the physical world through IoT sensors and actuators.
Telling a Lie: Analyzing the Language of Information and Misinformation during Global Health Events (2022.lrec-1)

Copied to clipboard

Challenge: a new dataset is available to stimulate research on health misinformation . linguistic characteristics of health misinfonia are unique to COVID-19 and other events .
Approach: They propose a new dataset that analyzes health misinformation at scale . it includes 2.8 million news articles and social media posts covering diseases . authors propose an annotation framework that allows for strong agreement between annotators .
Outcome: The proposed dataset is based on 2.8 million news articles and social media posts spanning 1900s to present . it shows that the proposed model is robust and can be used to detect misinformation .
Validity Assessment of Legal Will Statements as Natural Language Inference (2022.findings-emnlp)

Copied to clipboard

Challenge: This study introduces a dataset that focuses on the validity of statements in legal wills.
Approach: They propose a dataset that focuses on the validity of statements in legal wills.
Outcome: The proposed model achieves 80% macro F1 and accuracy, but group accuracy is in mid 80s at best, suggesting that the models’ understanding of the task remains superficial.
BizBench: A Quantitative Reasoning Benchmark for Business and Finance (2024.acl-long)

Copied to clipboard

Challenge: Answering questions within business and finance requires reasoning, precision, and a wide-breadth of technical knowledge.
Approach: They propose a benchmark for evaluating models’ ability to reason about realistic financial problems by focusing on question-answering over financial data via program synthesis.
Outcome: The proposed benchmark evaluates models' financial background knowledge, ability to parse financial documents, and capacity to solve complex problems with code.
WiCE: Real-World Entailment for Claims in Wikipedia (2023.emnlp-main)

Copied to clipboard

Challenge: Textual entailment models are increasingly used in fact-checking, presupposition verification in question answering, or summary evaluation.
Approach: They propose a new fine-grained textual entailment dataset built on natural claim and evidence pairs extracted from Wikipedia that provides en-tailment judgments over sub-sentence units of the claim and a minimal subset of evidence sentences that support each subclaim.
Outcome: The proposed dataset improves on multiple datasets at test time and shows that real claims involve verification and retrieval problems that existing models fail to address.
MultiFC: A Real-World Multi-Domain Dataset for Evidence-Based Fact Checking of Claims (D19-1)

Copied to clipboard

Challenge: Existing efforts to verify factual claims are limited by small datasets or artificially constructed datasets.
Approach: They propose to use the largest publicly available dataset of naturally occurring factual claims for automatic claim verification.
Outcome: The proposed model outperforms baseline models and evidence pages significantly.
On LLM-Based Scientific Inductive Reasoning Beyond Equations (2025.emnlp-main)

Copied to clipboard

Challenge: Existing research on inductive reasoning models emphasizes rule design without grounding them in specific scenarios.
Approach: They propose to use LLMs to learn underlying patterns from limited examples in entirely new environments.
Outcome: The proposed benchmark evaluates the inductive reasoning abilities of large language models in scientific settings.
Masked by Consensus: Disentangling Privileged Knowledge in LLM Correctness (2026.acl-long)

Copied to clipboard

Challenge: Recent research suggests large language models encode meta-information about their own outputs.
Approach: They investigate whether large language models possess similar privileged knowledge about answer correctness . they train correctness classifiers on question representations from a model’s hidden states and external models .
Outcome: The proposed model outperforms peer-model models in factual knowledge tasks, but shows no advantage in math reasoning.
Self-Evolving Multi-Agent Systems via Textual Backpropagation (2026.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) have proven effective for addressing complex, high-dimensional tasks, but current approaches rely on static, manually engineered multi-agent configurations.
Approach: They propose a framework that conceptualizes multi-agent collaboration as a layered neural network architecture.
Outcome: The proposed framework surpasses leading multi-agent baselines under the same configurations, showing consistent performance improvements.
Benefits of Intermediate Annotations in Reading Comprehension (2020.acl-main)

Copied to clipboard

Challenge: a lack of intermediate supervision makes learning difficult for complex compositional reading comprehension datasets . lack of supervision makes the learning difficult, leading to the lack of correct latent reasoning steps.
Approach: They propose to collect intermediate reasoning supervision along with the final answer during data collection . they find that this helps combat some of the biases introduced during the data collection process .
Outcome: The use of intermediate reasoning supervision improves model performance for complex compositional datasets.
Learning from Contrasts: Synthesizing Reasoning Paths from Diverse Search Trajectories (2026.acl-long)

Copied to clipboard

Challenge: MCTS methods retain only the single highest-reward trajectory, discarding comparative signals present in the many explored paths.
Approach: They propose a framework that transforms supervision extraction into a synthesis procedure.
Outcome: The proposed framework matches or exceeds baselines on 60K CRPS-synthesized examples on out-of-domain benchmarks.
Reasoning with Language Model is Planning with World Model (2023.emnlp-main)

Copied to clipboard

Challenge: Large language models (LLMs) have shown remarkable reasoning capabilities, particularly with Chain-of-Thought-style prompts.
Approach: They propose a framework that repurposes the LLM as both a world model and a reasoning agent and incorporates a principled planning algorithm (based on Monte Carlo Tree Search)
Outcome: The proposed framework repurposes the LLM as both a world model and a reasoning agent and incorporates a principled planning algorithm (based on Monte Carlo Tree Search) it achieves optimum balance between exploration and exploitation, while achieving high-reward reasoning paths efficiently.
DiVERT: Distractor Generation with Variational Errors Represented as Text for Math Multiple-choice Questions (2024.emnlp-main)

Copied to clipboard

Challenge: a new variational approach to distractors in multiple-choice questions is needed . high-quality distractors are crucial to the assessment and pedagogical value of MCQs . a variational method that learns the error behind distractors is more effective .
Approach: They propose a variational approach that learns an interpretable representation of errors behind distractors in math MCQs.
Outcome: The proposed method outperforms state-of-the-art approaches on distractors in math MCQs.
Ensembling Large Language Models with Process Reward-Guided Tree Search for Better Complex Reasoning (2025.naacl-long)

Copied to clipboard

Challenge: Existing methods for ensembling language models fail to address complex reasoning tasks.
Approach: They propose a framework for process-level ensembling of large language models using Monte Carlo tree search.
Outcome: The proposed framework outperforms both language model decoding and language model ensemble methods on five reasoning benchmarks.
Document-Level Event Argument Extraction With a Chain Reasoning Paradigm (2023.acl-long)

Copied to clipboard

Challenge: Document-level event argument extraction aims to identify event arguments beyond sentence level, where a significant challenge is to model long-range dependencies.
Approach: They propose a chain reasoning paradigm which captures long-range interdependence due to the chains’ compositional nature and generates decomposable first-order logic rules for reasoning.
Outcome: The proposed method outperforms previous methods on two benchmarks and is robust enough to defend against adversarial attacks.
ARise: Towards Knowledge-Augmented Reasoning via Risk-Adaptive Search (2025.acl-long)

Copied to clipboard

Challenge: Large language models (LLMs) have impressive capabilities but their application in open-ended, knowledge-intensive, complex reasoning scenarios is limited.
Approach: They propose a framework that integrates risk assessment of intermediate reasoning states with dynamic retrieval-augmented generation within a Monte Carlo tree search paradigm.
Outcome: The proposed framework outperforms the state-of-the-art KAR methods by up to 23.10% and the latest RAG-equipped large reasoning models by upto 25.37%.
A Survey of LLM-based Agents in Medicine: How far are we from Baymax? (2025.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) are transforming healthcare through their ability to understand and assist with medical tasks.
Approach: They analyze system profiles, clinical planning, medical reasoning frameworks, and external capacity enhancement.
Outcome: The findings highlight the future directions in medical reasoning, physical system integration, and training simulations.
Accelerating LLM Reasoning via Early Rejection with Partial Reward Modeling (2025.findings-emnlp)

Copied to clipboard

Challenge: Large Language Models (LLMs) are increasingly relied upon for solving complex reasoning tasks.
Approach: They propose to use Process Reward Models to scale inference time compute by generating in parallel . they propose to provide early signals that enable the rejection of suboptimal candidates before full generation of step is complete.
Outcome: The proposed method achieves 1.4 – 9 reduction in inference FLOPs without degrading final performance.
MCScript: A Novel Dataset for Assessing Machine Comprehension Using Script Knowledge (L18-1)

Copied to clipboard

Challenge: Various approaches for script knowledge extraction and processing have been proposed in recent years.
Approach: They propose a dataset to evaluate natural language understanding approaches based on commonsense knowledge.
Outcome: The proposed dataset provides test cases for the broader natural language understanding community.
World Models for Math Story Problems (2023.findings-acl)

Copied to clipboard

Challenge: Recent efforts to solve math story problems have lacked accurate representations of mathematical concepts.
Approach: They propose a graph-based semantic formalism for solving math story problems . they combine existing datasets and annotate a corpus of 1,019 problems with MathWorld .
Outcome: The proposed model can be used to solve math story problems with pre-trained language models . the model can also be used for generating new problems by using the model as a design space .
ACENet: Attention Guided Commonsense Reasoning on Hybrid Knowledge Graph (2022.emnlp-main)

Copied to clipboard

Challenge: Existing approaches estimate plausibility of candidate choices separately based on their respective KGs, without considering the interference among different choices.
Approach: They propose an Attention guided Commonsense rEasoning Network to integrate hybrid knowledge into the neural network.
Outcome: The proposed model outperforms existing methods on CommonsenseQA and OpenbookQA datasets and shows significant performance gains.
RLMEval: Evaluating Research-Level Neural Theorem Proving (2025.findings-emnlp)

Copied to clipboard

Challenge: RLMEval evaluates large language models for research-level neural theorem proving and proof autoformalization . the best model achieves only a 10.3% pass rate on existing benchmarks .
Approach: They propose a new evaluation suite for large language models . it evaluates research-level theorems from real-world Lean formalization projects .
Outcome: RLMEval evaluates research-level theorems from real-world Lean formalization projects.
Safe: Enhancing Mathematical Reasoning in Large Language Models via Retrospective Step-aware Formal Verification (2025.acl-long)

Copied to clipboard

Challenge: Chain-of-Thought prompting is a de facto method to elicit reasoning capabilities from large language models (LLMs).
Approach: They propose a step-aware formal verification framework Safe to address hallucinations in CoT prompting . they propose 'formal step' as a benchmark for step correctness theorem proving with 30,809 formal statements.
Outcome: The proposed framework shows significant performance improvement while offering interpretable and verifiable evidence.
Lite Unified Modeling for Discriminative Reading Comprehension (2022.acl-long)

Copied to clipboard

Challenge: generative and discriminative MRCs focus on answer generation, extractive MRC on answer extraction.
Approach: They propose a lightweight POS-Enhanced Iterative Co-Attention Network to handle diverse discriminative MRC tasks synchronously.
Outcome: The proposed model improves on four discriminative MRC benchmarks.
CriticSearch: Fine-Grained Credit Assignment for Search Agents via a Retrospective Critic (2026.findings-acl)

Copied to clipboard

Challenge: Existing search agent pipelines rely on sparse outcome rewards, leading to inefficient exploration and unstable training.
Approach: They propose a tool-integrated reasoning framework that provides turn-level feedback via a retrospective critic mechanism.
Outcome: The proposed framework outperforms baselines in multi-hop reasoning benchmarks and achieves faster convergence and training stability.
Voting or Consensus? Decision-Making in Multi-Agent Debate (2025.findings-acl)

Copied to clipboard

Challenge: Increasing the number of agents improves performance, while more discussion rounds before voting reduces it.
Approach: They propose two new methods to improve multi-agent debates by increasing agent diversity and reducing discussion rounds before voting.
Outcome: The proposed methods improve task performance by up to 3.3% with AAD and up to 7.4% with CI.
Combining Knowledge Hunting and Neural Language Models to Solve the Winograd Schema Challenge (P19-1)

Copied to clipboard

Challenge: Existing methods to solve Winograd Schema Challenge use only knowledge embedded in text . this limits the performance of such models on the WSC problems.
Approach: They propose to augment existing language models with a commonsense knowledge hunting module and an explicit reasoning module to extract the needed knowledge from text.
Outcome: The proposed system improves on the language model based methods by 5.53% and 7.7% on the dataset.
Beyond Surface Simplicity: Revealing Hidden Reasoning Attributes for Precise Commonsense Diagnosis (2025.acl-long)

Copied to clipboard

Challenge: Existing commonsense question answering benchmarks often treat these aspects in isolation, resulting in evaluation accuracy differences of up to 24.8% across different difficulty levels.
Approach: They propose a framework that reveals hidden reasoning attributes behind commonsense questions by leveraging the knowledge generated during the reasoning process.
Outcome: The proposed framework reveals hidden reasoning attributes behind commonsense questions by leveraging the knowledge generated during the reasoning process.
Towards Advanced Mathematical Reasoning for LLMs via First-Order Logic Theorem Proving (2025.emnlp-main)

Copied to clipboard

Challenge: Large language models (LLMs) have shown promising first-order logic (FOL) reasoning capabilities with applications in various areas, but their effectiveness in complex mathematical reasoning involving multi-step FOL deductions remains under-explored.
Approach: They propose a self-adaptive solution that enhances the Diversity and REAsonability of LLMs’ generation strategies by introducing an Axiom-Driven Strategy Diversification mechanism and a Sub-Proposition Error Feedback to help LLM reflect on and correct their proofs.
Outcome: The proposed model improves diversity and REAsonability of LLMs’ generation strategies by introducing an Axiom-Driven Strategy Diversification mechanism and a Sub-Proposition Error Feedback to help LLM reflect on and correct proofs.
ToolBeHonest: A Multi-level Hallucination Diagnostic Benchmark for Tool-Augmented Large Language Models (2024.emnlp-main)

Copied to clipboard

Challenge: Currently, tool-augmented large language models (LLMs) only achieve total scores of 45.3 and 37.0, respectively, on a scale of 100.
Approach: They propose a multi-level diagnostic process to assess the LLM's hallucinations through two perspectives: depth and breadth.
Outcome: The proposed diagnostic process assesses the hallucinations of large language models through two perspectives: depth and breadth.
A Few Bad Apples Spoil the Bunch: Preventing Global Entropy Collapse Driven by a Small Set of Tokens in LLM Reasoning (2026.findings-acl)

Copied to clipboard

Challenge: Reinforcement learning with verifiable rewards and Reinforced Learning from internal feedback fail to benefit from test-time compute due to entropy collapse and the resulting loss of reasoning diversity.
Approach: They propose a strategy that assigns each generated token a redistribution score and applies selective KL regularization to only the top 5% of tokens under this score.
Outcome: The proposed model improves on both RLVR and RLIF models on math reasoning benchmarks, showing that targeted entropy control at a vanishingly small subset of tokens is sufficient to sustain reasoning diversity and effective test-time scaling.
New Protocols and Negative Results for Textual Entailment Data Collection (2020.emnlp-main)

Copied to clipboard

Challenge: Natural language inference data has proven useful in benchmarking and as pretraining data for tasks requiring language understanding.
Approach: They propose four alternative protocols to improve annotation quality and diversity . they use 8.5k-example training sets to compare different protocols .
Outcome: The proposed protocols improve the ease of training and quality of the examples.
Incorporating Zoning Information into Argument Mining from Biomedical Literature (2022.lrec-1)

Copied to clipboard

Challenge: Argumentative zoning is a text zonation scheme that is used to segment text into zones that serve distinct functions.
Approach: They propose to use zoning information to incorporate into argument mining tasks . they add zonation labels predicted by an off-the-shelf model to the beginning of each sentence .
Outcome: The proposed models improve argument mining models without additional annotation cost.
Bridging the Memorization-Utilization Gap: Near-Lossless Context Compression via Reinforcement Learning (2026.acl-long)

Copied to clipboard

Challenge: Recent advances in context compression have failed to effectively utilize compressed representations for downstream tasks.
Approach: They propose a holistic training paradigm that uses outcome-based RL to enable implicit expansion.
Outcome: The proposed model outperforms previous models on NIAH, LongBench and multi-hop reasoning.
Compute Optimal Scaling of Skills: Knowledge vs Reasoning (2025.findings-acl)

Copied to clipboard

Challenge: Scaling laws are a critical component of the LLM development pipeline, but little is known about whether the COs of individual skills such as mathematical reasoning, question answering (QA) or coding, align with these APEs.
Approach: They examine knowledge-based QA and code generation to find out whether skill-dependent scaling is an artefact of the pretraining datamix.
Outcome: The proposed scaling laws are skill-dependent, and knowledge and code exhibit fundamental differences in scaling behaviour when corrected for datamix differences.
DiscoSense: Commonsense Reasoning with Discourse Connectives (2022.emnlp-main)

Copied to clipboard

Challenge: DiscoSense is a benchmark for commonsense reasoning using a wide variety of discourse connectives.
Approach: They propose a benchmark for commonsense reasoning by understanding a wide variety of discourse connectives.
Outcome: The proposed benchmark outperforms existing benchmarks on commonsense reasoning tasks.
Neutralizing Bias in LLM Reasoning using Entailment Graphs (2025.findings-acl)

Copied to clipboard

Challenge: Natural Language Inference (NLI) is a foundational understanding task in language understanding.
Approach: They propose a framework to construct counterfactual reasoning data and fine-tune LLMs to reduce attestation bias.
Outcome: The proposed framework reduces hallucinations from attestation bias on original and bias-neutralized datasets while keeping hypotheses unchanged.
Learning by Analogy: Diverse Questions Generation in Math Word Problem (2023.findings-acl)

Copied to clipboard

Challenge: Existing methods for solving math word problem (MWP) use shortcut learning to train solvers based on samples with a single question.
Approach: They propose to generate diverse yet consistent questions from a common scenario . they then feed the equations to a question generator to obtain the diverse questions . their method leads to performance improvement on the current benchmark Math23K .
Outcome: The proposed method generates diverse yet consistent questions with a variety of equations and questions . it improves on the current benchmark, which is based on the proposed method .
JECC: Commonsense Reasoning Tasks Derived from Interactive Fictions (2023.findings-acl)

Copied to clipboard

Challenge: Existing benchmarks focus on a single reasoning type and ask human annotators to write candidate statements related to the particular type of commonsense.
Approach: They propose a new commonsense reasoning dataset based on human’s Interactive Fiction (IF) gameplaywalkthroughs.
Outcome: The proposed dataset is challenging to previous machine reading models and large language models with a significant 20%performance gap compared to human experts.
MatCha: Enhancing Visual Language Pretraining with Math Reasoning and Chart Derendering (2023.acl-long)

Copied to clipboard

Challenge: Visual language models that are pretraining on natural images or image-text pairs crawled from the web perform poorly on visual language tasks such as ChartQA and ChartQA.
Approach: They propose to perform several pretraining tasks that cover plot deconstruction and numerical reasoning which are key capabilities in visual language modeling.
Outcome: The proposed model outperforms state-of-the-art methods on benchmarks such as PlotQA and ChartQA by as much as 20%.
Revisiting Chain-of-Thought Prompting: Zero-shot Can Be Stronger than Few-shot (2025.findings-emnlp)

Copied to clipboard

Challenge: In-Context Learning (ICL) is an essential emergent ability of Large Language Models (LLMs).
Approach: They introduce CoT to exemplars of ICL to enhance the reasoning capability . however, it remains unclear whether CoT exemplar is still beneficial for recent, stronger models in such tasks.
Outcome: The enhanced exemplars fail to improve the model’s reasoning performance, despite being constructed using answers from advanced models such as Qwen2.5-Max and DeepSeek-R1.
CAT: A Contextualized Conceptualization and Instantiation Framework for Commonsense Reasoning (2023.acl-long)

Copied to clipboard

Challenge: HKUST-KnowComp proposes a framework for commonsense reasoning that can be used to conceptualize commonsence knowledge bases at scale.
Approach: They propose a framework that integrates event conceptualization and instantiation to conceptualize commonsense knowledge bases at scale.
Outcome: The proposed framework achieves state-of-the-art on two conceptualization tasks and the acquired abstract commonsense knowledge significantly improves commonsence inference modeling.
What Can We Learn from Collective Human Opinions on Natural Language Inference Data? (2020.emnlp-main)

Copied to clipboard

Challenge: Despite the subjective nature of many NLU evaluations, little attention has been paid to the distribution of human opinions.
Approach: They use a dataset with 464,500 annotations to study Collective HumAn OpinionS . they argue that models lack the ability to recover the distribution over human labels .
Outcome: The proposed dataset examines the distribution of human opinions in NLU evaluation datasets.
Better Process Supervision with Bi-directional Rewarding Signals (2025.findings-acl)

Copied to clipboard

Challenge: Existing processes that reward for each step are one-directional and lack a mechanism to model the distance to the final target.
Approach: They propose a process supervision model that evaluates the correctness of previous steps and the probability of future success.
Outcome: The proposed model outperforms existing supervision models like ORM and PRM on reasoning tasks and improves solution re-design.
Towards Medical Complex Reasoning with LLMs through Medical Verifiable Problems (2025.findings-acl)

Copied to clipboard

Challenge: OpenAI o1 has been a significant milestone in large language model development . however, most research in reasoning has focused on mathematical tasks . medical domains require robust reasoning to provide reliable answers .
Approach: They propose a method to verify medical reasoning using a medical verifier . they also propose RL and reinforcement learning to enhance reasoning .
Outcome: The proposed method outperforms general and medical-specific baselines using only 40K verifiable problems.
A Real-Time System for Credibility on Twitter (2020.lrec-1)

Copied to clipboard

Challenge: Using neural networks, we can analyze Twitter in real-time to determine whether users are credible and false.
Approach: They propose to analyze Twitter in real-time using neural networks to determine credibility of tweets and users who posted them.
Outcome: The proposed method analyzes Twitter in real-time to determine which users are credible and which are not, what is false or what is true on the Internet.
Eyes Show the Way: Modelling Gaze Behaviour for Hallucination Detection (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing methods for hallucination detection depend on knowledge sources that are explicit such as Wikipedia or knowledge graphs.
Approach: They propose a cognitive approach that leverages gaze signals from humans to detect hallucinations in natural language processing (NLP) they collect and introduce an eye tracking corpus consisting of 500 instances, annotated by five annotators for hallucinism detection.
Outcome: The proposed approach achieves a balanced accuracy of 87.1% on a FactCC dataset.
ArgAnalysis35K : A large-scale dataset for Argument Quality Analysis (2023.acl-long)

Copied to clipboard

Challenge: Existing datasets in argument quality detection lack quality, quantity and diversity of topics and arguments.
Approach: They propose a dataset that adds a detailed explanation of why the argument made is true, applicable or impactful.
Outcome: The proposed dataset covers 34,890 high-quality argument-analysis pairs and is the largest of its kind to our knowledge.
Leveraging LLM Reasoning Enhances Personalized Recommender Systems (2024.findings-acl)

Copied to clipboard

Challenge: Recent advances have showcased the potential of Large Language Models (LLMs) in executing reasoning tasks, particularly facilitated by Chain-of-Thought (CoT) prompting.
Approach: They propose to use Large Language Models to perform tasks with subjectivity and personalized preferences as inputs to RecSys.
Outcome: The proposed framework aligns with real human judgment on the coherence and faithfulness of LLM reasoning responses.
Pierce the Mists, Greet the Sky: Decipher Knowledge Overshadowing via Knowledge Circuit Analysis (2025.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) are hampered by hallucinations, a particularly challenging variant, knowledge overshadowing, which can lead to erroneous outputs even with high-quality training data.
Approach: They propose a framework to analyze and detect knowledge overshadowing by using knowledge circuit analysis to dissect the function of key components in the circuit and how attention pattern dynamics contribute to the phenomenon.
Outcome: Extensive experiments show that the framework can detect and analyze knowledge overshadowing and improves on existing models.
Few-shot Dialogue Strategy Learning for Motivational Interviewing via Inductive Reasoning (2024.findings-acl)

Copied to clipboard

Challenge: Motivational Interviewing (MI) requires a system that can infer how to motivate users to adopt positive lifestyle changes.
Approach: They propose a framework that can learn and apply conversation strategies from expert demonstrations by using natural language inductive rules.
Outcome: The proposed framework outperforms in-context demonstrations that are over 50 times longer and can learn natural language strategies from demonstrations.
Dropping Experts, Recombining Neurons: Retraining-Free Pruning for Sparse Mixture-of-Experts LLMs (2025.findings-emnlp)

Copied to clipboard

Challenge: Sparse Mixture-of-Experts (SMoE) architectures require loading all expert parameters . previous work focused on expert pruning and merging but focused on neuron-level structure .
Approach: They propose a task-agnostic framework for expert pruning and reconstruction . it prunes redundant experts using router statistics, then decomposes them into neuron-level expert segments .
Outcome: The proposed framework reduces the number of experts and memory usage, making it easier to deploy.
Enhanced Data Synthesis for LLM through Reasoning Structures Generated by Hierarchical GFlowNet (2025.findings-acl)

Copied to clipboard

Challenge: Existing methods to optimize instruction-response pairs lack a systematic design for the underlying reasoning structure.
Approach: They propose a Reasoning Structure driven data Synthesis method that leverages a coarse-to-fine directed acyclic graph to construct reasoning structures efficiently.
Outcome: The proposed method outperforms existing methods in 48.50%, 84.00%, 79.90% of the synthetic datasets trained on the proposed model.
TRN-R1-Zero: Text-rich Network Reasoning via LLMs with Reinforcement Learning Only (2026.acl-long)

Copied to clipboard

Challenge: Recent large language model-based approaches often overlook graph context or depend on distillation from larger models, limiting generalisation.
Approach: They propose a framework for zero-shot reasoning on text-rich networks . they use a Neighbour-aware Group Relative Policy Optimisation objective .
Outcome: The proposed framework optimises base LLMs using a Neighbour-aware group relative policy optimisation objective based on a novel margin gain metric for the informativeness of neighbouring signals .
HalluGuard: Evidence-Grounded Small Reasoning Models to Mitigate Hallucinations in Retrieval-Augmented Generation (2026.findings-acl)

Copied to clipboard

Challenge: Large Language Models excel at NLP tasks but remain prone to hallucinations . small language models can achieve competitive results in specific tasks .
Approach: They propose a 4B-parameter Small Reasoning Model (SRM) that can be used to classify document-claim pairs as grounded or hallucinated in closed-book, document-grounded settings.
Outcome: The proposed model achieves 84.4% balanced accuracy on the RAGTruth subset of the LLM-AggreFact benchmark, surpassing specialized models, MiniCheck (7B; 84.0%) and Granite Guardian 3.3 (82.2%) Across the benchmark, it reaches 77.1% BAcc, surpasses larger general-purpose LLMs such as GPT-4o (75.9%).
Calibration-Aware Policy Optimization for Reasoning LLMs (2026.acl-long)

Copied to clipboard

Challenge: Existing approaches to model calibration are limited or sacrifice gains in reasoning accuracy.
Approach: They propose a method that improves calibration by 15% while boosting accuracy by 5% . they propose GRPO-style algorithms that misalign uncertainty-agnostic advantage estimation .
Outcome: The proposed approach improves calibration by 15% while achieving comparable to or better than GRPO on multiple mathematical reasoning benchmarks.
HalluMeasure: Fine-grained Hallucination Measurement Using Chain-of-Thought Reasoning (2024.emnlp-main)

Copied to clipboard

Challenge: HalluMeasure is a new LLM-based hallucination detection mechanism that decomposes an LLM response into atomic claims and evaluates each claim against the provided reference context.
Approach: They propose a new LLM-based hallucination detection mechanism that decomposes an LLM response into atomic claims and evaluates each atomic claim against the provided reference context.
Outcome: The proposed model can detect 3 major categories of hallucinations and 10 more specific subtypes which help to identify reasons behind the hallucinian errors.
Draft Model Knows When to Stop: Self-Verification Speculative Decoding for Long-Form Generation (2025.emnlp-main)

Copied to clipboard

Challenge: Conventional speculative decoding methods use a predefined length policy for proposing drafts, but the reality deviates from this assumption.
Approach: They propose a self-verification length policy that adaptively determines the lengths of draft sequences by referring to the draft entropy.
Outcome: The proposed method achieves 17% speedup on MT-Bench and 22% speedup in long-form reasoning.
ThinkEdit: Interpretable Weight Editing to Mitigate Overly Short Thinking in Reasoning Models (2025.emnlp-main)

Copied to clipboard

Challenge: Recent studies have shown that Large Language Models (LLMs) augmented with chain-of-thought (CoT) reasoning demonstrate impressive problem-solving abilities.
Approach: They propose a weight-editing approach to reduce overly short reasoning by steering the model along a linear direction in the representation space.
Outcome: The proposed model reduces overly short reasoning and yields significant accuracy gains on multiple math benchmarks.
When to Trust Tools? Adaptive Tool Trust Calibration For Tool-Integrated Math Reasoning (2026.findings-acl)

Copied to clipboard

Challenge: Large reasoning models have performance enhancements but still suffer from shortcomings due to limitations of the underlying language models.
Approach: They propose a framework that allows the model to choose when to trust or ignore the tool results based on the confidence score of generated code blocks.
Outcome: The proposed framework reduces the "Tool Ignored" issue by 4.1% to 7.5%.
Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization (2026.findings-acl)

Copied to clipboard

Challenge: Existing RLVR algorithms suffer from entropy collapse, leading to premature determinism and unstable optimization.
Approach: They propose an adaptive entropy flow balancing mechanism that rescales entropic-increasing and enotro-decreazing updates according to their contributions to enthroy change.
Outcome: The proposed method outperforms existing RLVR algorithms on six reasoning benchmarks.
RFS-Guard: Detecting Reasoning Hallucinations via Cross-Phase Routing Focus in Large Reasoning Models (2026.acl-long)

Copied to clipboard

Challenge: Large reasoning models (LRMs) generate intermediate reasoning traces before the final answer, yet they remain vulnerable to reasoning hallucinations such as subtle arithmetic errors.
Approach: They propose a Routing Focus Score (RFS) that measures how strongly cross-step attention routing aligns with semantic proximity derived from hidden-state cosine similarity.
Outcome: The proposed framework detects and localizes hallucinations without external tools or repeated sampling.
IS-CoT: Breaking the Long-form Generation Collapse via Interleaved Structural Thinking (2026.acl-long)

Copied to clipboard

Challenge: Existing models with reasoning capabilities suffer from a severe length collapse in open-ended writing .
Approach: They propose a framework that embeds a dynamic plan-write-reflect cycle into the generation process and train a model with interleaved reasoning traces.
Outcome: The proposed framework achieves state-of-the-art performance on long-form benchmarks compared to other models on the same dataset.
It Ain’t Over: A Multi-aspect Diverse Math Word Problem Dataset (2023.emnlp-main)

Copied to clipboard

Challenge: Existing studies lack diversity in problem types, lexical usage patterns, languages, and intermediate solution forms for the math word problem.
Approach: They propose a new MWP dataset with a wide range of diversity in problem types, lexical usage patterns, languages, and intermediate solutions.
Outcome: The proposed dataset provides an opportunity to evaluate the capability of large language models.
PhysicsArena: The First Multimodal Physics Reasoning Benchmark Exploring Variable, Process, and Solution Dimensions (2025.findings-emnlp)

Copied to clipboard

Challenge: Current physics benchmarks focus on text-only inputs or only on problem-solving . current physics reasoning benchmarks neglect critical intermediate steps of variable identification and process formulation.
Approach: a new benchmark evaluates multimodal large language models in physics reasoning . the benchmark measures variables, process formulations, and solution derivation .
Outcome: PhysicsArena is the first multimodal physics reasoning benchmark . it evaluates MLLMs across three critical dimensions: variable identification, process formulation, and solution derivation.
OMS: On-the-fly, Multi-Objective, Self-Reflective Ad Keyword Generation via LLM Agent (2025.emnlp-main)

Copied to clipboard

Challenge: Keyword decision in Sponsored Search Advertising is critical to the success of ad campaigns.
Approach: They propose a keyword generation framework that is On-the-fly and Multi-objective to automate keyword generation.
Outcome: Experiments show that OMS outperforms existing methods in keyword generation . relying on large-scale query-keyword data is a major limitation, authors say .
VIVA+: Human-Centered Situational Decision-Making (2025.findings-emnlp)

Copied to clipboard

Challenge: Multimodal Large Language Models (MLLMs) show promising results in complex, human-centered environments, yet evaluating their capacity for nuanced, humanlike reasoning and decision-making remains challenging.
Approach: They introduce VIVA+, a cognitively grounded benchmark for evaluating the reasoning and decision-making of MLLMs in human-centered situations.
Outcome: The VIVA+ model is based on 1,317 real-world situations paired with 6,373 multiple-choice questions . it consists of three core abilities for decision-making: (1) Foundational Situation Comprehension, (2) Context-Driven Action Justification, and (3) Reflective Reasoning.
Representation-based Broad Hallucination Detectors Fail to Generalize Out of Distribution (2025.findings-emnlp)

Copied to clipboard

Challenge: a spurious correlation between hallucination detection methods and data is limiting the current SOTA.
Approach: They propose a set of guidelines for hallucination detection and its evaluation.
Outcome: The proposed method performs no better than supervised linear probes on the RAGTruth dataset .
CU-MAM: Coherence-Driven Unified Macro-Structures for Argument Mining (2025.acl-long)

Copied to clipboard

Challenge: Argument Mining (AM) involves the automatic identification of argument structure in natural language.
Approach: They propose an approach that captures local and global coherence to identify argument structures by modeling macro-structure.
Outcome: The proposed approach shows superior performance on heterogeneous datasets and on unseen datasets.
V-ALPHASOCIAL: Benchmark and Self-Reflective Chain-of-Thought Generation for Visual Social Commonsense Reasoning (2025.findings-acl)

Copied to clipboard

Challenge: Social commonsense reasoning is a multimodal task that requires both textual and visual cues.
Approach: They propose a method that integrates visual cues into social commonsense reasoning tasks.
Outcome: The proposed method improves social commonsense reasoning on a multimodal foundation model.
X-Router: Decoupling Knowledge and Reasoning for Cost-Effective LLM Inference (2026.findings-acl)

Copied to clipboard

Challenge: Existing adaptive methods focus on a single axis, overlooking evidence need and reasoning depth are only partially correlated.
Approach: They propose a dual-axis routing framework that separates retrieval necessity from reasoning necessity under a user-defined cost–quality trade-off.
Outcome: The proposed framework reduces token usage and latency while improving answer quality over strong baselines.
Tree of Problems: Improving structured problem solving with compositionality (2024.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) have demonstrated remarkable performance across multipletasks through in-context learning.
Approach: They propose a Tree of Problems (ToP) that is a simpler version of Tree of Thoughts (toT) they propose 'in-context learning' is the ability of Large Language Models (LLMs) to perform a task with the help of a few demonstrations within their context.
Outcome: The proposed approach outperforms ToT and GoT and performs better on complex reasoning tasks.
On the Universal Truthfulness Hyperplane Inside LLMs (2024.emnlp-main)

Copied to clipboard

Challenge: Recent studies have explored hallucinations through the lens of internal representations, proposing mechanisms to decipher LLMs’ adherence to facts.
Approach: They propose to train a universal truthfulness hyperplane that distinguishes the model’s factually correct and incorrect outputs on a diverse collection of over 40 datasets and examine its cross-task, cross-domain, and in-domain generalization.
Outcome: The proposed model is able to distinguish factual outputs from incorrect outputs on a diverse collection of over 40 datasets.
Internal states before wait modulate reasoning patterns (2025.findings-emnlp)

Copied to clipboard

Challenge: Prior work has shown that a significant driver of performance in reasoning models is their ability to reason and self-correct.
Approach: They address the question whether model’s latents preceding wait tokens contain relevant information for modulating the subsequent reasoning process.
Outcome: The proposed model's latents preceding wait tokens contain relevant information for modulating the subsequent reasoning process.
Veri-R1: Toward Precise and Faithful Claim Verification via Online Reinforcement Learning (2026.findings-acl)

Copied to clipboard

Challenge: Existing approaches to online claim verification rely on prompt engineering or pre-designed reasoning workflows.
Approach: They propose an online reinforcement learning framework that enables an LLM to interact with a search engine and receive reward signals that explicitly shape its planning, retrieval, and reasoning behaviors.
Outcome: Empirical results show that Veri-R1 improves joint accuracy by 30% and doubles evidence score, often surpassing larger-scale model counterparts.
GRACE: Discriminator-Guided Chain-of-Thought Reasoning (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing language models (LMs) can assign a high likelihood to incorrect steps . Existing models (LLMs), however, struggle with complex multi-step reasoning.
Approach: They propose a stepwise decoding approach that steers the decoding process towards producing correct reasoning steps.
Outcome: The proposed approach outperforms existing methods on math and symbolic reasoning tasks.
CoMAT: Chain of Mathematically Annotated Thought Improves Mathematical Reasoning (2025.emnlp-main)

Copied to clipboard

Challenge: Mathematical reasoning remains a significant challenge for large language models (LLMs), despite advances in prompting techniques such as Chain-of-Thought (CoT).
Approach: They propose a framework that enhances reasoning through two stages: Symbolic Conversion and Reasoning Execution.
Outcome: The proposed framework outperforms traditional CoT on six out of seven benchmarks across four LLMs.
Context as a Tool: Context Management for Long-Horizon SWE-Agents (2026.findings-acl)

Copied to clipboard

Challenge: Existing large language models rely on append-only context maintenance or passively triggered compression heuristics, leading to context explosion, semantic drift, and degraded reasoning in long-running interactions.
Approach: They propose a new context management paradigm that elevates context maintenance to a callable tool . they propose 'cat' framework that injects context-management actions into complete interaction trajectories .
Outcome: The proposed model outperforms ReAct-based agents and static compression baselines on SWE-Verified tests.
Veracity Bias and Beyond: Uncovering LLMs’ Hidden Beliefs in Problem-Solving Reasoning (2025.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) have been aligned to avoid harmful biases and stereotypes, but recent studies have revealed the superficial nature of this alignment.
Approach: They propose to use large language models to avoid harmful biases and stereotypes by assigning personas to LLMs to observe decision discrepancies in social scenarios or asking them to associate specific attributes with social targets.
Outcome: The proposed models attribute fewer correct solutions and more incorrect ones to African-American groups in math and coding, while Asian authorships are least preferred in writing evaluation.
Weak2Wise: An Automated, Lightweight Framework for Weak-LLM-Friendly Reasoning Synthesis (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing approaches to finetuning large language models rely on expensive manual annotations or auxiliary models and fail to address the unique constraints of smaller "weak" LLMs.
Approach: Weak2Wise is a fully automated framework for synthesizing highquality, weak-LLM-friendly reasoning traces.
Outcome: Weak2Wise is a fully automated, lightweight framework for synthesizing highquality, weak-LLM-friendly reasoning traces.
VerifyMatch: A Semi-Supervised Learning Paradigm for Natural Language Inference with Confidence-Aware MixUp (2024.emnlp-main)

Copied to clipboard

Challenge: Natural language inference (NLI) is a key task for evaluating a model's ability to perform natural language understanding and reasoning.
Approach: They propose to construct pseudo-generated samples using class-specific fine-tuned large language models (LLMs) . they retain all pseudo-labeled samples, but use MixUp to ensure unlabele .
Outcome: The proposed approach achieves competitive accuracy compared to strong baselines for NLI datasets in low-resource settings.
Memorization or Reasoning? Exploring the Idiom Understanding of LLMs (2025.emnlp-main)

Copied to clipboard

Challenge: idioms have long posed a challenge due to their unique linguistic properties, which set them apart from other common expressions.
Approach: They propose to use a large-scale dataset of idioms in six languages to evaluate LLMs' idiomatic processing ability.
Outcome: The proposed model integrates contextual cues and reasoning to improve idiom understanding in LLMs, suggesting that their performance is influenced by memorization and reasoning.
Decoupling the Effect of Chain-of-Thought Reasoning: A Human Label Variation Perspective (2026.findings-acl)

Copied to clipboard

Challenge: Reasoning-tuned large language models (LLMs) with long Chain-of-Thought excel at single-answer tasks, yet their ability to model Human Label Variation remains underexplored.
Approach: They conduct systematic disentanglement experiments to isolate the effect of reasoning text from intrinsic model priors on distribution-based tasks.
Outcome: The proposed model improves distributional alignment, but distributional ranking is governed by model priors.
Format-Adapter: Improving Reasoning Capability of LLMs by Adapting Suitable Format (2026.findings-acl)

Copied to clipboard

Challenge: Prior work showed that multiple reasoning formats outperform a single format when generating multiple answers.
Approach: They propose a method to measure reasoning error when generating multiple answers . they propose 'formatadapter' which generates and selects suitable reasoning formats .
Outcome: The proposed method achieves a 4.3% performance improvement over previous works on math and commonsense reasoning tasks.
Reason-KE++: Aligning the Process, Not Just the Outcome, for Faithful LLM Knowledge Editing (2026.findings-acl)

Copied to clipboard

Challenge: Current methods for modifying parameters to integrate new knowledge are not accurate enough.
Approach: They propose an SFT+RL framework that instills process-level faithfulness by a stage-aware Reward mechanism and a Stage-assisted Reward Mechanism.
Outcome: The proposed framework instills process-level faithfulness while boosting final accuracy.
GeoLaux: A Benchmark for Evaluating MLLMs’ Geometry Performance on Long-Step Problems Requiring Auxiliary Lines (2026.acl-long)

Copied to clipboard

Challenge: Existing benchmarks for Geometry problem solving lack fine-grained evaluation for long-step problems necessitating auxiliary line construction.
Approach: They present a fine-grained annotated dataset with long-step reasoning and auxiliary line construction that provides a detailed evaluation of 23 leading MLLMs.
Outcome: The proposed model performs significantly worse on long-step problems than short-step ones, with 18 models showing a performance drop of over 50%.
MED-COREASONER: Reducing Language Disparities in Medical Reasoning via Language-Informed Co-Reasoning (2026.acl-long)

Copied to clipboard

Challenge: Existing models that use English and local languages have a multilingual gap . a language-informed co-reasoning framework can be used to improve multilingual reasoning .
Approach: They propose a language-informed co-reasoning framework that elicits parallel English and local-language reasoning and abstracts them into structured concepts.
Outcome: Experiments show that Med-CoReasoner improves multilingual reasoning performance by 5% . the framework produces clinically sound and culturally grounded reasoning traces .
TPA: Next Token Probability Attribution for Detecting Hallucinations in RAG (2026.acl-long)

Copied to clipboard

Challenge: Existing approaches attribute hallucinations to a binary conflict between internal knowledge stored in FFNs and the retrieved context.
Approach: They propose a framework which mathematically attributes each next-token probability to seven distinct sources and aggregates source attributions by POS tags to quantify contribution of each model component to the generation of specific linguistic categories within a response.
Outcome: Extensive experiments show that the proposed framework achieves state-of-the-art performance.
Two Pathways to Truthfulness: On the Intrinsic Encoding of LLM Hallucinations (2026.acl-long)

Copied to clipboard

Challenge: Previous work shows that large language models generate hallucinations, yet the origins and mechanisms of these signals remain unclear.
Approach: They propose to validate and disentangle two different pathways for truthfulness cues . they also propose to use the same mechanism to derive self-contained evidence from the generated answer .
Outcome: The proposed applications improve hallucination detection performance by integrating two different inputs.
Temp-R1: A Unified Autonomous Agent for Complex Temporal KGQA via Reverse Curriculum Reinforcement Learning (2026.acl-long)

Copied to clipboard

Challenge: Existing methods rely on fixed workflows and expensive closed-source APIs, limiting flexibility and scalability.
Approach: They propose a temporal reasoning agent that trains on difficult questions first . they expand the action space with specialized internal actions alongside external action .
Outcome: The proposed agent improves 19.8% over baselines on complex questions and multi-tasks.
Empowering Math Problem Generation and Reasoning for Large Language Model via Synthetic Data based Continual Learning Framework (2025.emnlp-main)

Copied to clipboard

Challenge: Existing learning frameworks for large language models (LLMs) for math problem generation are limited and lack quality data.
Approach: They propose a synthetic data based continual learning framework to improve LLMs ability for MPG and math reasoning.
Outcome: The proposed framework improves performance on large language models and math reasoning using supervised fine-tuning, data synthesis and direct preference optimization.
MentalSeek-Dx: Towards Progressive Hypothetico-Deductive Reasoning for Real-world Psychiatric Diagnosis (2026.acl-long)

Copied to clipboard

Challenge: Mental health disorders represent a burgeoning global public health challenge . lack of ecological validity and fine-grained diagnostic supervision limits their utility .
Approach: They propose a medical-specialized LLM trained to internalize clinical reasoning process through supervised trajectory construction and curriculum-based reinforcement learning.
Outcome: The proposed model achieves state-of-the-art with only 14B parameters, establishing a clinically grounded framework for reliable psychiatric diagnosis.
FOLIO: Natural Language Reasoning with First-Order Logic (2024.emnlp-main)

Copied to clipboard

Challenge: Existing benchmarks for logical reasoning in large language models lack language naturalness or limited complexity.
Approach: They propose to use first-order logic annotations to evaluate logical reasoning capabilities of large language models.
Outcome: The proposed dataset evaluates the FOL reasoning ability of supervised fine-tuning on medium-sized language models.
Hallucination Detection in LLMs Using Spectral Features of Attention Maps (2025.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) have demonstrated remarkable performance across tasks but remain prone to hallucinations.
Approach: They propose a method that uses attention maps to detect hallucinations . they propose to use top-k eigenvalues of the attention maps as input to probes .
Outcome: The proposed method achieves state-of-the-art hallucination detection performance among attention-based methods.
Can LLMs Estimate Cognitive Complexity of Reading Comprehension Items? (2026.acl-long)

Copied to clipboard

Challenge: Existing methods for estimating the cognitive complexity of reading comprehension items are expensive, time-consuming, and subject to rater variability.
Approach: They propose to use two dimensions to estimate cognitive complexity of RC items to focus on evidence Scope and transformation level to estimate the cognitive complexity.
Outcome: The proposed models can estimate the cognitive complexity of items by focusing on two dimensions—Evidence Scope and Transformation Level—that indicate the degree of cognitive burden involved in reasoning about the answer.
Full-Step-DPO: Self-Supervised Preference Optimization with Step-wise Rewards for Mathematical Reasoning (2025.findings-acl)

Copied to clipboard

Challenge: Existing approaches to improve long-chain mathematical reasoning focus on the first erroneous step, but ignore all other steps and rely heavily on external signals.
Approach: They propose a DPO framework that leverages step-wise rewards from the entire reasoning chain instead of optimizing only the first erroneous step.
Outcome: The proposed framework improves on in-domain and out-of-domain mathematical reasoning benchmarks.
How Can Input Reformulation Improve Tool Usage Accuracy in a Complex Dynamic Environment? A Study on tau-bench (2025.findings-emnlp)

Copied to clipboard

Challenge: Recent advances in reasoning and planning capabilities of large language models have enabled their potential as autonomous agents capable of tool use in dynamic environments.
Approach: They propose an input-reformulation multi-agent framework that reformulates user queries .
Outcome: The proposed framework outperforms ReAct, Function Calling, and Self-Reflection in overall pass5 scores.
RealFin: How Well Do LLMs Reason About Finance When Users Leave Things Unsaid? (2026.findings-acl)

Copied to clipboard

Challenge: General-purpose models tend to over-commit and guess, while most finance-specialized models fail to clearly identify missing premises.
Approach: They propose a bilingual benchmark that removes premises from exam-style questions while keeping them linguistically plausible.
Outcome: The proposed model overcommits and guesses while most finance-specialized models fail to clearly identify missing premises.
Neo-Classic: A Benchmark for Evaluating Linguistic-Aesthetic Reasoning in Classical Chinese Poetry (2026.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) achieve high accuracy on established Classical Chinese Poetry benchmarks, but it remains challenging to distinguish transferable Linguistic-Aesthetic Reasoning from reliance on familiar pre-training patterns.
Approach: They propose a benchmark that combines a constructionist Out-of-Sample dataset with reverse understanding probes to evaluate large-scale large-format models.
Outcome: The proposed model performs well on classical Chinese poetry benchmarks, but a performance gap persists . the model can complete famous couplets and can be used to understand a variety of texts.
OptiCo: Adaptive Distributed Training Optimization via Collaborative Agent Reasoning (2026.acl-long)

Copied to clipboard

Challenge: Existing distributed training frameworks are plagued by over-reliance on prior profiling and poor generalization across models/hardware.
Approach: They propose a model-driven multi-agent framework that leverages Large Language Models to enable automatic and explainable distributed training strategy configuration.
Outcome: The proposed framework outperforms expert-designed training strategies within 20 iterations.
REST: Stress Testing Large Reasoning Models by Asking Multiple Problems at Once (2026.acl-long)

Copied to clipboard

Challenge: Recent Large Reasoning Models (LRMs) lack a narrow evaluation paradigm . a single-question evaluation setup suffers from two major limitations .
Approach: They propose a stress-testing framework that exposes LRMs to multiple problems simultaneously.
Outcome: The proposed framework outperforms existing models on reasoning benchmarks and state-of-the-art models.
VecCISC: Improving Confidence-Informed Self-Consistency with Reasoning Trace Clustering and Candidate Answer Selection (2026.findings-acl)

Copied to clipboard

Challenge: Weighted majority voting requires a critic to evaluate each candidate’s reasoning trace to produce the answer’s confidence score.
Approach: They propose a lightweight framework that uses a measure of semantic similarity to filter reasoning traces that are semantically equivalent to others, degenerate, or hallucinated.
Outcome: The proposed framework reduces token usage by 47% while maintaining or exceeding the accuracy of CISC.
Incentivizing In-depth Reasoning over Long Contexts with Process Advantage Shaping (2026.findings-acl)

Copied to clipboard

Challenge: Reinforcement Learning with Verifiable Rewards (RLVR) has proven effective in enhancing LLMs’ short-context reasoning but falters in long-contemporal scenarios requiring precise grounding and multi-hop reasoning.
Approach: They propose a framework that constructs high-difficulty, multi-hop long-context QA pairs with inherent reasoning chains to overcome this bottleneck.
Outcome: The proposed framework outperforms RLVR baselines and matches frontier LLMs while using far fewer parameters.
Where Reasoning Breaks: Logic-Aware Path Selection by Controlling Logical Connectives in LLMs Reasoning Chains (2026.findings-acl)

Copied to clipboard

Challenge: a single transition error can propagate through the entire reasoning chain, leading to unstable performance.
Approach: They propose a framework that intervenes at logical connective junctions to improve LLMs' reasoning.
Outcome: The proposed framework achieves favorable accuracy–efficiency trade-off compared to global inference time scaling methods like beam search and self-consistency.
Exposing the Achilles’ Heel: Evaluating LLMs Ability to Handle Mistakes in Mathematical Reasoning (2025.acl-long)

Copied to clipboard

Challenge: Existing evaluations focus on final accuracy, neglecting the critical aspect of reasoning capabilities.
Approach: They propose to evaluate LLMs’ abilities to detect and correct reasoning mistakes by using rule-based methods and smaller language models.
Outcome: The proposed model outperforms existing models such as GPT-4o and GPT4 in both accuracy and accuracy, but lacks data contamination and memorization concerns.
From RAG to Agentic RAG for Faithful Islamic Question Answering (2026.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) are increasingly used for Islamic question answering, where ungrounded responses may carry serious religious consequences.
Approach: They propose a bilingual, bilingual, Arabic/English benchmark with atomic single-gold answers that measures hallucination and abstention.
Outcome: The proposed model improves accuracy and robustness even with a small model.
Beyond Monolithic Rewards: Hybrid Multi-Aspect Reward Optimization (2026.findings-acl)

Copied to clipboard

Challenge: Existing approaches to optimize for multimodal learning use a single reward mechanism, but they lack confidence calibration across domains.
Approach: They propose a hybrid reward and multi-aspect reward modeling framework that integrates model-based and rule-based reward paradigms for accuracy and confidence calibration.
Outcome: The proposed model improves accuracy and confidence calibration across multimodal tasks and introduces a generalized length-penalty reward to stabilize training and improve performance.
Logical Reasoning with Outcome Reward Models for Test-Time Scaling (2025.emnlp-main)

Copied to clipboard

Challenge: Logical reasoning is a critical benchmark for evaluating the capabilities of large language models (LLMs), but it is under-explored in deductive reasoning.
Approach: They propose to use Chain-of-Thought to generate data using single and multiple samples to train ORMs.
Outcome: The proposed model expands the type of errors covered in the training dataset, covering previously unexplored error types.
FLAMES: Improving LLM Math Reasoning via a Fine-Grained Analysis of the Data Synthesis Pipeline (2025.findings-emnlp)

Copied to clipboard

Challenge: Recent work improving LLM math reasoning with synthetic data uses unique setups, making comparison of data synthesis strategies impractical.
Approach: They propose a framework for LLM assessment of math reasoning with synthetic data . they use 10 existing data synthesis strategies and multiple other factors to study performance .
Outcome: The proposed data synthesis strategies outperform public datasets on OlympiadBench, CollegeMath, GSMPlus and MATH.
Uncertainty-Aware Answer Selection for Improved Reasoning in Multi-LLM Systems (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing approaches to selecting reliable responses from multiple LLMs often depend on external verifiers, human evaluators, or self-consistency techniques.
Approach: They propose a calibrated log-likelihood-based selection framework to improve multi-LLM performance.
Outcome: The proposed method outperforms majority voting and exceeds self-consistency performance when using a large number of model calls.
Learning from Mistakes: Negative Reasoning Samples Enhance Out-of-Domain Generalization (2026.acl-long)

Copied to clipboard

Challenge: Recent studies show that supervised fine-tuning (SFT) is a common approach for reasoning in large language models.
Approach: They propose to use supervised fine-tuning (SFT) on chain-of-thought trajectories demonstrations . they find that incorporating negative traxories yields substantial OOD generalization gains .
Outcome: The proposed scheme yields 5.51% OOD gain over positive-only training.
Chronos: Learning Temporal Dynamics of Reasoning Chains for Test-Time Scaling (2026.findings-acl)

Copied to clipboard

Challenge: Existing methods for testing time scales treat reasoning traces or tokens equally, ignoring substantial variations in trajectory quality and localized logical failures.
Approach: They propose a chronological reasoning scorer that models each trajectory as a time series.
Outcome: The proposed method achieves relative improvements of 34.21% over Pass@128 and 22.70% over Maj@135 on HMMT25, highlighting its effectiveness.
Dynamic Sampling that Adapts: Self-Aware Iterative Data Persistent Optimization for Mathematical Reasoning (2026.findings-acl)

Copied to clipboard

Challenge: Current data selection paradigms rely on static, externally defined metrics, which fail to adapt to the evolving capabilities of models during training.
Approach: They propose a dynamic sampling framework that aligns training data with the model's intrinsic competence by iterating on real-time feedback.
Outcome: Extensive experiments on eight benchmarks show that SAI-DPO outperforms static baselines at most nearly 6 points, achieving state-of-the-art efficiency with significantly less data.
Pun Unintended: LLMs and the Illusion of Humor Understanding (2025.emnlp-main)

Copied to clipboard

Challenge: Existing models for pun detection lack nuanced grasp typical of human interpretation.
Approach: They analyze existing pun detection benchmarks and human evaluation across recent LLMs to find subtle changes in puns that mislead LLM.
Outcome: The proposed models lack the nuance typical of human interpretation and lack the depth of their analysis to detect puns.
ReCreate: Reasoning and Creating Domain Agents Driven by Experience (2026.acl-long)

Copied to clipboard

Challenge: Large Language Model (LLM) agents are reshaping the industrial landscape, but tasks differ widely, making them labor-intensive to build.
Approach: They propose an experience-driven framework for the automatic creation of domain agents . they leverage agent interaction histories to provide rich concrete signals on success or failure .
Outcome: The proposed framework outperforms human-designed agents and existing methods in experiments across diverse domains.
VANE: Guiding High-Value Exploration in RLVR via Outcome-Process Novelty Shaping (2026.findings-acl)

Copied to clipboard

Challenge: Extensive experiments on large-scale mathematical reasoning and out-of-distribution tasks demonstrate the effectiveness and generalization of the proposed method.
Approach: They propose a method that quantifies novelty across the outcome space and semantic process space by using reward or solution divergence.
Outcome: Experiments on Qwen2.5-Math-7B demonstrate the proposed method is general and efficient.
Towards Generalizable and Faithful Logic Reasoning over Natural Language via Resolution Refutation (2024.lrec-main)

Copied to clipboard

Challenge: Large language models (LLMs) have achieved significant performance in various natural language reasoning tasks, but struggle with performing first-order logic reasoning over formal logical theories expressed in natural language.
Approach: They propose a framework which introduces the paradigm of resolution refutation to solve first-order logic reasoning problems by extending reasoning rules and employing the principle of proof by contradiction.
Outcome: The proposed framework outperforms existing models while maintaining performance in simple scenarios.
Investigating the interaction of linguistic and mathematical reasoning in language models using multilingual number puzzles (2025.emnlp-main)

Copied to clipboard

Challenge: Across languages, numeral systems vary widely in how they construct and combine numbers.
Approach: They conduct experiments to examine the linguistic and mathematical aspects of numbers in language.
Outcome: The models can't solve linguistic-mathematical puzzles involving cross-linguistic numeral systems, the authors found . they lack the ability to flexibly infer compositional rules from implicit patterns in human-scale data.
Logit Space Constrained Fine-Tuning for Mitigating Hallucinations in LLM-Based Recommender Systems (2025.emnlp-main)

Copied to clipboard

Challenge: Existing LLM-based recommender systems rely on standard fine-tuning methodologies, often ignoring hallucination issues during the fine-uning process.
Approach: They propose a logit space constraint-based fine-tuning framework to mitigate hallucination in LLM-based recommenders by incorporating Kullback–Leibler divergence into the training objective.
Outcome: Experiments on two recommendation models with distinct LLM backbones and four real-world datasets show that LCFT reduces hallucination and enhances recommendation performance.
Subtle Errors in Reasoning: Preference Learning via Error-injected Self-editing (2025.acl-long)

Copied to clipboard

Challenge: Existing studies to improve mathematical ability typically involve applying preference learning to step-wise solution pairs, but they overlook critical subtle errors.
Approach: They propose a preference learning framework that injects predefined subtle errors into pivotal tokens to construct hard pairs for error mitigation.
Outcome: Extensive experiments show that the proposed framework improves on Qwen2-7B-Instruct and MATH with 4.5K training samples.
How Do Answer Tokens Read Reasoning Traces? Self-Reading Patterns in Thinking LLMs for Quantitative Reasoning (2026.findings-acl)

Copied to clipboard

Challenge: Prior work on activation steering has focused on shaping reasoning traces, but it remains unclear how answer tokens actually read and integrate the reasoning to produce reliable outcomes.
Approach: They propose a training-free steering method that uses self-reading quality scores to guide inference toward benign self-readiness and away from uncertain and disorganized reading.
Outcome: The proposed method yields consistent accuracy gains in the reasoning traces generated by thinking LLMs.
Schoenfeld’s Anatomy of Mathematical Reasoning by Language Models (2026.acl-long)

Copied to clipboard

Challenge: Large language models expose reasoning traces, yet their underlying cognitive structure and steps remain difficult to identify and analyze beyond surface-level statistics.
Approach: They propose a framework that explicitly abstracts reasoning traces into functional reasoning steps such as Analysis, Explore, Implement, Verify, etc.
Outcome: The proposed framework reveals reproducible thinking dynamics and structural differences between reasoning and non-reasoning models, which are not apparent from token-level views.
Analysing Chain of Thought Dynamics: Active Guidance or Unfaithful Post-hoc Rationalisation? (2025.emnlp-main)

Copied to clipboard

Challenge: Recent work has demonstrated that using chain of thought (CoT) on soft-reasoning tasks can yield limited or even negative performance gains.
Approach: They investigate how chain of thought (CoT) is used in soft-reasoning tasks across instruction-tuned, reasoning and reasoning-distilled models.
Outcome: The proposed model can steer predictions without faithfully reflecting reasoning, indicating a disconnect between CoT influence and faithfulness.
Reasoning in a Combinatorial and Constrained World: Benchmarking LLMs on Natural-Language Combinatorial Optimization (2026.findings-acl)

Copied to clipboard

Challenge: NLCO evaluates large language models for combinatorial optimization (CO) . existing evaluations emphasize relatively simple reasoning competencies .
Approach: They propose a combinatorial optimization benchmark that evaluates large language models on CO reasoning.
Outcome: The proposed model can handle combinatorial optimization without writing code or calling external solvers.
DARL: Encouraging Diverse Answers for General Reasoning without Verifiers (2026.findings-acl)

Copied to clipboard

Challenge: Recent efforts such as RLPR have extended RLVR to general domains, enabling training on broader datasets and achieving improvements over RL PR.
Approach: They propose a framework that encourages the generation of diverse answers within a controlled deviation range from the reference while preserving alignment with it.
Outcome: Extensive experiments on 13 benchmarks show that DARL surpasses RLPR in both reasoning accuracy and output diversity.
WikiSplit++: Easy Data Refinement for Split and Rephrase (2024.lrec-main)

Copied to clipboard

Challenge: Existing text simplification methods rely on encoder-decoder models to achieve this task.
Approach: They propose a text-to-text generation approach that applies encoder-decoder models to a large-scale dataset to improve Split and Rephrase.
Outcome: The proposed approach improves Split and Rephrase readability and performance on large datasets, but still suffers from hallucinations and under-splitting.
Improving Rule-based Reasoning in LLMs using Neurosymbolic Representations (2025.emnlp-main)

Copied to clipboard

Challenge: Large language models (LLMs) face challenges in reliably solving reasoning tasks, especially when solving tasks that require strict rule following.
Approach: They propose a method that encodes hidden states into neurosymbolic vectors and decodes them into a neurosample vector space to enable problem-solving within a neural space.
Outcome: The proposed method shows an average of 88.6% lower cross-entropy loss and 15.4 times more problems correctly solved on a suite of mathematical reasoning tasks compared to chain-of-thought prompting and supervised fine-tuning (LoRA).
Training LLMs for Divide-and-Conquer Reasoning Elevates Test-Time Scalability (2026.acl-long)

Copied to clipboard

Challenge: Large language models have demonstrated strong reasoning capabilities through step-by-step chain-of-thought (CoT) reasoning, but their strictly sequential nature constrains test-time scalability.
Approach: They propose an end-to-end reinforcement learning framework to enhance LLMs' DAC-style reasoning capacity by decomposing a problem into subproblems and solving them sequentially.
Outcome: The proposed model surpasses CoT by 8.6% and 6.3% on competition-level benchmarks and is available at the [github.com/MasterVito/DAC-RL].
SMART: Evaluating LLMs’ Mathematical Reasoning via a Human Cognitive Process-Inspired Benchmark (2026.acl-long)

Copied to clipboard

Challenge: Existing evaluation methods focus on the final answer or on the intermediate reasoning steps, overlooking its inherently multi-stage and multi-dimensional nature.
Approach: They propose a benchmark that decomposes mathematical problem-solving into four cognitive dimensions and introduces dimension-specific tasks to measure their cognitive processes.
Outcome: The proposed model decomposes mathematical problem-solving into four cognitive dimensions and introduces dimension-specific tasks to measure their cognitive processes.
Beyond High-Entropy Exploration: Correctness-Aware Low-Entropy Segment-Based Advantage Shaping for Reasoning LLMs (2026.findings-acl)

Copied to clipboard

Challenge: Recent work studies RLVR through token entropy, arguing that high-entropies drive exploration and should receive stronger updates.
Approach: They propose a correctness-aware reinforcement framework that performs fine-grained advantage modulation over low-entropy segments.
Outcome: The proposed framework improves accuracy over strong RL baselines across three backbones and six math benchmarks while maintaining high-entropy exploration.
PruneCD: Contrasting Pruned Self Model to Improve Decoding Factuality (2025.emnlp-main)

Copied to clipboard

Challenge: Existing approaches to decode early exit logits in large language models are lacking in factuality and accuracy.
Approach: They propose a contrastive decoding method that constructs the amateur model via layer pruning rather than early exit.
Outcome: The proposed method improves factuality with minimal inference overhead and is robust and practical.
Word Salad Chopper: Reasoning Models Waste A Ton Of Decoding Budget On Useless Repetitions, Self-Knowingly (2025.emnlp-main)

Copied to clipboard

Challenge: Large Reasoning Models (LRMs) are often bottlenecked by the high cost of output tokens.
Approach: They propose a lightweight, turnkey component for Large Reasoning Models that is minimally invasive to its reasoning trajectory.
Outcome: The proposed component is lightweight and low overhead, and lacks semantic value.
Beyond Reasoning Gains: Mitigating General-Capability Forgetting in Large Reasoning Models (2026.findings-acl)

Copied to clipboard

Challenge: Reinforcement learning with verifiable rewards (RLVR) has delivered impressive gains in mathematical and multimodal reasoning . however, the recipe introduces a significant risk of capability regression, where models forget foundational skills after prolonged training without employing regularization strategies.
Approach: They propose a replay strategy with dynamic objective reweighting for general knowledge preservation using short-horizon signals of convergence and instability.
Outcome: The proposed method preserves general capabilities and improves reasoning . it can be applied to existing RLVR pipelines without training additional models or tuning .
CCD: Mitigating Hallucinations in Radiology MLLMs via Clinical Contrastive Decoding (2026.findings-acl)

Copied to clipboard

Challenge: Multimodal large language models generate medical hallucinations due to over-sensitivity to clinical sections.
Approach: They propose a framework that integrates structured clinical signals from task-specific radiology expert models.
Outcome: The proposed framework improves overall performance on radiology report generation (RRG) on the MIMIC-CXR dataset, it yields up to 17% improvement in RadGraph-F1.
FLAIR: Steering LLM Mathematical Problem Solving based on A Fuzzy-Logic-AssIsted Reasoner (2026.acl-long)

Copied to clipboard

Challenge: Existing approaches to mathematical reasoning rely on static heuristics or pre-determined reasoning strategies.
Approach: They propose an adaptive framework that integrates fuzzy theory into LLM-based mathematical reasoning.
Outcome: The proposed framework outperforms state-of-the-art models while offering effective and interpretable diagnostics of intermediate problem-solving states.
HalluAudio: A Comprehensive Benchmark for Hallucination Detection in Large Audio-Language Models (2026.acl-long)

Copied to clipboard

Challenge: Existing studies on hallucination focus on text or vision, while few audio-oriented studies are limited in scale, modality coverage, and diagnostic depth.
Approach: They propose a large-scale benchmark for evaluating hallucinations across speech, sound, and music.
Outcome: The proposed model improves hallucination rate, yes/no bias, error-type analysis, and refusal rate.
KnowRL: Exploring Knowledgeable Reinforcement Learning for Factuality (2026.acl-long)

Copied to clipboard

Challenge: Existing Reinforcement Learning approaches rely on outcome-oriented rewards to reinforce fabricated reasoning paths when the final answer is correct.
Approach: They propose a framework that integrates factual supervision directly into reasoning . they propose to decompose chain of thought into atomic facts and verify them against ground-truth knowledge .
Outcome: The proposed framework reduces the Incorrect Rate on SimpleQA by 20.3% while maintaining strong performance on complex reasoning benchmarks.
Assessing the Belief Consistency of Large Language Models on the Logical Conversation Process (2026.acl-long)

Copied to clipboard

Challenge: Large language models have been shown remarkable ability to understand given contexts.
Approach: They propose a method to evaluate whether beliefs held by LLMs remain consistent . they propose to use multiple choice question answering format to assess belief consistency .
Outcome: The proposed method evaluates the consistency of LLMs in a multiple-choice question answering format.
Decoupled Reasoning with Implicit Fact Tokens (DRIFT): A Dual-Model Framework for Efficient Long-Context Inference (2026.findings-acl)

Copied to clipboard

Challenge: Existing solutions to integrate extensive, dynamic knowledge into Large Language Models (LLMs) are constrained by finite context windows, retriever noise, or the risk of catastrophic forgetting.
Approach: They propose a dual-model architecture that explicitly decouples knowledge extraction from the reasoning process by compressing document chunks into implicit fact tokens conditioned on the query.
Outcome: The proposed architecture significantly outperforms strong baselines among comparably sized models on long-context tasks while maintaining inference accuracy.
A Survey of Reasoning-Intensive Retrieval: Progress and Challenges (2026.acl-long)

Copied to clipboard

Challenge: Reasoning-Intensive Retrieval (RIR) targets retrieval settings where relevance is mediated by latent inferential links between a query and supporting evidence, rather than semantic similarity.
Approach: They propose a taxonomy that categorizes methods based on where and how reasoning is integrated into the retrieval pipeline.
Outcome: The proposed method framework provides a detailed analysis of the current landscape and its trade-offs and practical applications.
Reducing Hallucinations in LLMs via Factuality-Aware Preference Learning (2026.findings-acl)

Copied to clipboard

Challenge: Preference alignment methods can reinforce hallucinations when preference judgments reward fluency and confidence over factual correctness.
Approach: They propose a method that corrects misordered preference pairs and adds a factuality-aware margin to emphasize pairs with clear correctness differences.
Outcome: The proposed method improves factuality and reduces hallucination rates across seven open-weight LLMs.
STaD: Scaffolded Task Design for Identifying Compositional Skill Gaps in LLMs (2026.findings-acl)

Copied to clipboard

Challenge: Existing benchmarks often treat complex tasks as monolithic, resulting in inconsistent performance and inconsistent explanations.
Approach: They propose a framework for creating controlled variations of benchmark tasks based on the concept of scaffolding, which introduces structured, incremental support in a step-by-step manner.
Outcome: The proposed framework enables systematic probing of model behavior by identifying the specific reasoning skill compositions they lack.
LLMs as Lab Engineers: A Benchmark for Analytical Method Lifecycle Management (2026.findings-acl)

Copied to clipboard

Challenge: General-purpose commercial models outperform domain-specialized ones, while RAG and reasoning significantly improve performance.
Approach: They propose a benchmark to evaluate LLMs' capabilities in analytical chemistry scenarios.
Outcome: The proposed framework outperforms existing benchmarks focused on factual knowledge and provides practical guidance for analytical chemistry challenges.
ParaSuite: Boosting LLM Reasoning via Paradox Resolution (2026.acl-long)

Copied to clipboard

Challenge: Existing benchmarks for paradox research focus on checking basic logical consistency and not reflective reasoning.
Approach: They propose a pipeline dedicated to paradox research that automates data synthesis, evaluation, and training.
Outcome: The proposed pipeline improves paradoxical and general STEM reasoning.
Reinforcement Learning for Diffusion LLMs via Energy-Based Gibbs Alignment (2026.acl-long)

Copied to clipboard

Challenge: Diffusion Large Language Models (dLLMs) offer parallel decoding and bidirectional context modeling . aligning dLLms with reinforcement learning (RL) remains a challenge .
Approach: They propose a variational framework that reformulates RL for dLLMs as a distribution matching problem.
Outcome: The proposed framework reformulates RL for dLLMs as a distribution matching problem.
DiNO: Disinformation Narrative Observer (2026.acl-long)

Copied to clipboard

Challenge: Disinformation is an escalating global threat, making it essential to understand its content, dissemination, and evolution.
Approach: They propose a method to extract disinformation narratives from news articles . they evaluated how well their topics and stances aligned with a recognized disinformation dataset.
Outcome: The proposed method outperforms other narrative mining methods in analyzing disinformation narratives.
Tackling the Root of Misinformation by Teaching Laypeople about Logical Fallacies via Socratic Questioning and Critical Argumentation (2026.acl-long)

Copied to clipboard

Challenge: Existing systems that detect logical fallacies in public discourse do not help people recognize them independently.
Approach: They propose an intelligent tutoring system which uses large language models to help humans learn about logical fallacies.
Outcome: The proposed system outperforms baseline LLMs lacking such pedagogical strategies.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations