Challenge: FinDVer is a benchmark to evaluate the explainable claim verification capabilities of LLMs . financial documents are typically long, intricate and dense, and they include both quantita and numerical reasoning.
Approach: They propose a benchmark to evaluate the explainable claim verification capabilities of LLMs . they assess 25 LLM systems under long-context and RAG settings .
Outcome: The proposed benchmark can be used to evaluate the explainable claim verification capabilities of LLMs in financial documents.

Similar Papers

Claim Verification in the Age of Large Language Models: A Survey (2026.acl-srw)

Copied to clipboard

Challenge: Recent election cycles have seen a large number of false information spread across social media and news platforms.
Approach: They propose a framework for automated claim verification using Large Language Models and Retrieval Augmented Generation.
Outcome: The proposed frameworks are based on large-scale models and new methods such as Retrieval Augmented Generation (RAG).
Assessing the Reasoning Capabilities of LLMs in the context of Evidence-based Claim Verification (2025.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) have shown remarkable proficiency in complex tasks where reasoning capabilities are paramount.
Approach: They propose a framework to break down claims into atomic reasoning types needed for verification.
Outcome: The proposed framework breaks down claims into atomic reasoning types needed for verification.
Towards Effective Extraction and Evaluation of Factual Claims (2025.acl-long)

Copied to clipboard

Challenge: Lack of a standardized evaluation framework impedes assessment and comparison of claim extraction methods.
Approach: They propose a framework for evaluating claim extraction in the context of fact-checking . they also introduce Claimify, an LLM-based claim extraction method .
Outcome: The proposed evaluation framework outperforms existing methods in the evaluation of claim extraction methods.
FinDABench: Benchmarking Financial Data Analysis Ability of Large Language Models (2025.coling-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) have demonstrated impressive capabilities across a wide range of tasks, but their proficiency and reliability in the specialized domain of financial data analysis remain uncertain.
Approach: FinDABench is a benchmark designed to evaluate the financial data analysis capabilities of Large Language Models (LLMs) it comprises 15,200 training instances and 8,900 test instances, all meticulously crafted by human experts.
Outcome: FinDABench measures the financial data analysis capabilities of large language models (LLMs) across three dimensions: 1) Core Ability; 2) Analytical Ability; 3) Technical Ability.
SciVer: Evaluating Foundation Models for Multimodal Scientific Claim Verification (2025.acl-long)

Copied to clipboard

Challenge: Existing scientific claim verification benchmarks focus on textual content alone or on verifying claims based on a single table.
Approach: They propose to use SciVer to evaluate the ability of foundation models to verify claims within a multimodal scientific context.
Outcome: The proposed model outperforms 21 state-of-the-art models and human experts on SciVer.
Are Large Language Models Reliable Reviewers? A Benchmark for Error Detection in Financial Documents (2026.findings-acl)

Copied to clipboard

Challenge: Existing LLMs struggle to identify errors in financial documents, a study shows . 18% of financial practitioners make errors daily, one-third make errors several times weekly, and 59% make errors multiple times monthly.
Approach: They introduce FinED-Bench, a publicly available Benchmark for financial error detection . it covers nine real-world financial scenarios and includes over 900 documents in 2025 . supervised fine-tuning can significantly improve the performance of weaker LLMs, they show .
Outcome: The proposed benchmark covers nine real-world financial scenarios and includes over 900 documents reported in 2025 that are unseen by existing language models.
MultiFC: A Real-World Multi-Domain Dataset for Evidence-Based Fact Checking of Claims (D19-1)

Copied to clipboard

Challenge: Existing efforts to verify factual claims are limited by small datasets or artificially constructed datasets.
Approach: They propose to use the largest publicly available dataset of naturally occurring factual claims for automatic claim verification.
Outcome: The proposed model outperforms baseline models and evidence pages significantly.
A Systematic Survey of Claim Verification: Corpora, Systems, and Case Studies (2025.findings-emnlp)

Copied to clipboard

Challenge: This survey analyses 198 studies published between January 2022 and March 2025 .
Approach: This survey synthesizes recent advances in CV corpus creation and system design.
Outcome: The results of this study are synthesized from 198 studies published between January 2022 and March 2025.
ClaimVer: Explainable Claim-Level Verification and Evidence Attribution of Text Through Knowledge Graphs (2024.findings-emnlp)

Copied to clipboard

Challenge: Despite the fact that many fact-checking tools lack granularity and explainability, they lack the ability to be useful in various contexts.
Approach: They propose a text validation framework that provides granular explanations for each claim and localizes the specific problematic content to reduce cognitive load.
Outcome: The proposed framework provides granular explanations for each claim prediction and localizes and educates users on the specific content.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations