Challenge: Text generated by Large Language Models (LLMs) may contain plausible but incorrect information known as hallucinations.
Approach: They extend the label set for verdict prediction to capture claim-evidence relationships humans would commonly interpret as supported or refuted.
Outcome: The proposed system improves F1 by 4 percentage points compared to baseline.

Similar Papers

EuroVerdict: A Multilingual Dataset for Verdict Generation Against Misinformation (2025.findings-acl)

Copied to clipboard

Challenge: a global issue that shapes public discourse shapes opinion and decision-making . many multilingual work has focused on claim verification rather than generating explanatory verdicts .
Approach: They propose a multilingual dataset designed for verdict generation covering eight European languages.
Outcome: The EuroVerdict dataset covers claims, manual verdicts, and supporting evidence . it is compared with other datasets in eight European languages .
Annotation Study of Japanese Judgments on Tort for Legal Judgment Prediction with Rationales (2022.lrec-1)

Copied to clipboard

Challenge: An annotation scheme for Japanese judgment documents is proposed to provide a reliable dataset for Legal Judgment Prediction (LJP) the anticipated cost of LJP will be much lower than that of human legal professionals.
Approach: They propose to build an annotation scheme for legal judgment prediction, especially for torts, which extracts decisions and rationales at character-level.
Outcome: The proposed annotation scheme can produce a dataset of Japanese LJP at reasonable reliability.
Factcheck-Bench: Fine-Grained Evaluation Benchmark for Automatic Fact-checkers (2024.findings-emnlp)

Copied to clipboard

Challenge: Large language models generate naturally sounding answers over a broad range of human inquiries, but they often generate answers that contradict real-world facts.
Approach: They propose a framework for annotating and evaluating the factuality of large language models . they propose 'factcheck-bench' which provides a multi-stage annotation scheme .
Outcome: The proposed framework outperforms several popular LLM fact-checkers in claim, sentence, and document levels.
AFaCTA: Assisting the Annotation of Factual Claim Detection with Reliable LLM Annotators (2024.acl-long)

Copied to clipboard

Challenge: generative AI is a counter-measure to misinformation, but factual claim detection suffers from inconsistency in definitions and high cost of manual annotation.
Approach: They propose a framework that assists in the annotation of factual claims with the help of large language models.
Outcome: The proposed framework can be used to annotate factual claims with the help of large language models and can work with or without expert supervision.
A Comprehensive Evaluation of Large Language Models on Legal Judgment Prediction (2023.findings-emnlp)

Copied to clipboard

Challenge: Large language models (LLMs) have demonstrated great potential for domain-specific applications, such as the law domain.
Approach: They propose a framework to investigate LLMs' competence in the law domain by using similar cases and multi-choice options.
Outcome: The proposed solutions can be extended to other domains to facilitate evaluations in other domain.
FactCG: Enhancing Fact Checkers with Graph-Based Multi-Hop Data (2025.naacl-long)

Copied to clipboard

Challenge: Prior research on training grounded factuality classification models to detect hallucinations in large language models (LLMs) has relied on public natural language inference (NLI) data and synthetic data.
Approach: They propose a method that leverages multi-hop reasoning on context graphs extracted from documents to generate complex multi-level claims without relying on LLMs to decide data labels.
Outcome: The proposed model outperforms GPT-4-o on the LLM-Aggrefact benchmark with much smaller model size.
Annotation Artifacts in Natural Language Inference Data (N18-2)

Copied to clipboard

Challenge: Large-scale datasets for natural language inference are created by crowdsourcing annotations . authors show that success of natural language models to date has been overestimated .
Approach: They propose a method for crowdsourcing annotations to generate 3 new sentences based on a sentence (premise) they show that a simple text categorization model can correctly classify the hypothesis alone in about 67% of SNLI and 53% of MultiNLI .
Outcome: The proposed model can classify the hypothesis alone in 67% of SNLI and 53% of MultiNLI datasets.
Validity Assessment of Legal Will Statements as Natural Language Inference (2022.findings-emnlp)

Copied to clipboard

Challenge: This study introduces a dataset that focuses on the validity of statements in legal wills.
Approach: They propose a dataset that focuses on the validity of statements in legal wills.
Outcome: The proposed model achieves 80% macro F1 and accuracy, but group accuracy is in mid 80s at best, suggesting that the models’ understanding of the task remains superficial.
NLI4CT: Multi-Evidence Natural Language Inference for Clinical Trial Reports (2023.emnlp-main)

Copied to clipboard

Challenge: Clinical trial reports (CTRs) are indispensable for the development of personalized medicine.
Approach: They propose a resource to help researchers interpret clinical trial reports . they use natural language inference to compute textual entailment .
Outcome: The proposed resource is the first to cover interpretation of full clinical trial reports . it includes tasks to determine inference relation between natural language statements and CTRs .
FactReasoner: A Probabilistic Approach to Long-Form Factuality Assessment for Large Language Models (2025.findings-emnlp)

Copied to clipboard

Challenge: Large language models often fail to ensure factual accuracy of outputs thus limiting reliability in real-world applications.
Approach: They propose a neuro-symbolic based factuality assessment framework that employs probabilistic reasoning to evaluate the truthfulness of long-form generated responses.
Outcome: The proposed framework outperforms state-of-the-art prompt-based methods in factual accuracy and recall.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations