Challenge: Current evaluations of FEC models that depend on factuality metrics are not reliable and detailed enough.
Approach: They propose a fine-grained evaluation framework that automatically evaluates FEC models on different error categories.
Outcome: The proposed evaluation framework compares models on different error categories and finds the best training modes and significant differences in the performance of existing models.

Similar Papers

Annotating and Detecting Fine-grained Factual Errors for Dialogue Summarization (2023.acl-long)

Copied to clipboard

Challenge: Existing work on factual inconsistency in abstractive summarization addresses this problem.
Approach: They propose a dataset with fine-grained factual error annotations named DIASUMFACT and an unsupervised model named ENDERANKER.
Outcome: The proposed model performs on par with the state-of-the-art models while requiring fewer resources.
Understanding Factual Errors in Summarization: Errors, Summarizers, Datasets, Error Detectors (2023.acl-long)

Copied to clipboard

Challenge: Abstractive summarization systems still include factual errors in generated summaries despite recent improvements in factuality detection .
Approach: They aggregate factuality error annotations from nine existing datasets and stratify them according to the underlying summarization model.
Outcome: The proposed method improves on the ChatGPT-based model and shows that it is not superior for all error types.
Annotating and Modeling Fine-grained Factuality in Summarization (2021.naacl-main)

Copied to clipboard

Challenge: Recent abstractive summarization systems produce factual errors that are not faithful to the input . current methods are lacking in identifying what errors are most important to target .
Approach: They use synthetic and human-labeled data to identify factual errors in summarization and train models on the factuality detection task.
Outcome: The proposed model detects factual errors on word, dependency, and sentence levels.
Evaluating Factuality in Cross-lingual Summarization (2023.findings-acl)

Copied to clipboard

Challenge: Existing evaluation metrics for monolingual summarization require translation to evaluate the factuality of cross-lingual summmarization.
Approach: They propose to analyze cross-lingual factuality by collecting annotations and generated summaries from models at summary level and sentence level.
Outcome: The proposed dataset shows that over 50% of generated summaries contain factual errors with different characteristics from monolingual summarization.
Understanding Factuality in Abstractive Summarization with FRANK: A Benchmark for Factuality Metrics (2021.naacl-main)

Copied to clipboard

Challenge: Modern summarization models generate fluent but often factually unreliable outputs.
Approach: They propose to use human annotations to identify different categories of factual errors and benchmark factuality metrics to improve summarization evaluation.
Outcome: The proposed method identifies the proportion of different categories of factual errors and benchmarks their human judgements as well as their specific strengths and weaknesses.
CONFIT: Toward Faithful Dialogue Summarization with Linguistically-Informed Contrastive Fine-tuning (2022.naacl-main)

Copied to clipboard

Challenge: Factual inconsistencies in generated summaries severely limit the practical applications of abstractive dialogue summarization.
Approach: They propose a typology of factual errors to better understand hallucinations generated by current models and a contrastive fine-tuning strategy to improve the factual consistency and overall quality of summaries.
Outcome: The proposed model significantly reduces all kinds of factual errors on both SAMSum dialogue summarization and AMI meeting summarizing datasets.
Shortcomings of Question Answering Based Factuality Frameworks for Error Localization (2023.eacl-main)

Copied to clipboard

Challenge: Abstractive summarization systems often generate summaries with factual errors . many approaches to detect these errors have been proposed, but this capability has not been evaluated in past research .
Approach: They propose to use question answering-based factuality metrics to detect errors in summaries . they find that QA-based frameworks fail to correctly identify error spans in generated summary .
Outcome: The proposed methods outperform trivial exact match baselines in localizing errors in summaries.
ReFEree: Reference-Free and Fine-Grained Method for Evaluating Factual Consistency in Real-World Code Summarization (2026.acl-long)

Copied to clipboard

Challenge: Existing methods for evaluating factual consistency are primarily designed for short summaries of isolated code snippets.
Approach: They propose a reference-free and fine-grained method for evaluating factual consistency in real-world code summaries.
Outcome: The proposed method achieves highest correlation with human judgment among 13 baselines, improving 15-18% over the previous state-of-the-art.
Verify with Caution: The Pitfalls of Relying on Imperfect Factuality Metrics (2025.findings-acl)

Copied to clipboard

Challenge: Recent advances in large language models have led to optimism that they can serve as reliable evaluators of natural language outputs.
Approach: They propose to use factuality metrics to evaluate natural language outputs . they find they misestimate the factual accuracy of NLG systems .
Outcome: The proposed metrics are inconsistent with each other and often misestimate the factual accuracy of NLG systems, causing biases against paraphrased outputs and outputs that draw upon faraway parts of the source documents.
Exploring the Factual Consistency in Dialogue Comprehension of Large Language Models (2024.naacl-long)

Copied to clipboard

Challenge: LLMs generate responses following user's instructions, which requires high dialogue comprehension ability.
Approach: They propose to evaluate LLMs' dialogue comprehension ability using a dialogue summarization task to derive factual questions from the generated summaries and use them as a more flexible measurement of dialogue comprehension.
Outcome: The proposed model reduces the error rate by 11% on the dialogue summarization task.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations