Challenge: FactPICO is a factuality benchmark for plain language summarization of medical texts describing randomized controlled trials . existing metrics for factual summarizing medical evidence are poorly correlated with expert judgments on the instance level.
Approach: They propose a factuality benchmark for plain language summarization of medical texts . they assess factuality of critical elements of RCTs in those summaries .
Outcome: The proposed benchmark assesses the factuality of medical summaries using LLMs . the summary summators are based on 345 plain language summaires with fine-grained evaluation .

Similar Papers

Understanding LLMs’ summarization capabilities: an analysis of biomedical abstract and lay summary generation (2026.findings-acl)

Copied to clipboard

Challenge: Abstracts use technical language for academic audiences, while lay summaries aim to make findings accessible to non-specialists.
Approach: They evaluate the performance of lightweight LLMs in generating biomedical abstracts and lay summaries in a zero-shot setting.
Outcome: The proposed models perform well in generating biomedical abstracts and lay summaries in a zero-shot setting.
Summarizing, Simplifying, and Synthesizing Medical Evidence using GPT-3 (with Varying Success) (2023.acl-short)

Copied to clipboard

Challenge: Large language models are capable of producing high quality summaries of general domain news articles in few- and zero-shot settings, but it is unclear whether they are similarly capable in more specialized domains such as biomedicine.
Approach: They use GPT-3 to generate single- and multi-document summaries of biomedical articles, given no supervision, using a set of annotations.
Outcome: The proposed model outperforms fully supervised models in generic news summarization, but struggles to synthesize evidence across multiple documents.
SummEdits: Measuring LLM Ability at Factual Reasoning Through The Lens of Summarization (2023.emnlp-main)

Copied to clipboard

Challenge: Existing factual consistency benchmarks are inadequate to detect factual inconsistencies in LLMs.
Approach: They propose a protocol for inconsistency detection benchmark creation and implement it in a 10-domain benchmark called SummEdits.
Outcome: The proposed method is 20 times more cost-effective per sample and highly reproducible, as it estimates inter-annotator agreement at about 0.9.
Evaluating the Factuality of Zero-shot Summarizers Across Varied Domains (2024.eacl-short)

Copied to clipboard

Challenge: Recent work has shown that large language models can generate zero-shot summaries without explicit supervision that are often comparable or even preferred to manually composed reference summary.
Approach: They evaluate large language models (LLMs) that generate zero-shot summaries without explicit supervision that are often comparable to manual reference summary . they acquire annotations from domain experts to identify inconsistencies in summaires and categorize errors.
Outcome: The proposed model outperforms fine-tuned models in biomedical articles and legal bills across specialized domains.
Factcheck-Bench: Fine-Grained Evaluation Benchmark for Automatic Fact-checkers (2024.findings-emnlp)

Copied to clipboard

Challenge: Large language models generate naturally sounding answers over a broad range of human inquiries, but they often generate answers that contradict real-world facts.
Approach: They propose a framework for annotating and evaluating the factuality of large language models . they propose 'factcheck-bench' which provides a multi-stage annotation scheme .
Outcome: The proposed framework outperforms several popular LLM fact-checkers in claim, sentence, and document levels.
Factuality Beyond Coherence: Evaluating LLM Watermarking Methods for Medical Texts (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing benchmarks focus on detection-quality tradeoffs and overlook factual risks.
Approach: They propose a method that assesses factual accuracy and coherence . they use a factor-weighted score to prioritize factual accurate beyond coherency .
Outcome: The proposed method assesses factual accuracy and coherence in medical text . it shows current watermarking methods substantially compromise medical factuality .
Understanding Faithfulness and Reasoning of Large Language Models on Plain Biomedical Summaries (2024.findings-emnlp)

Copied to clipboard

Challenge: Generating plain biomedical summaries with Large Language Models (LLMs) can enhance access to biomedically knowledge.
Approach: They propose a benchmark dataset with expert-annotated Faithfulness and Reasoning on plain biomedical summaries.
Outcome: The proposed dataset shows that LLMs perform poorly in generating faithful biomedical summaries and that abstractiveness and faithfulness are negatively correlated.
Evaluating Factuality in Cross-lingual Summarization (2023.findings-acl)

Copied to clipboard

Challenge: Existing evaluation metrics for monolingual summarization require translation to evaluate the factuality of cross-lingual summmarization.
Approach: They propose to analyze cross-lingual factuality by collecting annotations and generated summaries from models at summary level and sentence level.
Outcome: The proposed dataset shows that over 50% of generated summaries contain factual errors with different characteristics from monolingual summarization.
FENICE: Factuality Evaluation of summarization based on Natural language Inference and Claim Extraction (2024.findings-acl)

Copied to clipboard

Challenge: Recent advances in text summarization have shown remarkable performance, but a significant number of summaries exhibit factual inconsistencies, such as hallucinations.
Approach: They propose a factuality-oriented metric that evaluates text summarization for accuracy . they use a human annotation process to examine the accuracy of automatically generated summaries .
Outcome: The proposed metric sets a new state-of-the-art on AGGREFACT, the de-facto benchmark for factuality evaluation.
Evaluation of LLMs in Medical Text Summarization: The Role of Vocabulary Adaptation in High OOV Settings (2025.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) have been successful in medical text summarization . however, they do not perform fine-grained evaluations under difficult settings .
Approach: They show that large language models show a significant performance drop for data points with high concentration of out-of-vocabulary words or with high novelty.
Outcome: The proposed model shows a significant performance drop for data points with high concentration of out-of-vocabulary words or with high novelty.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations