Papers with NPR

4 papers
ClaimRank: Detecting Check-Worthy Claims in Arabic and English (N18-5)

Copied to clipboard

Challenge: ClaimRank is an online system for detecting check-worthy claims . it can be used to prioritize the claims fact-checkers should consider first .
Approach: ClaimRank is an online system for detecting check-worthy claims . it is originally trained on political debates, but can work for any kind of text . authors propose to make automated fact-checking easier by prioritizing claims based on annotations from reputable fact- checking organizations.
Outcome: ClaimRank is an online system for detecting check-worthy claims . it can mimic the sentence selection strategies of reputable fact-checking organizations .
What’s Mine becomes Yours: Defining, Annotating and Detecting Context-Dependent Paraphrases in News Interview Dialogs (2024.emnlp-main)

Copied to clipboard

Challenge: a dataset of utterance pairs from NPR and CNN is used to classify paraphrases in dialog.
Approach: They propose a dataset annotated for context-dependent paraphrases and develop a training for crowd-workers to classify paraphrase in dialog.
Outcome: The proposed dataset contains 5,581 annotations on 600 utterance pairs.
MediaSum: A Large-scale Media Interview Dataset for Dialogue Summarization (2021.naacl-main)

Copied to clipboard

Challenge: Existing datasets for dialogue summarization are limited to their small sizes and are built from a narrow domain.
Approach: They propose a large-scale media interview dataset consisting of 463.6K transcripts with abstractive summaries.
Outcome: The proposed dataset is larger and contains multi-party conversations from multiple domains.
NewsInterview: a Dataset and a Playground to Evaluate LLMs’ Grounding Gap via Informational Interviews (2025.acl-long)

Copied to clipboard

Challenge: Existing large datasets (1k-10k transcripts) are generated via crowdsourcing and are inherently unnatural.
Approach: They curate a dataset of 40,000 two-person informational interviews from NPR and CNN . they find that LLMs are significantly less likely than human interviewers to use acknowledgements and pivot to higher-level questions.
Outcome: The proposed model is based on 40,000 interviews with journalists and CNN .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations