Challenge: Existing semantic textual similarity (STS) datasets are not suitable for news similarity detection due to their specificity to a single topic.
Approach: They propose to segment news similarity datasets into topics to improve model training . they propose four different levels of complexity specifically designed for news similarities detection task .
Outcome: The proposed dataset includes seven topics: Crime & Law, Culture & Entertainment, Disasters & Accidents, Economy & Business, Politics / Conflicts, Science & Technology, Sports.

Similar Papers

Query-focused Scenario Construction (D19-1)

Copied to clipboard

Challenge: Stronger neural network models and harder synthetic training settings are important to achieve high performance.
Approach: They propose a query-based system that extracts compatible sets of events from news data . stronger neural network models and harder synthetic training settings are important to achieve high performance .
Outcome: The proposed system outperforms baselines on a human-curated dataset of scenarios about real-world news topics.
Fakeddit: A New Multimodal Benchmark Dataset for Fine-grained Fake News Detection (2020.lrec-1)

Copied to clipboard

Challenge: Prior fake news datasets lack multimodal text and image data, metadata, comment data, and fine-grained classification at the scale and breadth of their datasets.
Approach: They propose to use a multimodal dataset to build a machine learning classification model that uses text and image data to classify fake news.
Outcome: The proposed model is based on a multimodal dataset consisting of over 1 million samples from multiple categories of fake news.
Hidden Biases in Unreliable News Detection Datasets (2021.eacl-main)

Copied to clipboard

Challenge: Recent studies show that automatic unreliable news detection models only use the article itself without resorting to fact-checking mechanisms.
Approach: They propose to use a simple model as a difficulty/bias probe instead of a complex one . they observe a significant drop in accuracy for all models tested in a clean split .
Outcome: The proposed model can achieve good performance by memorizing site-label mapping instead of modeling the real task.
When it Rains, it Pours: Modeling Media Storms and the News Ecosystem (2023.findings-emnlp)

Copied to clipboard

Challenge: Occasionally, an event triggers a media storm, with coverage lasting weeks instead of days.
Approach: They develop a pairwise article similarity model to identify story clusters in news corpora and build a corpus of media storms over a nearly two year period.
Outcome: The proposed model validates theories about storm evolution and topical distribution and supports hypotheses about storm influence on media coverage and intermedia agenda setting.
Different Absorption from the Same Sharing: Sifted Multi-task Learning for Fake News Detection (D19-1)

Copied to clipboard

Challenge: Existing methods for detecting fake news use shared features as complementarity features without selection.
Approach: They propose a sifted multi-task learning method with a selected sharing layer for fake news detection.
Outcome: The proposed method boosts the F1-score by more than 0.87%, 1.31% on two public and widely used competition datasets.
BREAKING! Presenting Fake News Corpus for Automated Fact Checking (P19-2)

Copied to clipboard

Challenge: a new study shows that fake news spreads faster than mainstream articles on the same topic . however, there is no dataset containing compelling fake and questionable news articles .
Approach: They introduce manually verified corpus of compelling fake and questionable news articles on the USA politics . they plan to extend the corpus in the future and use it for automated fake news detection.
Outcome: The proposed model is based on linguistic features and will be extended in the future . it will be used to improve the existing model and improve the tools in the field of fake news detection .
Two Heads Are Better Than One: Improving Fake News Video Detection by Correlating with Neighbors (2023.findings-acl)

Copied to clipboard

Challenge: Existing frameworks for detecting fake news videos are limited . a new approach is proposed to integrate neighborhood information of new videos .
Approach: They propose a framework for automatically detecting fake news videos . it integrates neighborhood relationship of new videos belonging to same event .
Outcome: The proposed framework improves performance of existing detectors and graph aggregation and debunking rectification modules.
A Multi-View Media Profiling Suite: Resources, Evaluation, and Analysis (2026.findings-acl)

Copied to clipboard

Challenge: a large-scale label set for media outlets from Media Bias/Fact Check (MBFC) is lacking in the field.
Approach: They propose to use a large-scale label set to analyze outlets' representations . they also propose to evaluate embedding views and fusion strategies .
Outcome: The proposed method achieves state-of-the-art results on ACL-2020 and establishes strong benchmarks on MBFC-2025.
Multi-News: A Large-Scale Multi-Document Summarization Dataset and Abstractive Hierarchical Model (P19-1)

Copied to clipboard

Challenge: Multi-document summarization (MDS) of news articles has been limited to datasets of a couple of hundred examples.
Approach: They propose a model which integrates a traditional extractive summarization model with a standard SDS model and achieves competitive results on MDS datasets.
Outcome: The proposed model achieves competitive results on large-scale datasets.
End-to-End Segmentation-based News Summarization (2022.findings-acl)

Copied to clipboard

Challenge: Existing summarization systems only provide one genetic summary of the whole article, making it difficult for users to navigate the reading.
Approach: They propose a task of segmenting a news article into multiple sections and generating the corresponding summary to each section.
Outcome: The proposed model outperforms state-of-the-art models on a 27k news article dataset . it can jointly segment a document and produce the summary for each section .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations