Challenge: Contextual large language model embeddings are often monolingual, do not scale, and struggle in multilingual settings.
Approach: They propose a hierarchical approach to embed news articles and social media data using Matryoshka embeddings that can determine story similarity at varying levels of granularity based on which subset of dimensions is examined.
Outcome: The proposed model achieves state-of-the-art performance on the SemEval 2022 task 8 dataset.

Similar Papers

Multilingual Clustering of Streaming News (D18-1)

Copied to clipboard

Challenge: a novel method for clustering news across languages is proposed . a key challenge in handling news streams is that they must be generated on the fly .
Approach: They propose a method for clustering news across languages into monolingual and crosslingual clusters . they use real news datasets in multiple languages to find an ever growing number of cluster labels .
Outcome: The proposed method produces state-of-the-art results on real news datasets in German, English and Spanish.
Matryoshka-Adaptor: Unsupervised and Supervised Tuning for Smaller Embedding Dimensions (2024.emnlp-main)

Copied to clipboard

Challenge: Embeddings from Large Language Models (LLMs) have emerged as critical components in information retrieval applications.
Approach: They propose a tuning framework for the customization of LLM embeddings.
Outcome: The proposed framework reduces embedding dimensions while maintaining comparable performance levels.
PolyNarrative: A Multilingual, Multilabel, Multi-domain Dataset for Narrative Extraction from News Articles (2025.acl-long)

Copied to clipboard

Challenge: a new dataset of news articles annotated for narratives provides a framework for narrative detection . recurring narratives can propagate with very high velocity across audiences, languages and countries .
Approach: They propose a multilingual dataset annotated for narratives using two-level taxonomies . they define narrative as a recurring, repetitive, overt or implicit claim that promotes a specific interpretation or viewpoint on an ongoing topic .
Outcome: The proposed dataset will foster research in narrative detection and enable new research directions . the authors identify multiple narratives in the same article, and the results are published online .
A Structured Clustering Approach for Inducing Media Narratives (2026.acl-long)

Copied to clipboard

Challenge: Existing approaches to modeling media narratives miss subtle narrative patterns through coarse-grained analysis or require domain-specific taxonomies that limit scalability.
Approach: They propose a framework for inducing rich narrative schemas by jointly modeling events and characters via structured clustering.
Outcome: The proposed framework produces explainable narrative schemas that align with established framing theory while scaling to large corpora without exhaustive manual annotation.
Multilingual Multifaceted Understanding of Online News in Terms of Genre, Framing, and Persuasion Techniques (2023.acl-long)

Copied to clipboard

Challenge: a new dataset of news articles is presented that covers genre, framing, and persuasion techniques.
Approach: They propose a multilingual multifacet dataset of news articles annotated for genre, framing and persuasion techniques.
Outcome: The proposed dataset contains 1,612 news articles covering recent news on current topics of public interest in six European languages.
Enhancing Event-centric News Cluster Summarization via Data Sharpening and Localization Insights (2025.acl-long)

Copied to clipboard

Challenge: Existing work on text summarization approaches are approaching or exceeding human excellence .
Approach: They propose a framework that optimizes the balance between information volume and entropy in input texts.
Outcome: The proposed framework optimizes information volume and entropy in input texts, achieving notable improvements in localized contexts.
Entity Framing and Role Portrayal in the News (2025.findings-acl)

Copied to clipboard

Challenge: a dataset of news articles containing 22 fine-grained characters is annotated for entity framing and role portrayal . the dataset includes 1,378 recent news articles in five languages focusing on the Ukraine-Russia War and climate change .
Approach: They propose a multilingual and hierarchical corpus annotated for entity framing and role portrayal in news articles.
Outcome: The proposed dataset includes 1,378 recent news articles in five languages focusing on the Ukraine-Russia War and climate change . the authors report evaluation results on state-of-the-art multilingual transformers and hierarchical zero-shot learning using LLMs at the level of a document, paragraph, and sentence .
MassiveSumm: a very large-scale, very multilingual, news summarisation dataset (2021.emnlp-main)

Copied to clipboard

Challenge: Current research in automatic summarisation is expensive to create, posing a challenge for any language.
Approach: They propose to use a large-scale multilingual summarisation dataset with articles in 92 languages and more than 35 writing scripts to generate a multilingual dataset.
Outcome: The proposed method is the largest, most inclusive, existing dataset and one of the largest and most inclusive datasets for any NLP task.
LLMTaxo: Leveraging Large Language Models for Constructing Taxonomy of Factual Claims from Social Media (2025.findings-acl)

Copied to clipboard

Challenge: Social media's global reach and ease of use have transformed how millions of users exchange opinions, news, and factual claims in real-time, making it fertile ground for misinformation.
Approach: They propose a framework that leverages large language models to construct taxonomies of factual claims from social media by generating topics at multiple levels of granularity.
Outcome: The proposed framework produces clear, coherent, and comprehensive taxonomies on three diverse datasets and outperforms other frameworks in most metrics.
Multilingual Fine-Grained News Headline Hallucination Detection (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing models to generate news headlines often suffer from the "hallucination" problem, where the produced headline does not fully align with the source article's content.
Approach: They propose to use a multilingual, fine-grained dataset to detect news headlines in 5 languages using supervised fine-tuning techniques and coarse-to-fine prompting to boost the few-shot detection performance.
Outcome: The proposed methods boost the few-shot hallucination detection performance in terms of the example-F1 metric.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations