Papers with clustering

141 papers
TOPICAL: TOPIC Pages AutomagicaLly (2024.naacl-demo)

Copied to clipboard

Challenge: Topic pages aggregate useful information about an entity or concept into a single concise article.
Approach: They propose a web app that generates topic pages for biomedical entities on demand . they use large language models and retrieval-augmented generation to generate high-quality topics .
Outcome: The proposed method is based on a human evaluation of 150 biomedical topics . it uses large language models and retrieval-augmented generation (RAG)
MaintNet: A Collaborative Open-Source Library for Predictive Maintenance Language Resources (2020.coling-demos)

Copied to clipboard

Challenge: Maintenance record logbooks are an emerging text type in NLP. maintenance record logbook data is often written in non-standard language with many domain specific technical terms, abbreviations, and non-standardized spelling and grammar.
Approach: They propose to create a collaborative open-source library of technical and domain-specific language resources for maintenance record logbooks.
Outcome: The proposed library provides tools to aid in their (pre-)processing and clustering.
Unsupervised Multimodal Clustering for Semantics Discovery in Multimodal Utterances (2024.acl-long)

Copied to clipboard

Challenge: Existing methods for semantics discovery focus on text, video, and audio, failing to leverage the rich multimodal information in the real world.
Approach: They propose a method to construct augmentation views for multimodal data and use them to perform pre-training to establish well-initialized representations for subsequent clustering.
Outcome: The proposed method improves on benchmark multimodal intent and dialogue act datasets by 2-6% over state-of-the-art methods.
Triad-based Neural Network for Coreference Resolution (C18-1)

Copied to clipboard

Challenge: Entity coreference resolution aims to identify mentions that refer to the same entity.
Approach: They propose a triad-based neural network system that generates affinity scores between entity mentions for coreference resolution.
Outcome: The proposed system generates affinity scores between mentions for coreference resolution.
Language Clustering for Multilingual Named Entity Recognition (2021.findings-emnlp)

Copied to clipboard

Challenge: Recent work in multilingual natural language processing has shown progress on tasks such as natural language inference and joint multilingual translation.
Approach: They propose a technique that groups similar languages together by embeddings from a pre-trained masked language model and automatically discovering language clusters in this embeddable space.
Outcome: The proposed technique outperforms baselines on 15 languages in the WikiAnn dataset showing meaningful multilingual transfer for low-resource languages (Swahili and Yoruba).
NLP Tools for Predictive Maintenance Records in MaintNet (2020.aacl-demo)

Copied to clipboard

Challenge: Maintenance logbooks often contain free text fields with domain specific terms, abbreviations, and non-standard spelling . most standard NLP pipelines for pre-processing and annotation are trained on standard contemporary corpora.
Approach: They propose to create an open-source library and data repository for predictive maintenance language datasets and to evaluate the tools available at MaintNet.
Outcome: The proposed tools improve the performance of existing pipelines and improve the quality of the existing ones.
Semantic Diversity for Natural Language Understanding Evaluation in Dialog Systems (2020.coling-industry)

Copied to clipboard

Challenge: a dialog system is used to evaluate NLU models using aggregated metrics on a large number of utterances.
Approach: They propose a method to generate a test set with high semantic diversity for NLU evaluation in dialog systems.
Outcome: The proposed test sets are based on high diversity of utterances from different regions of the utteration embedding space.
Deep Bayesian Natural Language Processing (P19-4)

Copied to clipboard

Challenge: Introduction to deep Bayesian learning for natural language addresses the fundamentals of statistical models and neural networks.
Approach: This tutorial addresses the advances in deep Bayesian learning for natural language . it focuses on advanced Bayessian models and deep models . authors present case studies and domain applications to tackle different issues .
Outcome: This tutorial focuses on advanced Bayesian models and deep models for natural language . case studies and domain applications are presented to tackle different issues in deep Bayessian processing, learning and understanding.
Template-guided Grammatical Error Feedback Comment Generation (2023.eacl-srw)

Copied to clipboard

Challenge: Writing corrective feedback on learner text is widespread in language education, but it can be time-consuming for teachers.
Approach: They propose to use feedback comment generation to generate explanatory notes for learners by categorizing comments and constraining outputs of noisy classes.
Outcome: The proposed scheme can be used to generate feedback comment corpora using a broader scope than existing typologies focused on error correction.
A Closer Look at Clustering Bilingual Comparable Corpora (2024.lrec-main)

Copied to clipboard

Challenge: Existing methods for clustering comparable corpora are not suitable for bilingual corpors.
Approach: They propose new clustering models fully adapted to comparable corpora based on a deep variant of Kmeans . they illustrate their behavior on bilingual collections created from Wikipedia .
Outcome: The proposed models show that they can cluster comparable corpora on bilingual collections . the proposed models are based on a state-of-the-art deep variant of Kmeans .
Exploring Semantic Spaces for Detecting Clustering and Switching in Verbal Fluency (2022.coling-1)

Copied to clipboard

Challenge: Existing evaluations of word/concept representations on verbal fluency tasks rely on human annotations of clusters and switches between sub-categories.
Approach: They analyze word/concept representations in an experimental verbal fluency dataset . they find that ConceptNet embeddings outperforms other semantic representations .
Outcome: The proposed method outperforms other semantic representations by a large margin.
The SUMMA Platform: A Scalable Infrastructure for Multi-lingual Multi-media Monitoring (P18-4)

Copied to clipboard

Challenge: The SUMMA Platform is a highly scalable distributed architecture for monitoring a large number of media broadcasts in parallel, with a lag behind actual broadcast time of at most a few minutes.
Approach: The open-source SUMMA Platform is a highly scalable distributed architecture for monitoring a large number of media broadcasts in parallel . it offers a fully automated media ingestion pipeline capable of recording live broadcasts, detection and transcription of spoken content, translation of all text (original or transcribed) into English, recognition and linking of Named Entities, topic detection, clustering and cross-lingual multi-document summarization of related media items and extraction and storage of factual claims in these news items.
Outcome: The SUMMA Platform is a highly scalable distributed architecture for monitoring a large number of media broadcasts in parallel, with a lag behind actual broadcast time of at most a few minutes.
Improving Word Sense Induction through Adversarial Forgetting of Morphosyntactic Information (2024.starsem-1)

Copied to clipboard

Challenge: Contextualized word representations from pre-trained language models encode more information than is necessary for the identification of word senses and some of this information affect performance negatively in unsupervised settings.
Approach: They propose to use a framework to erase specific information from pre-trained word models and create feature-invariant representations that are invariant to these ‘nuisance features’.
Outcome: The proposed framework erases information from the representations of pre-trained language models, thereby creating feature-invariant representations.
Event Ontology Completion with Hierarchical Structure Evolution Networks (2023.emnlp-main)

Copied to clipboard

Challenge: Existing methods for event detection require predefined schemas, but manual defining is expensive and labor-intensive.
Approach: They propose a task to achieve event clustering, hierarchy expansion and type naming . they propose 'neighbor Contrastive Clustering' module and a Hierarchy-Aware Linking module .
Outcome: The proposed method outperforms baseline methods on three datasets.
New Intent Discovery with Pre-training and Contrastive Learning (2022.acl-long)

Copied to clipboard

Challenge: Existing methods for identifying intents from unlabeled utterances are label-intensive, inefficient, and inaccurate.
Approach: They propose a multi-task strategy to leverage unlabeled data and external labeled data for representation learning.
Outcome: The proposed method outperforms state-of-the-art methods on three intent recognition benchmarks.
Multi-lingual Common Semantic Space Construction via Cluster-consistent Word Embedding (D18-1)

Copied to clipboard

Challenge: a new approach to multilingual word embedding is needed to achieve this goal . a multilingual common semantic space is a language-agnostic semantic continuous space .
Approach: They propose a multilingual common semantic space where words from multiple languages are mapped into a shared space so that resources and knowledge can be shared across languages.
Outcome: The proposed approach achieves 14.6% absolute F-score gain over state-of-the-art methods on cross-lingual direct transfer.
Theoretical Linguistics Rivals Embeddings in Language Clustering for Multilingual Named Entity Recognition (2023.acl-srw)

Copied to clipboard

Challenge: Existing studies have used descriptive typological features and a coarse language family classification as baselines for language clustering.
Approach: They propose two types of language groupings based on morpho-syntactic features in a nominal domain and one based upon a head parameter.
Outcome: The proposed methods outperform state-of-the-art embedding-based models in multilingual named entity recognition (NER) . their results suggest that theoretical linguistics plays a significant role in multi-lingual learning tasks.
QuickGraph: A Rapid Annotation Tool for Knowledge Graph Extraction from Technical Text (2022.acl-demo)

Copied to clipboard

Challenge: Acquiring high-quality annotated corpora for complex multi-task information extraction (MT-IE) is an arduous and costly process for human-annotators.
Approach: They propose a supervised MT-IE annotation tool built with indirect weak supervision and clustering to maximise annotator productivity.
Outcome: The proposed tool is compared with existing tools in the field of MT-IE and aims to increase annotator productivity.
Answer is All You Need: Instruction-following Text Embedding via Answering the Question (2024.acl-long)

Copied to clipboard

Challenge: Existing methods for encoding instruction information fail to be sensitive to clearer criteria like “evaluate similarity based on emotion” . instead, we propose a different approach, which treats the instruction as a “question” about the input text and encodes the expected answers to obtain the representation accordingly.
Approach: They propose a text embedder that captures characteristics of texts specified by user instructions clarifying the similarity criterion.
Outcome: The proposed model improves instruction-following capabilities when applied to large language models and encoder-based LMs.
Identifying Linguistic Areas for Geolocation (D19-55)

Copied to clipboard

Challenge: a recent study shows that social media posts are often given as continuous coordinates . but, the resulting discrete coordinates do not always correspond to existing linguistic areas .
Approach: They propose an algorithm for clustering coordinates and associating them with towns using point-to-city (P2C) they compare accuracy of a state-of-the-art geolocation model with P2C labels to one with regular k-d tree labels.
Outcome: The proposed method improves accuracy at 100 miles, but degrades for finer-grained distinctions . iterative k-d tree-based method can cluster coordinates and associate them with towns .
No, you’re not alone: A better way to find people with similar experiences on Reddit (D19-55)

Copied to clipboard

Challenge: a probabilistic clustering algorithm can help users find posts that discuss experiences similar to their own . a recent study shows that probabilistic Clustering can yield a better performance than baseline clustering methods .
Approach: They propose a probabilistic clustering algorithm that can help Reddit users find posts that discuss experiences similar to their own.
Outcome: The proposed algorithm can find posts that discuss experiences similar to their own . it performs better than baseline clustering methods due to high runtime overhead .
Controllable Clustering with LLM-driven Embeddings (2025.emnlp-industry)

Copied to clipboard

Challenge: Unsupervised text clustering is unlikely to produce groupings that work across use cases . authors present techniques to effectively control text embeddings with minimal human input .
Approach: They propose techniques to control text embeddings with minimal human input . they evaluate clustering performance for datasets with multiple independent labels .
Outcome: The proposed techniques improve clustering for one perspective or use case, but at a tradeoff in performance for another use case.
Focusing Condition: Inference-Time Self-Contrastive Steering Elicits Better Conditional Text Embeddings in LLMs (2026.acl-long)

Copied to clipboard

Challenge: Existing methods for extracting conditional text embeddings from large language models (LLMs) relying on prompts often fails to produce high-quality conditional embeddables, resulting in degradation of quality.
Approach: They propose a plug-and-play method that constructs unconditional general text embeddings and uses them to refine conditional text embeds.
Outcome: The proposed method improves performance of prompt-based methods on clustering, Semantic Textual Similarity, and triplet alignment datasets.
Improving Hierarchical Text Clustering with LLM-guided Multi-view Cluster Representation (2024.emnlp-industry)

Copied to clipboard

Challenge: a multi-stage approach to hierarchical clustering of interaction drivers in contact centers is proposed . silhouette score and human preference score are improved by 36.7% for top-level clusters compared to standard agglomerative clustering .
Approach: They propose a multi-stage approach that introduces different perspectives or views to improve the quality of hierarchical clustering of interaction drivers in a contact center.
Outcome: The proposed approach improves the quality of generated clusters on public datasets with minimal query time compared to the current state-of-the-art approaches.
A Hybrid Supervised-LLM Pipeline for Actionable Suggestion Mining in Unstructured Customer Reviews (2026.eacl-industry)

Copied to clipboard

Challenge: Existing approaches to extract actionable suggestions from customer reviews are often mixed-intent, unstructured text.
Approach: They propose a hybrid pipeline that uses a RoBERTa classifier and a precision–recall surrogate to extract actionable suggestions from customer reviews.
Outcome: The proposed pipeline outperforms prompt-only, rule-based, and classifier-only baselines in extraction accuracy and cluster coherence.
ASK: Aspects and Retrieval based Hybrid Clarification in Task Oriented Dialogue Systems (2025.acl-industry)

Copied to clipboard

Challenge: Ambiguous user queries pose a challenge in task-oriented dialogue systems . Large Language Models (LLMs) rely on the top-k retrieved documents for clarification . traditional approaches lack principled mechanisms to determine when to use broad domain knowledge vs specific retrieved document context for clarification.
Approach: They propose a hybrid approach that dynamically chooses between document-based or aspect-based clarification based on query ambiguity.
Outcome: The proposed approach shows significant improvements over baselines on product troubleshooting and product search datasets.
Claim-Guided Textual Backdoor Attack for Practical Applications (2025.findings-naacl)

Copied to clipboard

Challenge: a novel backdoor attack is based on textual claims to trick models into misbehaving on targeted claims.
Approach: a new backdoor attack is designed to trick models into misbehaving on targeted claims . the code and data will be available at https://github.com/minkyoo9/CGBA .
Outcome: a new backdoor attack exploits the power of textual claims to trick models into misbehaving on claims without affecting their performance on clean data.
DeCLUTR: Deep Contrastive Learning for Unsupervised Textual Representations (2021.acl-long)

Copied to clipboard

Challenge: Sentence embeddings are an important component of many natural language processing systems.
Approach: They propose a self-supervised objective for learning universal sentence embeddings that does not require labelled training data.
Outcome: The proposed approach closes the performance gap between unsupervised and supervised pretraining for universal sentence encoders.
Contrastive Bootstrapping for Label Refinement (2023.acl-short)

Copied to clipboard

Challenge: Existing methods for fine-grained classification categorize texts into coarse-gritty classes, but they are suboptimal in real-world scenarios.
Approach: They propose a lightweight contrastive clustering-based bootstrapping method to iteratively refine the labels of passages.
Outcome: The proposed method outperforms the state-of-the-art methods by a large margin on NYT and 20News datasets.
Leveraging Only the Category Name for Aspect Detection through Prompt-based Constrained Clustering (2022.findings-emnlp)

Copied to clipboard

Challenge: Aspect category detection (ACD) aims to automatically identify user-concerned aspects from online reviews.
Approach: They propose a method that relies on the category name of each aspect and a pretrained language model to generate constraints for clustering.
Outcome: The proposed framework performs better than existing weakly supervised methods on nine benchmark datasets.
Disentangled Learning of Stance and Aspect Topics for Vaccine Attitude Detection in Social Media (2022.naacl-main)

Copied to clipboard

Challenge: Existing approaches to detect vaccine attitudes on social media require abundant annotations and pre-defined aspect categories.
Approach: They propose a semi-supervised approach to detect vaccine attitudes on social media . they use an autoencoding architecture to learn from unlabelled data the topical information of the domain .
Outcome: The proposed model outperforms existing aspect-based models on stance detection and tweet clustering.
Automated Tone Transcription and Clustering with Tone2Vec (2024.findings-emnlp)

Copied to clipboard

Challenge: Lexical tones play a crucial role in Sino-Tibetan languages, but current phonetic fieldwork relies on manual effort.
Approach: They propose a pitch-based similarity representations for tone transcription called Tone2Vec . they propose an open-source package that facilitates automated fieldwork and analysis .
Outcome: Experiments on dialect clustering and variance show that Tone2Vec captures fine-grained tone variation.
Hybrid Inverted Index Is a Robust Accelerator for Dense Retrieval (2023.emnlp-main)

Copied to clipboard

Challenge: Inverted file structure is a common technique for accelerating dense retrieval, but its lossy nature degrades it.
Approach: They propose a hybrid index where embedding clusters and salient terms work collaboratively to accelerate dense retrieval.
Outcome: The proposed method achieves lossless retrieval quality with competitive efficiency across index settings.
GenEOL: Harnessing the Generative Power of LLMs for Training-Free Sentence Embeddings (2025.findings-naacl)

Copied to clipboard

Challenge: Training-free embedding methods focus on optimizing embeddable prompts . previous methods have overlooked the benefits of utilizing generative abilities of LLMs - GenEOL .
Approach: They propose a method that leverages pretrained large language models to embed text . they propose generating diverse transformations of a sentence that preserve its meaning .
Outcome: The proposed method outperforms existing training-free embedding methods by 2.85 points on the sentence semantic text similarity (STS) benchmark.
Dialect Clustering with Character-Based Metrics: in Search of the Boundary of Language and Dialect (2020.lrec-1)

Copied to clipboard

Challenge: 'A language is a dialect with an army and navy' is attributed to sociologist Max Weinrich.
Approach: They propose a universal character-based method for representing sentences so that one can calculate the distance between any two sentence pairs.
Outcome: The proposed method can be used to calculate distance between two sentences by clustering a dialect/sub-language mixed corpus into sub-groups and to partially answer the question of what separates languages from dialects.
An Unsupervised Sentence Embedding Method by Mutual Information Maximization (2020.emnlp-main)

Copied to clipboard

Challenge: Sentence BERT is inefficient for sentence-pair tasks as it needs to evaluate combinatorially many sentence pairs which is very time-consuming.
Approach: They propose a lightweight extension on top of BERT and a self-supervised learning objective to derive meaningful sentence embeddings in an unsupervised manner.
Outcome: The proposed method outperforms baselines on common semantic textual similarity tasks and downstream supervised tasks and achieves performance competitive with supervised methods on various tasks.
Hierarchical Level-Wise News Article Clustering via Multilingual Matryoshka Embeddings (2025.acl-long)

Copied to clipboard

Challenge: Contextual large language model embeddings are often monolingual, do not scale, and struggle in multilingual settings.
Approach: They propose a hierarchical approach to embed news articles and social media data using Matryoshka embeddings that can determine story similarity at varying levels of granularity based on which subset of dimensions is examined.
Outcome: The proposed model achieves state-of-the-art performance on the SemEval 2022 task 8 dataset.
Facts That Matter (D18-1)

Copied to clipboard

Challenge: Existing methods to discover facts from natural language text are based on relation extraction and open information extraction.
Approach: They propose a task of generating a machine-readable representation of the most prominent information in a text document as a set of facts.
Outcome: The proposed system outperforms baselines and text summarizers in a supervised evaluation of salience tasks.
Intent Detection and Discovery from User Logs via Deep Semi-Supervised Contrastive Clustering (2022.naacl-main)

Copied to clipboard

Challenge: Existing approaches to intent detection rely on epoch wise clustering and classification based on labeled and unlabeled data.
Approach: They propose an end-to-end deep contrastive clustering algorithm that jointly updates model parameters and cluster centers via supervised and self-supervised learning.
Outcome: The proposed approach outperforms baselines on five public datasets and human-in-the-loop variant for practical deployment.
Zero-Shot End-to-End Spoken Language Understanding via Cross-Modal Selective Self-Training (2024.eacl-long)

Copied to clipboard

Challenge: End-to-end (E2E) spoken language understanding models are constrained by the cost of collecting speech-semantics pairs.
Approach: They propose a model that learns E2E SLU without speech-semantics pairs . they propose cross-modal selective self-training (CMSST) to address imbalance and noise issues .
Outcome: The proposed model learns E2E SLU without speech-semantics pairs . the proposed model requires the domains of speech-text and text-sensitization to match .
Model Performance-Guided Evaluation Data Selection for Effective Prompt Optimization (2025.findings-acl)

Copied to clipboard

Challenge: Existing prompt engineering methods rely on randomly selected evaluation subsets, leading to suboptimal prompts.
Approach: They propose an iterative evaluation data selection approach for effective prompt optimization using real time model performance.
Outcome: The proposed approach improves effectiveness by 1.6% to 3.1% and stability by 50% to 55.5% on two datasets BIG-bench and LIAR and two models GPT-3.5 and GPT-4o-mini.
MTEB: Massive Text Embedding Benchmark (2023.eacl-main)

Copied to clipboard

Challenge: Existing text embeddings are evaluated on a small set of datasets, not covering their possible applications to other tasks.
Approach: They propose a benchmarking framework that evaluates 8 embedding tasks covering 58 datasets and 112 languages.
Outcome: The proposed model is the most comprehensive benchmark of text embeddings to date.
AdapterSoup: Weight Averaging to Improve Generalization of Pretrained Language Models (2023.findings-eacl)

Copied to clipboard

Challenge: Pretrained language models often need to specialize to specific domains.
Approach: They propose an approach that performs weight-space averaging of adapters trained on different domains.
Outcome: The proposed approach improves performance to new domains without extra training.
FNSCC: Fuzzy Neighborhood-Aware Self-Supervised Contrastive Clustering for Short Text (2025.findings-emnlp)

Copied to clipboard

Challenge: Short texts pose significant challenges for clustering due to semantic sparsity, limited context and fuzzy category boundaries.
Approach: proposed framework incorporates neighborhood information at instance and cluster levels . a cluster-level framework introduces fuzzy neighborhood-aware weighting .
Outcome: The proposed framework outperforms state-of-the-art models on short texts . it excludes neighbors from negative sample set to enhance inter-cluster separability .
A Lightweight Mixture-of-Experts Neural Machine Translation Model with Stage-wise Training Strategy (2024.findings-naacl)

Copied to clipboard

Challenge: Using mixture-of-experts (MoE) to deal with language heterogeneity is a challenge in neural machine translation (NMT).
Approach: They propose a lightweight MoE-based NMT model that is trained via an elaborate stage-wise training strategy.
Outcome: The proposed model achieves stable improvements in translation tasks by introducing fewer extra parameters compared to baseline models.
Efficient Cluster-Based k-Nearest-Neighbor Machine Translation (2022.acl-long)

Copied to clipboard

Challenge: k-Nearest-Neighbor Machine Translation (kNN-MT) is a non-parametric solution for domain adaptation . previous studies have shown that kNN retrieval is at the expense of high latency .
Approach: They propose to use clustering to improve retrieval efficiency by combining a non-parametric MT with an in-domain feature-based retrieval module.
Outcome: The proposed method reduces translation latency by 57% while maintaining the most useful information of the original datastore.
LOGAN: Local Group Bias Detection by Clustering (2020.emnlp-main)

Copied to clipboard

Challenge: a number of machine learning models inherit and amplify the societal biases in data.
Approach: a new bias detection technique based on clustering is proposed to detect local biases in data . authors propose to use LOGAN to analyze local bias in data.
Outcome: The proposed technique detects bias in a local region and allows better analysis of model predictions.
Improved Training of Deep Text Clustering (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing methods for deep clustering optimization with shallow models have limited performance due to poor power of feature learning.
Approach: They propose a general deep clustering optimization method that leverages information feedback to construct generalized labels to optimize the deep model.
Outcome: The proposed method reduces the impact of noise on the clustering process by using correlation relationship between the samples.
Compressing then Matching: An Efficient Pre-training Paradigm for Multimodal Embedding (2026.acl-long)

Copied to clipboard

Challenge: Recent approaches demonstrate that MLLMs can be adapted into competitive embedding models via large-scale contrastive learning.
Approach: They propose a compressed pre-training phase which serves as a warm-up stage for contrastive learning.
Outcome: The proposed model achieves state-of-the-art among MLLMs of comparable size on the MMEB, realizing optimization in both efficiency and effectiveness.
Comprehensive Abstractive Comment Summarization with Dynamic Clustering and Chain of Thought (2024.findings-acl)

Copied to clipboard

Challenge: Recent work on news comment summarization has focused on extractive methods within constraints.
Approach: They propose an enhanced fast clustering algorithm that maintains a dynamic similarity threshold to ensure high density of each comment cluster being built.
Outcome: The proposed method improves the baseline methods and the test suite on real-world news comments.
Learning Invariant Representations of Social Media Users (D19-1)

Copied to clipboard

Challenge: Existing methods for learning to compare social media users fail to generalize to new users or even to previously known users.
Approach: They propose a procedure to learn a mapping from short episodes of user activity to a vector space in which the distance between points captures the similarity of the corresponding users’ invariant features.
Outcome: The proposed procedure can be applied to users not seen at training time and enables efficient comparisons of users in the resulting vector space.
FinMTEB: Finance Massive Text Embedding Benchmark (2025.emnlp-main)

Copied to clipboard

Challenge: Existing text embedding benchmarks for financial domains are inadequately addressing the nuanced requirements of specialized domains like finance.
Approach: They propose a finance-adapted embedding model that outperforms general-purpose models . they also introduce a new model, Fin-E5, which is also open-sourced .
Outcome: The proposed framework outperforms general-purpose models on financial embedding tasks.
Text-Guided Image Clustering (2024.eacl-long)

Copied to clipboard

Challenge: Current image clustering methods neglect the use of generated textual descriptions.
Approach: They propose to use image captioning and visual question-answering to cluster images . they propose a new approach to inject task- or domain knowledge into image clustering .
Outcome: The proposed method outperforms existing methods on eight image clustering datasets.
Quality Assessment of Tabular Data using Large Language Models and Code Generation (2025.emnlp-industry)

Copied to clipboard

Challenge: Data quality is vital for business decisions; poor data quality costs organizations an average of $12.9 million annually.
Approach: They propose a framework that combines statistical inliner detection with LLM-driven rule and code generation.
Outcome: The proposed framework produces semantically valid quality rules and validates them with retrieval-augmented generation (RAG) Extensive evaluations on benchmark datasets confirm the effectiveness of the proposed framework.
Event-Driven News Stream Clustering using Entity-Aware Contextual Embeddings (2021.eacl-main)

Copied to clipboard

Challenge: a novel method for online news stream clustering is proposed . a user can scour the many news sources multiple times a day to find news articles .
Approach: They propose a method for online news stream clustering that is a variant of the streaming K-means algorithm.
Outcome: The proposed model achieves state-of-the-art on a standard stream clustering dataset of English documents.
Contextualized Embeddings for Enriching Linguistic Analyses on Politeness (2020.coling-main)

Copied to clipboard

Challenge: Current word embeddings in natural language processing do capture context and thus can be leveraged to enrich linguistic analyses.
Approach: They propose a model which leverages pre-trained BERT to cluster contextualized representations of a word based on context in which it appears and labels of items it occurs in.
Outcome: The proposed model can detect interpretable, finer-grained context patterns associated with (im)polite language.
Clustering-based Inference for Biomedical Entity Linking (2021.naacl-main)

Copied to clipboard

Challenge: Existing approaches to linking entities ignore relationships between entities in biomedical knowledge bases.
Approach: They propose a model which can link mentions of unseen entities using learned representations of entities.
Outcome: The proposed model improves on the largest publicly available biomedical dataset by 3.0 points of accuracy and 2.3 points of reliability.
Intent Contrastive Learning Based on Multi-view Augmentation for Sequential Recommendation (2025.coling-main)

Copied to clipboard

Challenge: Existing work on intent-related models fails to capture long-term dependencies in user behavior and fails to effectively utilize item relevance.
Approach: They propose a sequential recommendation framework that combine temporal variability with position encoding that has extrapolation properties to encode sequences, thereby expanding the model’s view of user behavior.
Outcome: The proposed model improves on three real datasets by 0.8% to 14.7% compared to baselines.
Analyzing Encoded Concepts in Transformer Language Models (2022.naacl-main)

Copied to clipboard

Challenge: a new framework to analyze how latent concepts are encoded in representations learned in pre-trained lan-guage models is proposed . conceptX uses clustering to discover the encoded concepts and align them with a large set of human-defined concepts.
Approach: They propose a framework to analyze how latent concepts are encoded in representations learned within pre-trained lan-guage models.
Outcome: The proposed framework explains encoded concepts by aligning with human-defined concepts.
EnsLM: Ensemble Language Model for Data Diversity by Semantic Clustering (2021.acl-long)

Copied to clipboard

Challenge: Existing studies have shown that data diversity affects the performance of LMs if we train a single LM over the entire dataset.
Approach: They propose an autoencoding topic model with a mixture prior to perform clustering for the data.
Outcome: The proposed model can learn knowledge from different samples while extracting cluster-specific features.
Text is All You Need: LLM-enhanced Incremental Social Event Detection (2025.acl-long)

Copied to clipboard

Challenge: Existing state-of-the-art (SOTA) SED models rely on graph neural networks (GNNs) Existing SED frameworks rely heavily on GNNs, which require complex graph construction and time-consuming training processes.
Approach: They propose a framework that leverages the rich background knowledge of large language models to formalize and disambiguate short texts by completing abbreviations and summarizing informal expressions.
Outcome: The proposed framework outperforms existing models on two challenging real-world datasets.
Co-Evolving LLMs and Embedding Models via Density-Guided Preference Optimization for Text Clustering (2025.emnlp-main)

Copied to clipboard

Challenge: Existing methods for text clustering use static pseudo-oracles, i.e., unidirectionally querying them for similarity assessment or data augmentation.
Approach: They propose a training framework that enables bidirectional refinement between LLMs and embedding models by using task-aware prompts to guide the LLM in generating interpretations for the input texts.
Outcome: Experiments on 14 benchmark datasets across 5 tasks demonstrate the effectiveness of the proposed training framework.
X-Class: Text Classification with Extremely Weak Supervision (2021.naacl-main)

Copied to clipboard

Challenge: Weak supervision is a problem in text classification, but it requires corpusspecific knowledge.
Approach: They propose a framework for extremely weak supervision that can be used to train a text classifier.
Outcome: The proposed framework outperforms seed-driven weakly supervised methods on 7 benchmark datasets.
TAGCOS: Task-agnostic Gradient Clustered Coreset Selection for Instruction Tuning Data (2025.findings-naacl)

Copied to clipboard

Challenge: Existing methods for instruction tuning are limited due to the increasing volume of instruction datasets and the increased computational costs.
Approach: They propose to extract a small and highly informative subset of training samples from a large dataset that achieves comparable performance to the full dataset.
Outcome: The proposed algorithm outperforms other unsupervised methods and achieves comparable performance to the full dataset.
Combining Information-Weighted Sequence Alignment and Sound Correspondence Models for Improved Cognate Detection (C18-1)

Copied to clipboard

Challenge: a new approach to cognate detection is proposed to capture the remaining similarities between cognate word forms after thousands of years of divergence.
Approach: They propose a method which uses information weighting and sound correspondence modeling to improve cognate detection.
Outcome: The proposed approach improves on the measure of form similarity and distance-based cognate clustering.
A Probabilistic Generative Model for Typographical Analysis of Early Modern Printing (2020.acl-main)

Copied to clipboard

Challenge: Scholars often need to go beyond textual analysis for establishing provenance of historical documents.
Approach: They propose a deep and interpretable probabilistic generative model to analyze glyph shapes in printed Early Modern documents by generating a latent vector responsible for inking variations, jitter, noise and other unforeseen phenomena.
Outcome: The proposed model outperforms interpretable clustering baselines and overly-flexible deep generative models on the task of completely unsupervised discovery of typefaces in mixed-fonts documents.
Actively Supervised Clustering for Open Relation Extraction (2023.acl-long)

Copied to clipboard

Challenge: Existing methods for Open Relation Extraction (OpenRE) use a two-stage pipeline, which learns relation representations and assignments in the first stage, then manually labels relation for each cluster.
Approach: They propose a method that performs relation learning and relation labeling simultaneously without a significant increase in human effort.
Outcome: The proposed method improves existing SOTA methods by 13.8% and 10.6% on two datasets.
Self-Supervised Neural Topic Modeling (2021.findings-emnlp)

Copied to clipboard

Challenge: Topic models are useful tools for analyzing and interpreting the main underlying themes of large corpora of text.
Approach: They propose a self-supervised neural topic model that learns a topic representation jointly from three co-occurring words and a document that the triple originates from.
Outcome: The proposed model outperforms existing topic models in coherence metrics and document clustering accuracy.
Modeling Frames in Argumentation (D19-1)

Copied to clipboard

Challenge: In argumentation, framing is used to emphasize a specific aspect of a topic while concealing others.
Approach: They propose an unsupervised method for framing arguments into non-overlapping frames . authors propose a corpus of 12, 326 debate-portal arguments organized along the frames of debates' topics .
Outcome: The proposed method outperforms baselines on the argumentation task by 0.28 points.
Transformer-based Causal Language Models Perform Clustering (2025.findings-naacl)

Copied to clipboard

Challenge: Recent studies have shown great improvements in instruction-following capability through additional training for instruction- following tasks.
Approach: They propose to use a Transformer-based causal language model to study instruction-following capabilities.
Outcome: The proposed model learns task-specific information by clustering data within its hidden space, with this clustering process evolving dynamically during learning.
Towards a More Generalized Approach in Open Relation Extraction (2025.acl-long)

Copied to clipboard

Challenge: Existing OpenRE methods assume unlabeled data is a mixture of known and novel instances.
Approach: They propose a generalized OpenRE setting that considers unlabeled data as a mixture of known and novel instances.
Outcome: The proposed framework outperforms baselines in relation classification and clustering on three benchmark datasets.
On the Relationship Between RNN Hidden-State Vectors and Semantic Structures (2024.findings-acl)

Copied to clipboard

Challenge: Using hidden-state vectors of recurrent neural networks (RNNs) we examine the assumption that hidden- state vectors tend to form clusters of semantically similar vectors, which we dub the clustering hypothesis.
Approach: They propose to use recurrent neural networks (RNNs) that model processes with internal states to test their hypothesis.
Outcome: The proposed model is based on a set of RNNs that were trained to recognize regular languages and a context-free language.
D-CALM: A Dynamic Clustering-based Active Learning Approach for Mitigating Bias (2023.findings-acl)

Copied to clipboard

Challenge: Infusing clustering with active learning with AL can overcome the bias issue of both AL and traditional annotation methods while exploiting AL’s annotation efficiency.
Approach: They propose an algorithm that dynamically adjusts clustering and annotation efforts in response to an estimated classifier error-rate.
Outcome: The proposed algorithm outperforms baseline AL approaches with pretrained transformers and traditional Support Vector Machines on eight datasets for emotion, hatespeech, dialog act, and book type detection tasks.
Exploring Alignment in Shared Cross-lingual Spaces (2024.acl-long)

Copied to clipboard

Challenge: a new study examines the degree of alignment between languages in multilingual embeddings . cross-lingual embeds are designed to encode linguistic concepts that bridge equivalent semantic meaning . a comprehensive approach is needed to address these questions.
Approach: They employ clustering to uncover latent concepts within multilingual models . they introduce two metrics to quantify alignment and overlap of these concepts .
Outcome: The proposed model can capture linguistic nuances across languages, but is not language-agnostic? the proposed model is able to capture nuances in multiple languages, the authors say.
Zero-shot Script Parsing (2022.coling-1)

Copied to clipboard

Challenge: Existing resources cover only a small number of tasks, limiting its practical usefulness.
Approach: They propose a zero-shot learning approach to script parsing which enables us to acquire script knowledge without domain-specific annotations.
Outcome: The proposed model outperforms a previous model with scenario-specific supervision and achieves 68.1/74.4 average F1 for event / participant parsing.
An Active Learning Framework for Inclusive Generation by Large Language Models (2025.coling-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) exhibit bias toward underrepresented groups, despite advances in active learning.
Approach: They propose a clustering-based active learning framework enhanced with knowledge distillation that transforms the intermediate outputs of the learner model to yield more representative models without prior knowledge of underlying data distribution.
Outcome: The proposed framework improves performance across data subgroups and lexical diversity, underscoring the model’s resilience to skewness in available data.
Comparison of Diverse Decoding Methods from Conditional Language Models (P19-1)

Copied to clipboard

Challenge: Conditional language models can generate a diverse set of outputs, but for open-ended tasks, beam search is ill-suited to generating a set of diverse sequences.
Approach: They propose a method where we over-sample candidates and use clustering to remove similar sequences to achieve high diversity without sacrificing quality.
Outcome: The proposed method over-samples candidates and removes similar sequences to achieve high diversity without sacrificing quality.
Autoencoding Keyword Correlation Graph for Document Clustering (2020.acl-main)

Copied to clipboard

Challenge: Existing representation learning models do not capture the intra-sentential and inter-sententential features of long-text.
Approach: They propose a graph-based representation for document clustering that builds a Graph Autoencoder on a Keyword Correlation Graph.
Outcome: The proposed graph autoencoder can achieve better clustering performance than existing features.
Does Vision-and-Language Pretraining Improve Lexical Grounding? (2021.findings-emnlp)

Copied to clipboard

Challenge: Large pretrained language models (LMs) have been criticized for lack of grounding, i.e., connecting words to their meanings in the physical world.
Approach: They compare vision-and-language (VL) models trained jointly on text and image or video data to find out how they compare to text-only counterparts.
Outcome: The proposed model outperforms the text-only variants on a commonsense question answering task.
CoMave: Contrastive Pre-training with Multi-scale Masking for Attribute Value Extraction (2023.findings-acl)

Copied to clipboard

Challenge: Existing methods to extract product features from unstructured text still suffer from problems . e-commerce platforms are focusing on multi-scale values, which can be confusing .
Approach: They propose a pre-training technique to automatically obtain attribute value pairs from product descriptions to aid e-commerce.
Outcome: The proposed method improves on the existing token-level masking strategy and achieves state-of-the-art on four benchmarks.
A Data-Driven Method for Analyzing and Quantifying Lyrics-Dance Motion Relationships (2025.naacl-long)

Copied to clipboard

Challenge: Existing studies have not explored the relationships between lyrics and dance motions . previous studies focused on synthesizing or retrieving dance motion from lyrics .
Approach: They propose a method to detect parts of songs where meaningful relationships exist . they use clustering to transform lyrics and dance motions into symbols .
Outcome: The proposed method outperforms existing methods on prose and non-dance dance motions.
Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks (D19-1)

Copied to clipboard

Challenge: Existing methods for finding similar sentences require multiple inferences . a modern GPU requires 65 hours to find the most similar pair in 10,000 sentences .
Approach: They propose a modification of the pretrained BERT network that uses siamese and triplet networks to derive semantically meaningful sentence embeddings.
Outcome: The proposed method outperforms existing methods on sentence-pair regression tasks.
Strategic Demonstration Selection for Improved Fairness in LLM In-Context Learning (2024.emnlp-main)

Copied to clipboard

Challenge: Recent studies highlight the effectiveness of using in-context learning (ICL) to steer large language models in processing tabular data.
Approach: They propose a method that uses clustering and evolutionary strategies to curate a representative sample set from training data.
Outcome: The proposed method significantly improves fairness across various metrics, showing its efficacy in real-world scenarios.
Actively Learn from LLMs with Uncertainty Propagation for Generalized Category Discovery (2024.naacl-long)

Copied to clipboard

Challenge: Generalized category discovery (GCD) is a crucial task in open-world computing, where new categories frequently emerge, necessitating models that can adapt and learn continually.
Approach: They propose to integrate the feedback from LLMs into an active learning paradigm to simplify the labeling task and minimize the spread of inaccurate feedback.
Outcome: The proposed approach significantly improves baseline models at a nominal average cost.
Corpus-based Open-Domain Event Type Induction (2021.emnlp-main)

Copied to clipboard

Challenge: Existing event extraction methods require predefined event types and their annotations to learn event extractors.
Approach: They propose to represent each event type as a cluster of predicate sense, object head> pairs.
Outcome: The proposed method can discover salient and high-quality event types on three datasets from different domains.
Pseudo-Label Enhanced Prototypical Contrastive Learning for Uniformed Intent Discovery (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing methods focus on transferring in-domain (IND) prior knowledge to out-of-domain data through pre-training and clustering.
Approach: They propose a Pseudo-Label enhanced Prototypical Contrastive Learning model for uniformed intent discovery that integrates supervised and pseudo signals from IND and OOD data.
Outcome: The proposed method has been proven effective in two different settings of discovering new intents.
Interpretable Semantic Gradients in SSD: A PCA Sweep Approach and a Case Study on AI Discourse (2026.findings-acl)

Copied to clipboard

Challenge: Supervised Semantic Differential (SSD) is a mixed quantitative–interpretive method that models how text meaning varies with continuous individual-difference variables . currently no systematic method exists for choosing the number of retained components, introducing avoidable researcher degrees of freedom in the analysis pipeline.
Approach: They propose a PCA sweep procedure that treats dimensionality selection as a joint criterion over representation capacity, gradient interpretability, and stability across nearby values of K.
Outcome: The proposed method is based on a corpus of short posts about artificial intelligence written by Prolific participants who also completed Admiration and Rivalry narcissism scales.
Capturing Regional Variation with Distributed Place Representations and Geographic Retrofitting (D18-1)

Copied to clipboard

Challenge: Dialects are one of the main drivers of language variation, a major challenge for natural language processing tools.
Approach: They use a corpus of 16.8M anonymous online posts to learn continuous document representations of cities.
Outcome: The proposed method matches dialect areas at different granularities against an existing dialect map.
TR-MTEB: A Comprehensive Benchmark and Embedding Model Suite for Turkish Sentence Representations (2025.findings-emnlp)

Copied to clipboard

Challenge: TR-MTEB is the first large-scale, task-diverse benchmark for sentence embedding models for Turkish.
Approach: a new benchmark evaluates sentence embedding models for Turkish . TR-MTEB covers six core tasks and 26 high-quality datasets .
Outcome: The TR-MTEB benchmark covers six core tasks and includes 26 high-quality datasets . the models achieve competitive performance across most tasks and significantly improve on baseline models.
Multilingual Clustering of Streaming News (D18-1)

Copied to clipboard

Challenge: a novel method for clustering news across languages is proposed . a key challenge in handling news streams is that they must be generated on the fly .
Approach: They propose a method for clustering news across languages into monolingual and crosslingual clusters . they use real news datasets in multiple languages to find an ever growing number of cluster labels .
Outcome: The proposed method produces state-of-the-art results on real news datasets in German, English and Spanish.
GALA: Geometric Data Selection with Strategic Prospecting for Large Language Model Self-training (2026.findings-acl)

Copied to clipboard

Challenge: Existing approaches to self-training are based on reject sampling and lack quality reasoning paths.
Approach: They propose a framework for self-training using a generate-and-filter paradigm . they propose to identify diverse and informative samples from redundant data and exploit them more strategically.
Outcome: The proposed framework exploits informative samples from redundant data and improves reasoning trajectory prospecting.
Coarse-to-Fine Pre-training for Named Entity Recognition (2020.emnlp-main)

Copied to clipboard

Challenge: Named Entity Recognition (NER) is a task of discovering information entities and identifying their corresponding categories.
Approach: They propose a NER-specific framework to inject coarse-to-fine named entity knowledge into pre-trained models by using a remote supervision strategy.
Outcome: The proposed framework achieves significant improvements against several pre-trained base-lines, demonstrating its effectiveness in label-few and low-resource scenarios.
Cluster & Tune: Boost Cold Start Performance in Text Classification (2022.acl-long)

Copied to clipboard

Challenge: Existing methods to fine-tune pre-trained models for text classification are poor in practice.
Approach: They propose to add an intermediate unsupervised classification task between pre-training and fine-tuning phases to boost performance of pre-trained models.
Outcome: The proposed method improves performance on topical classification tasks when labeled data is scarce.
Crowdsourcing Regional Variation Data and Automatic Geolocalisation of Speakers of European French (L18-1)

Copied to clipboard

Challenge: a crowdsourcing platform is used to collect linguistic data and document language use, with a focus on regional variation in European French.
Approach: They propose a crowdsourcing platform to collect linguistic data and document language use with a special focus on regional variation in European French.
Outcome: The proposed platform collects linguistic data and documents language use with a special focus on regional variation in European French.
Multilingual and Explainable Text Detoxification with Parallel Corpora (2025.coling-main)

Copied to clipboard

Challenge: Existing approaches to manage toxic speech on social platforms are limited . however, there is a need for more proactive moderation of abusive speech.
Approach: They extend parallel text detoxification corpus to new languages to test the approach . they propose a method that combines toxic and non-toxic sentences into a more neutral form .
Outcome: The proposed method integrates the descriptive features of toxic and non-toxic sentences into a more neutral or non- toxic form.
Creating Large-Scale Multilingual Cognate Tables (L18-1)

Copied to clipboard

Challenge: Low-resource languages often suffer from a lack of high-coverage lexical resources.
Approach: They propose a method to generate cognate tables by clustering words from existing lexical resources.
Outcome: The proposed method outperforms baselines on the Romance and Turkic language families.
Enhancing Transformation from Natural Language to Signal Temporal Logic Using LLMs with Diverse External Knowledge (2025.findings-acl)

Copied to clipboard

Challenge: Temporal Logic (STL) is a formal specification tool for cyber-physical systems . but it is difficult to transform ambiguous and complex data into STL, a paper argues .
Approach: They propose a NL-STL dataset with 16,000 samples enriched with diverse patterns . they propose KGST framework to transform natural language into STL using a generate-then-refine process .
Outcome: The proposed dataset outperforms baseline models in diversity and accuracy . the proposed dataset contains 16,000 samples enriched with diverse patterns .
Disentangling language change: sparse autoencoders quantify the semantic evolution of indigeneity in French (2025.naacl-long)

Copied to clipboard

Challenge: Existing methods to measure semantic change with contextual word embeddings (CWEs) are not suitable for highly imbalanced datasets and pose challenges for interpretation.
Approach: They propose an interpretable, feature-level approach to analyzing language change using k-sparse autoencoders to trace the semantic evolution of the term "indigène(s)" between 1825 and 1950.
Outcome: The proposed approach can learn interpretable features from over 210,000 CWEs generated using sentences from the French National Library.
DeTiME: Diffusion-Enhanced Topic Modeling using Encoder-decoder based LLM (2023.findings-emnlp)

Copied to clipboard

Challenge: Neural Topic Models and Large Language Models (LLMs) primarily use contextual embeddings from LLMs, which are not optimal for clustering or topic generation.
Approach: They propose a framework that leverages Encoder-Decoders to generate highly clusterable embeddings that could generate topics that exhibit enhanced clusterability and enhanced semantic coherence compared to existing methods.
Outcome: The proposed framework is efficient to train and exhibits high adaptability, demonstrating its potential for a wide array of applications.
MIST: Mutual Information Maximization for Short Text Clustering (2024.acl-long)

Copied to clipboard

Challenge: Existing methods for clustering short texts are inadequate due to the limited amount of information provided by each text sample.
Approach: They propose a Mutual Information Maximization Framework for Short Text Clustering which maximizes mutual information between representations on sequence and token levels.
Outcome: The proposed framework outperforms the state-of-the-art method in terms of Accuracy or Normalized Mutual Information in most cases.
Improving Clustering with Positive Pairs Generated from LLM-Driven Labels (2025.emnlp-main)

Copied to clipboard

Challenge: Existing unsupervised clustering methods lack label knowledge, resulting in suboptimal performance.
Approach: They propose to use LLM-driven labels to generate positive pairs from embedded data and an embedder to obviate the need for negative pairs.
Outcome: The proposed framework surpasses state-of-the-art benchmarks on a range of datasets and generates interpretable labels for improved understanding of clustering results.
FaMTEB: Massive Text Embedding Benchmark in Persian Language (2025.findings-emnlp)

Copied to clipboard

Challenge: a comprehensive benchmark for Persian text embeddings is built upon the Massive Text Embedding Benchmark (MTEB) 63 datasets are included in the benchmark, including a novel task of summary retrieval.
Approach: They propose a benchmark for Persian (Farsi) text embeddings built upon the Massive Text Embedding Benchmark.
Outcome: The proposed framework includes 63 datasets spanning seven different tasks . the evaluation datasets were rigorously evaluated by humans and automated systems .
ChatMap: Mining Human Thought Processes for Customer Service Chatbots via Multi-Agent Collaboration (2025.findings-acl)

Copied to clipboard

Challenge: Existing methods for enhancing dialogue performance rely on summarizing behavior . e-commerce chatbots need to align their dialogue strategies with human behavior to achieve coherent, human-like conversations with customers.
Approach: They propose a method to extract core patterns from dialogue data and integrate them into models by mining service thought processes using a multi-agent aPproach.
Outcome: The proposed method outperforms manual methods and outperfies baselines on Taobao in China.
Large Language Models Meet Open-World Intent Discovery and Recognition: An Evaluation of ChatGPT (2023.emnlp-main)

Copied to clipboard

Challenge: Existing methods to fine-tune discriminative models address these challenges by focusing on in-domain intents.
Approach: They evaluate ChatGPT on OOD intent discovery and generalized intent discovery tasks . they outline the strengths and weaknesses of ChatGPt and outline their results .
Outcome: The proposed task aims to extend a closed intent classifier to open-world intent sets.
Semi-Supervised Clustering for Short Answer Scoring (L18-1)

Copied to clipboard

Challenge: Existing approaches to SAS use unsupervised clustering and have teachers label some items after clustering.
Approach: They propose to use semi-supervised clustering to provide structured groups of answers in addition to a score.
Outcome: The proposed method improves clustering performance from 0.504 kappa for unsupervised clustering to 0.566 kppa.
Semantic Geometry of Sentence Embeddings (2025.findings-emnlp)

Copied to clipboard

Challenge: Sentence embeddings are central to natural language processing, but their internal features are not interpretable and users lack fine-grained control for downstream tasks.
Approach: They propose a formal framework to characterize the organization of features in sentence embeddings . they show how they can be composed to capture richer semantic structures .
Outcome: The proposed method can be used to capture richer semantic structures.
𝜙-Decoding: Adaptive Foresight Sampling for Balanced Inference-Time Exploration and Exploitation (2025.acl-long)

Copied to clipboard

Challenge: Existing inference-time optimization strategies address the shortsightedness of auto-regressive generation, but the vast search space leads to excessive exploration and insufficient exploitation.
Approach: They propose a decoding strategy that approximates two distributions via foresight and clustering to provide an efficient estimation of step value.
Outcome: The proposed decoding strategy outperforms strong baselines in performance and efficiency.
MAKED: Multi-lingual Automatic Keyword Extraction Dataset (2022.lrec-1)

Copied to clipboard

Challenge: a large dataset of news articles spanning 20 languages is lacking for keyword extraction.
Approach: They propose a large-scale multi-lingual keyword extraction dataset for 11 of 20 languages . authors believe it will help advance the field of automatic keyword extraction .
Outcome: The proposed dataset is the first for 11 of 20 languages and is based on 540K+ news articles from the BBC News network.
The Paradigm Discovery Problem (2020.acl-main)

Copied to clipboard

Challenge: a paradigm discovery problem is a task of learning an inflectional morphological system from unannotated sentences.
Approach: They formalize the paradigm discovery problem and develop evaluation metrics for judging systems . they use word embeddings and string similarity to cluster forms by cell and by paradigm .
Outcome: The proposed system suggests clustering by cell across different inflection classes is the most pressing challenge for future work.
When Phrases Meet Probabilities: Enabling Open Relation Extraction with Cooperating Large Language Models (2024.acl-long)

Copied to clipboard

Challenge: Existing clustering-based open relation extraction methods use pre-trained language models . embeddings from language models are high-dimensional and anisotropic, so there is a gap .
Approach: They propose a framework that makes two LLMs work collaboratively to achieve clustering.
Outcome: The proposed framework outperforms existing methods by 1.4%3.13% on different datasets.
STSPL-SSC: Semi-Supervised Few-Shot Short Text Clustering with Semantic text similarity Optimized Pseudo-Labels (2024.findings-acl)

Copied to clipboard

Challenge: Existing methods for obtaining task-specific labels require prior knowledge of clustering categories and uncontrollable clustering centers.
Approach: They propose a framework for supervised clustering using a discrete process and a robust Contrastive Learning module.
Outcome: The proposed framework outperforms state-of-the-art models on a real-world dataset with just one label per class . the proposed framework is based on k-means clustering and a robust Contrastive Learning module .
Robust Text Classification: Analyzing Prototype-Based Networks (2024.findings-emnlp)

Copied to clipboard

Challenge: Language models exhibit a drop in performance on noisy data, which can cause classifiers to incorrectly change their predictions.
Approach: They propose to use Prototype-Based Networks to classify examples based on their similarity to prototypical examples of a class (prototypes) they show that PBNs offer more robustness under both targeted and static adversarial attacks.
Outcome: The proposed model is robust to noise and targets both targeted and static attacks.
Automatic Construction of an Annotated Corpus with Implicit Aspects (2022.lrec-1)

Copied to clipboard

Challenge: Aspect-based sentiment analysis (ABSA) is a task that involves classifying aspects of products or services described in user reviews.
Approach: They propose a method for constructing a corpus that is automatically annotated with implicit aspects by combining explicit and unlabeled sentences.
Outcome: The proposed method achieves a maximum accuracy of 82% on mobile phone reviews.
Improving Faithfulness of Large Language Models in Summarization via Sliding Generation and Self-Consistency (2024.lrec-main)

Copied to clipboard

Challenge: Abstractive summarization models (LLMs) have demonstrated impressive performance in various tasks, but they are still suffering from factual inconsistency problem called hallucination.
Approach: They propose to improve the faithfulness of large language models by impelling them to process the entire article more fairly and faithfully.
Outcome: The proposed strategy improves the faithfulness of large language models in summarization while maintaining their fluency and informativeness.
Improving Multi-view Document Clustering: Leveraging Multi-structure Processor and Hybrid Ensemble Clustering Module (2024.lrec-main)

Copied to clipboard

Challenge: Experimental results show that DMsECN outperforms existing models for document clustering .
Approach: They propose a multi-view document clustering model with a processor and hybrid module . they demonstrate that DMsECN outperforms existing models by creating a consensus structure from multiple clustering structures.
Outcome: The proposed model outperforms existing models on four multi-view document clustering datasets.
QA-Align: Representing Cross-Text Content Overlap by Aligning Question-Answer Propositions (2021.emnlp-main)

Copied to clipboard

Challenge: Existing approaches to consolidate textual inputs are difficult to implement . a recent study aims to capture content overlap by combining multiple textual elements .
Approach: They propose to align predicate-argument relations across texts to represent content overlap . their setting exploits QA-SRL, utilizing question-answer pairs to capture predicates .
Outcome: The proposed task captures content overlap beyond lexical similarity and complements cross-document coreference with proposition-level links, offering potential use for downstream tasks.
Cluster-Norm for Unsupervised Probing of Knowledge (2024.emnlp-main)

Copied to clipboard

Challenge: Empirical evidence suggests that simulated internal beliefs or knowledge can be extracted from language models but such methods require labels, which in some domains may not be readily provided due to human biases or because humans simply do not know the correct label.
Approach: They propose a method to minimize the impact of unrelated features in activation space by clustering and normalizing activations of contrast pairs before applying unsupervised probing techniques.
Outcome: Empirical evidence suggests that simulated internal beliefs or knowledge can be extracted from language model activations without human labels.
Enhancing Event-centric News Cluster Summarization via Data Sharpening and Localization Insights (2025.acl-long)

Copied to clipboard

Challenge: Existing work on text summarization approaches are approaching or exceeding human excellence .
Approach: They propose a framework that optimizes the balance between information volume and entropy in input texts.
Outcome: The proposed framework optimizes information volume and entropy in input texts, achieving notable improvements in localized contexts.
SPILL: Domain-Adaptive Intent Clustering based on Selection and Pooling with Large Language Models (2025.findings-acl)

Copied to clipboard

Challenge: Existing methods for intent clustering rely on labeled examples or unsupervised fine-tuning to optimize results for each new dataset.
Approach: They propose a method that uses an embedder to derive an embedding for each utterance and then pool them with the seed to improve the embeddable results.
Outcome: The proposed method outperforms embedding methods and is comparable to state-of-the-art methods.
Adversarial Robustness for Large Language NER models using Disentanglement and Word Attributions (2023.findings-emnlp)

Copied to clipboard

Challenge: Named Entity Recognition (NER) tasks are becoming more challenging due to the introduction of complex tagsets, which often leads to the failure of existing NER systems in accurately recognizing these entities.
Approach: They propose a novel attack which relies on disentanglement and word attribution techniques to learn an embedding and identifying important words across both components.
Outcome: The proposed approach improves the F1 score over the original LLM model by 8% and 18% on CoNLL-2003 and Ontonotes 5.0 datasets respectively.
Open-world Multi-label Text Classification with Extremely Weak Supervision (2024.emnlp-main)

Copied to clipboard

Challenge: Similar single-label XWS settings cannot be easily adapted for multi-l label classification.
Approach: They propose a novel method for open-world multi-label text classification under extremely weak supervision where the user provides a brief description without any labels or ground-truth label space.
Outcome: The proposed method exhibits a remarkable increase in ground-truth label space coverage on various datasets.
Comparative Study of Sentence Embeddings for Contextual Paraphrasing (2020.lrec-1)

Copied to clipboard

Challenge: Paraphrasing is an important aspect of natural-language generation that can produce more variety in the way specific content is presented.
Approach: They propose to use contextual paraphrasing to capture the meaning of a sentence while performing dialogue act clustering.
Outcome: The proposed task combines paraphrases with dialogue act clustering to capture such contextual paraphrasing.
ClusterLLM: Large Language Models as a Guide for Text Clustering (2023.emnlp-main)

Copied to clipboard

Challenge: Extensive experiments on 14 datasets show that ClusterLLM consistently improves clustering quality, at an average cost of $0.6 per dataset.
Approach: They propose a text clustering framework that leverages feedback from an instruction-tuned large language model, such as ChatGPT.
Outcome: Extensive experiments on 14 datasets show that ClusterLLM consistently improves clustering quality, at an average cost of $0.6 per dataset.
Intent Discovery with Frame-guided Semantic Regularization and Augmentation (2023.findings-acl)

Copied to clipboard

Challenge: Existing intent discovery methods focus on transferring prior knowledge of known intents to unknown ones.
Approach: They propose to use frame knowledge as conceptual semantic guidance to bridge the gap between known intents representation learning and unknown intents clustering.
Outcome: The proposed method outperforms solid baselines on two benchmark datasets.
LLM-Guided Semantic-Aware Clustering for Topic Modeling (2025.acl-long)

Copied to clipboard

Challenge: Experimental results show that topic modeling is competitive compared to closed-source methods.
Approach: They propose a semi-supervised topic modeling method that combines LLMs with clustering to improve topic generation and distribution.
Outcome: The proposed method outperforms state-of-the-art methods that utilize GPT-4 on topic alignment and exhibits competitive performance compared to Neural Topic Models on topic quality.
ZeroDL: Zero-shot Distribution Learning for Text Clustering via Large Language Models (2025.findings-acl)

Copied to clipboard

Challenge: Large language models (LLMs) have shown impressive performance on downstream tasks, but if they cannot be fully described in prompts, they could fail to perform the task.
Approach: They propose a method to contextualize a task toward a large language model (LLM) they use open-ended zero-shot inference from the entire dataset to aggregate the inference results and incorporate the aggregated meta-information for the actual task.
Outcome: The proposed method improves text clustering tasks and improves on several datasets.
ClusterUCB: Efficient Gradient-Based Data Selection for Targeted Fine-Tuning of LLMs (2025.findings-emnlp)

Copied to clipboard

Challenge: Gradient-based data influence approximation is not feasible in practice.
Approach: They propose a gradient-based data selection framework with clustering and a modified Upper Confidence Bound algorithm to solve this problem.
Outcome: The proposed framework can achieve comparable results to the original gradient-based data selection methods while reducing computational consumption.
Interpretable Text Embeddings and Text Similarity Explanation: A Survey (2025.emnlp-main)

Copied to clipboard

Challenge: Text embeddings are a fundamental component in many NLP tasks, but their interpretation and explanation remain challenging.
Approach: They propose a framework for interpretable text embeddings and text similarity explanation . they characterize the main ideas, approaches, and trade-offs and discuss lessons learned .
Outcome: The proposed methods are compared with existing models and compare them with existing ones.
RT-VQ2A2: Real Time Vector Quantized Question Answering with ASR (2024.lrec-main)

Copied to clipboard

Challenge: Existing frameworks for QA with large language models are difficult to implement due to noise, limited context length and latency.
Approach: They propose a model-agnostic framework to address problems in QA with large language models.
Outcome: The proposed framework reduces noise in the ASR output and the limited context length of LLMs and improves performance on the widely used Spoken-SQuAD dataset.
From Documents to Segments: A Contextual Reformulation for Topic Assignment (2026.findings-acl)

Copied to clipboard

Challenge: Traditional topic modeling treats each document as a single, coherent unit of topic.
Approach: They propose a paradigm that redefines topic assignment at the level of segments . they propose 'segment intrusion task' to extend word intrusion to the span level .
Outcome: The proposed paradigm improves topic purity, interpretability and applicability to multi-theme corpora.
D2CS - Documents Graph Clustering using LLM supervision (2025.findings-emnlp)

Copied to clipboard

Challenge: Document clustering does not inherently ensure thematic consistency.
Approach: They propose a framework that constructs a similarity graph over document embeddings and applies iterative graph-based clustering algorithms to partition the corpus into initial clusters.
Outcome: The proposed framework constructs a similarity graph over document embeddings and applies iterative graph-based clustering algorithms to partition the corpus into initial clusters.
Efficient Latent Semantic Clustering for Scaling Test-Time Computation of LLMs (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing methods for scaling test-time computation rely on external models that introduce substantial computational overhead and fail to capture context-aware semantics.
Approach: They propose a method that leverages the generator LLM’s internal hidden states for clustering, eliminating the need for external models.
Outcome: The proposed method improves the computational efficiency of test-time scaling while maintaining or exceeding the performance of existing methods.
Redundancy, Isotropy, and Intrinsic Dimensionality of Prompt-based Text Embeddings (2025.findings-acl)

Copied to clipboard

Challenge: Prompt-based text embedding models generate task-specific embeddables but have thousands of dimensions . dimensionality reductions for embedded text can result in performance degradations of only the first 25% of the dimensions resulting in a very small degradation .
Approach: They investigate how post-hoc dimensionality reduction affects performance of various tasks . they find that embeddings for classification and clustering exhibit lower intrinsic dimensionalities .
Outcome: The proposed model generates task-specific embeddings upon receiving tailored prompts, but has thousands of dimensions and high storage costs.
Sticking to the Mean: Detecting Sticky Tokens in Text Embedding Models (2025.acl-long)

Copied to clipboard

Challenge: Sticky tokens, when repeatedly inserted into sentences, pull sentence similarity toward a certain value, disrupting the normal distribution of embedding distances and degrading downstream performance.
Approach: They propose a method to detect “sticky tokens” by sentence and token filtering and apply it to 40 checkpoints across 14 model families.
Outcome: The proposed method detects 868 sticky tokens across 14 models and shows that their presence does not correlate with model size or vocabulary size.
Length-Induced Embedding Collapse in PLM-based Models (2025.acl-long)

Copied to clipboard

Challenge: In text embeddings from PLMs are essential for many NLP applications, but performance degrades on longer texts.
Approach: They propose a method which mitigates the phenomenon of Length Collapse . they propose TempScale to ensure more consistent embeddings across different text lengths .
Outcome: The proposed method improves performance on MTEB and LongEmbed by 0.94% on short and 1.10% on long texts.
Lemmatization as a Classification Task: Results from Arabic across Multiple Genres (2025.emnlp-main)

Copied to clipboard

Challenge: Existing tools for lemmatization in morphologically rich languages with ambiguous orthography face inconsistent standards and limited genre coverage.
Approach: They propose two new approaches that frame lemmatization as classification into a Lemma-POS-Gloss tagset, leveraging machine translation and semantic clustering.
Outcome: The proposed models perform better than existing models and are more interpretable, the authors show.
EvoMemKG: An Evolvable Memory Agent for Multi-hop Knowledge Graph Reasoning (2026.findings-acl)

Copied to clipboard

Challenge: Existing methods for integrating knowledge graphs with large language models lack continuous learning capabilities.
Approach: They propose an agent framework with a dynamic, evolvable memory mechanism specifically designed for KG reasoning.
Outcome: EvoMemKG achieves state-of-the-art performance without training or tools . it achieves improvements of up to 20% over baseline on multi-hop queries .
MPCG: Multi-Round Persona-Conditioned Generation for Modeling the Evolution of Misinformation with LLMs (2025.emnlp-main)

Copied to clipboard

Challenge: Misinformation evolves as it spreads, shifting in language, framing, and moral emphasis to adapt to new audiences.
Approach: They propose a multi-round, persona-conditioned framework that simulates how claims are iteratively reinterpreted by agents with distinct ideological perspectives.
Outcome: The proposed framework generates persona-specific claims across multiple rounds . it is based on an uncensored large language model and is scalable to multiple tasks .
TH-RAG : Topic-Based Hierarchical Knowledge Graphs for Robust Multi-hop Reasoning in Graph-based RAG Systems (2026.acl-long)

Copied to clipboard

Challenge: Retrieval-augmented generation (RAG) enables large language models to incorporate external knowledge at inference.
Approach: They propose a hierarchical framework that organizes triplets into subtopics and topics to enhance connectivity and integrate dispersed information.
Outcome: Experiments on abstractive and specific QA benchmarks show that TH-RAG outperforms strong baselines in accuracy and robustness while remaining efficient.
PL-MTEB: Polish Massive Text Embedding Benchmark (2026.findings-acl)

Copied to clipboard

Challenge: Text embeddings are used in many NLP tasks, including document clustering, semantic search, question answering, and classification.
Approach: They introduce the Polish Massive Text Embedding Benchmark (PL-MTEB) it is a comprehensive benchmark for text embeddings in the Polish language.
Outcome: The proposed model is based on 30 different NLP tasks in the Polish language.
An Answer is just the Start: Related Insight Generation for Open-Ended Document-Grounded QA (2026.findings-acl)

Copied to clipboard

Challenge: Existing QA benchmarks do not explicitly support document-grounded related insight generation . Existing document-based QA efforts focus on answering fact-based questions .
Approach: They propose a task to generate additional insights from a document collection that improves, extends or rethinks an initial answer to an open-ended question.
Outcome: The proposed task improves, extends, or rethinks an answer to an open-ended question.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations