Papers with clustering
Copied to clipboard
| Challenge: | Topic pages aggregate useful information about an entity or concept into a single concise article. |
| Approach: | They propose a web app that generates topic pages for biomedical entities on demand . they use large language models and retrieval-augmented generation to generate high-quality topics . |
| Outcome: | The proposed method is based on a human evaluation of 150 biomedical topics . it uses large language models and retrieval-augmented generation (RAG) |
Copied to clipboard
| Challenge: | Maintenance record logbooks are an emerging text type in NLP. maintenance record logbook data is often written in non-standard language with many domain specific technical terms, abbreviations, and non-standardized spelling and grammar. |
| Approach: | They propose to create a collaborative open-source library of technical and domain-specific language resources for maintenance record logbooks. |
| Outcome: | The proposed library provides tools to aid in their (pre-)processing and clustering. |
Copied to clipboard
| Challenge: | Existing methods for semantics discovery focus on text, video, and audio, failing to leverage the rich multimodal information in the real world. |
| Approach: | They propose a method to construct augmentation views for multimodal data and use them to perform pre-training to establish well-initialized representations for subsequent clustering. |
| Outcome: | The proposed method improves on benchmark multimodal intent and dialogue act datasets by 2-6% over state-of-the-art methods. |
Copied to clipboard
| Challenge: | Entity coreference resolution aims to identify mentions that refer to the same entity. |
| Approach: | They propose a triad-based neural network system that generates affinity scores between entity mentions for coreference resolution. |
| Outcome: | The proposed system generates affinity scores between mentions for coreference resolution. |
Copied to clipboard
| Challenge: | Recent work in multilingual natural language processing has shown progress on tasks such as natural language inference and joint multilingual translation. |
| Approach: | They propose a technique that groups similar languages together by embeddings from a pre-trained masked language model and automatically discovering language clusters in this embeddable space. |
| Outcome: | The proposed technique outperforms baselines on 15 languages in the WikiAnn dataset showing meaningful multilingual transfer for low-resource languages (Swahili and Yoruba). |
Copied to clipboard
| Challenge: | Maintenance logbooks often contain free text fields with domain specific terms, abbreviations, and non-standard spelling . most standard NLP pipelines for pre-processing and annotation are trained on standard contemporary corpora. |
| Approach: | They propose to create an open-source library and data repository for predictive maintenance language datasets and to evaluate the tools available at MaintNet. |
| Outcome: | The proposed tools improve the performance of existing pipelines and improve the quality of the existing ones. |
Copied to clipboard
| Challenge: | a dialog system is used to evaluate NLU models using aggregated metrics on a large number of utterances. |
| Approach: | They propose a method to generate a test set with high semantic diversity for NLU evaluation in dialog systems. |
| Outcome: | The proposed test sets are based on high diversity of utterances from different regions of the utteration embedding space. |
Copied to clipboard
| Challenge: | Introduction to deep Bayesian learning for natural language addresses the fundamentals of statistical models and neural networks. |
| Approach: | This tutorial addresses the advances in deep Bayesian learning for natural language . it focuses on advanced Bayessian models and deep models . authors present case studies and domain applications to tackle different issues . |
| Outcome: | This tutorial focuses on advanced Bayesian models and deep models for natural language . case studies and domain applications are presented to tackle different issues in deep Bayessian processing, learning and understanding. |
Copied to clipboard
| Challenge: | Writing corrective feedback on learner text is widespread in language education, but it can be time-consuming for teachers. |
| Approach: | They propose to use feedback comment generation to generate explanatory notes for learners by categorizing comments and constraining outputs of noisy classes. |
| Outcome: | The proposed scheme can be used to generate feedback comment corpora using a broader scope than existing typologies focused on error correction. |
Copied to clipboard
| Challenge: | Existing methods for clustering comparable corpora are not suitable for bilingual corpors. |
| Approach: | They propose new clustering models fully adapted to comparable corpora based on a deep variant of Kmeans . they illustrate their behavior on bilingual collections created from Wikipedia . |
| Outcome: | The proposed models show that they can cluster comparable corpora on bilingual collections . the proposed models are based on a state-of-the-art deep variant of Kmeans . |
Copied to clipboard
| Challenge: | Existing evaluations of word/concept representations on verbal fluency tasks rely on human annotations of clusters and switches between sub-categories. |
| Approach: | They analyze word/concept representations in an experimental verbal fluency dataset . they find that ConceptNet embeddings outperforms other semantic representations . |
| Outcome: | The proposed method outperforms other semantic representations by a large margin. |
Copied to clipboard
| Challenge: | The SUMMA Platform is a highly scalable distributed architecture for monitoring a large number of media broadcasts in parallel, with a lag behind actual broadcast time of at most a few minutes. |
| Approach: | The open-source SUMMA Platform is a highly scalable distributed architecture for monitoring a large number of media broadcasts in parallel . it offers a fully automated media ingestion pipeline capable of recording live broadcasts, detection and transcription of spoken content, translation of all text (original or transcribed) into English, recognition and linking of Named Entities, topic detection, clustering and cross-lingual multi-document summarization of related media items and extraction and storage of factual claims in these news items. |
| Outcome: | The SUMMA Platform is a highly scalable distributed architecture for monitoring a large number of media broadcasts in parallel, with a lag behind actual broadcast time of at most a few minutes. |
Copied to clipboard
| Challenge: | Contextualized word representations from pre-trained language models encode more information than is necessary for the identification of word senses and some of this information affect performance negatively in unsupervised settings. |
| Approach: | They propose to use a framework to erase specific information from pre-trained word models and create feature-invariant representations that are invariant to these ‘nuisance features’. |
| Outcome: | The proposed framework erases information from the representations of pre-trained language models, thereby creating feature-invariant representations. |
Copied to clipboard
| Challenge: | Existing methods for event detection require predefined schemas, but manual defining is expensive and labor-intensive. |
| Approach: | They propose a task to achieve event clustering, hierarchy expansion and type naming . they propose 'neighbor Contrastive Clustering' module and a Hierarchy-Aware Linking module . |
| Outcome: | The proposed method outperforms baseline methods on three datasets. |
Copied to clipboard
| Challenge: | Existing methods for identifying intents from unlabeled utterances are label-intensive, inefficient, and inaccurate. |
| Approach: | They propose a multi-task strategy to leverage unlabeled data and external labeled data for representation learning. |
| Outcome: | The proposed method outperforms state-of-the-art methods on three intent recognition benchmarks. |
Copied to clipboard
| Challenge: | a new approach to multilingual word embedding is needed to achieve this goal . a multilingual common semantic space is a language-agnostic semantic continuous space . |
| Approach: | They propose a multilingual common semantic space where words from multiple languages are mapped into a shared space so that resources and knowledge can be shared across languages. |
| Outcome: | The proposed approach achieves 14.6% absolute F-score gain over state-of-the-art methods on cross-lingual direct transfer. |
Copied to clipboard
| Challenge: | Existing studies have used descriptive typological features and a coarse language family classification as baselines for language clustering. |
| Approach: | They propose two types of language groupings based on morpho-syntactic features in a nominal domain and one based upon a head parameter. |
| Outcome: | The proposed methods outperform state-of-the-art embedding-based models in multilingual named entity recognition (NER) . their results suggest that theoretical linguistics plays a significant role in multi-lingual learning tasks. |
Copied to clipboard
| Challenge: | Acquiring high-quality annotated corpora for complex multi-task information extraction (MT-IE) is an arduous and costly process for human-annotators. |
| Approach: | They propose a supervised MT-IE annotation tool built with indirect weak supervision and clustering to maximise annotator productivity. |
| Outcome: | The proposed tool is compared with existing tools in the field of MT-IE and aims to increase annotator productivity. |
Copied to clipboard
| Challenge: | Existing methods for encoding instruction information fail to be sensitive to clearer criteria like “evaluate similarity based on emotion” . instead, we propose a different approach, which treats the instruction as a “question” about the input text and encodes the expected answers to obtain the representation accordingly. |
| Approach: | They propose a text embedder that captures characteristics of texts specified by user instructions clarifying the similarity criterion. |
| Outcome: | The proposed model improves instruction-following capabilities when applied to large language models and encoder-based LMs. |
Copied to clipboard
| Challenge: | a recent study shows that social media posts are often given as continuous coordinates . but, the resulting discrete coordinates do not always correspond to existing linguistic areas . |
| Approach: | They propose an algorithm for clustering coordinates and associating them with towns using point-to-city (P2C) they compare accuracy of a state-of-the-art geolocation model with P2C labels to one with regular k-d tree labels. |
| Outcome: | The proposed method improves accuracy at 100 miles, but degrades for finer-grained distinctions . iterative k-d tree-based method can cluster coordinates and associate them with towns . |
Copied to clipboard
| Challenge: | a probabilistic clustering algorithm can help users find posts that discuss experiences similar to their own . a recent study shows that probabilistic Clustering can yield a better performance than baseline clustering methods . |
| Approach: | They propose a probabilistic clustering algorithm that can help Reddit users find posts that discuss experiences similar to their own. |
| Outcome: | The proposed algorithm can find posts that discuss experiences similar to their own . it performs better than baseline clustering methods due to high runtime overhead . |
Copied to clipboard
| Challenge: | Unsupervised text clustering is unlikely to produce groupings that work across use cases . authors present techniques to effectively control text embeddings with minimal human input . |
| Approach: | They propose techniques to control text embeddings with minimal human input . they evaluate clustering performance for datasets with multiple independent labels . |
| Outcome: | The proposed techniques improve clustering for one perspective or use case, but at a tradeoff in performance for another use case. |
Copied to clipboard
| Challenge: | Existing methods for extracting conditional text embeddings from large language models (LLMs) relying on prompts often fails to produce high-quality conditional embeddables, resulting in degradation of quality. |
| Approach: | They propose a plug-and-play method that constructs unconditional general text embeddings and uses them to refine conditional text embeds. |
| Outcome: | The proposed method improves performance of prompt-based methods on clustering, Semantic Textual Similarity, and triplet alignment datasets. |
Copied to clipboard
| Challenge: | a multi-stage approach to hierarchical clustering of interaction drivers in contact centers is proposed . silhouette score and human preference score are improved by 36.7% for top-level clusters compared to standard agglomerative clustering . |
| Approach: | They propose a multi-stage approach that introduces different perspectives or views to improve the quality of hierarchical clustering of interaction drivers in a contact center. |
| Outcome: | The proposed approach improves the quality of generated clusters on public datasets with minimal query time compared to the current state-of-the-art approaches. |
Copied to clipboard
| Challenge: | Existing approaches to extract actionable suggestions from customer reviews are often mixed-intent, unstructured text. |
| Approach: | They propose a hybrid pipeline that uses a RoBERTa classifier and a precision–recall surrogate to extract actionable suggestions from customer reviews. |
| Outcome: | The proposed pipeline outperforms prompt-only, rule-based, and classifier-only baselines in extraction accuracy and cluster coherence. |
Copied to clipboard
| Challenge: | Ambiguous user queries pose a challenge in task-oriented dialogue systems . Large Language Models (LLMs) rely on the top-k retrieved documents for clarification . traditional approaches lack principled mechanisms to determine when to use broad domain knowledge vs specific retrieved document context for clarification. |
| Approach: | They propose a hybrid approach that dynamically chooses between document-based or aspect-based clarification based on query ambiguity. |
| Outcome: | The proposed approach shows significant improvements over baselines on product troubleshooting and product search datasets. |
Copied to clipboard
| Challenge: | a novel backdoor attack is based on textual claims to trick models into misbehaving on targeted claims. |
| Approach: | a new backdoor attack is designed to trick models into misbehaving on targeted claims . the code and data will be available at https://github.com/minkyoo9/CGBA . |
| Outcome: | a new backdoor attack exploits the power of textual claims to trick models into misbehaving on claims without affecting their performance on clean data. |
Copied to clipboard
| Challenge: | Sentence embeddings are an important component of many natural language processing systems. |
| Approach: | They propose a self-supervised objective for learning universal sentence embeddings that does not require labelled training data. |
| Outcome: | The proposed approach closes the performance gap between unsupervised and supervised pretraining for universal sentence encoders. |
Copied to clipboard
| Challenge: | Existing methods for fine-grained classification categorize texts into coarse-gritty classes, but they are suboptimal in real-world scenarios. |
| Approach: | They propose a lightweight contrastive clustering-based bootstrapping method to iteratively refine the labels of passages. |
| Outcome: | The proposed method outperforms the state-of-the-art methods by a large margin on NYT and 20News datasets. |
Copied to clipboard
| Challenge: | Aspect category detection (ACD) aims to automatically identify user-concerned aspects from online reviews. |
| Approach: | They propose a method that relies on the category name of each aspect and a pretrained language model to generate constraints for clustering. |
| Outcome: | The proposed framework performs better than existing weakly supervised methods on nine benchmark datasets. |
Copied to clipboard
| Challenge: | Existing approaches to detect vaccine attitudes on social media require abundant annotations and pre-defined aspect categories. |
| Approach: | They propose a semi-supervised approach to detect vaccine attitudes on social media . they use an autoencoding architecture to learn from unlabelled data the topical information of the domain . |
| Outcome: | The proposed model outperforms existing aspect-based models on stance detection and tweet clustering. |
Copied to clipboard
| Challenge: | Lexical tones play a crucial role in Sino-Tibetan languages, but current phonetic fieldwork relies on manual effort. |
| Approach: | They propose a pitch-based similarity representations for tone transcription called Tone2Vec . they propose an open-source package that facilitates automated fieldwork and analysis . |
| Outcome: | Experiments on dialect clustering and variance show that Tone2Vec captures fine-grained tone variation. |
Copied to clipboard
| Challenge: | Inverted file structure is a common technique for accelerating dense retrieval, but its lossy nature degrades it. |
| Approach: | They propose a hybrid index where embedding clusters and salient terms work collaboratively to accelerate dense retrieval. |
| Outcome: | The proposed method achieves lossless retrieval quality with competitive efficiency across index settings. |
Copied to clipboard
| Challenge: | Training-free embedding methods focus on optimizing embeddable prompts . previous methods have overlooked the benefits of utilizing generative abilities of LLMs - GenEOL . |
| Approach: | They propose a method that leverages pretrained large language models to embed text . they propose generating diverse transformations of a sentence that preserve its meaning . |
| Outcome: | The proposed method outperforms existing training-free embedding methods by 2.85 points on the sentence semantic text similarity (STS) benchmark. |
Copied to clipboard
| Challenge: | 'A language is a dialect with an army and navy' is attributed to sociologist Max Weinrich. |
| Approach: | They propose a universal character-based method for representing sentences so that one can calculate the distance between any two sentence pairs. |
| Outcome: | The proposed method can be used to calculate distance between two sentences by clustering a dialect/sub-language mixed corpus into sub-groups and to partially answer the question of what separates languages from dialects. |
Copied to clipboard
| Challenge: | Sentence BERT is inefficient for sentence-pair tasks as it needs to evaluate combinatorially many sentence pairs which is very time-consuming. |
| Approach: | They propose a lightweight extension on top of BERT and a self-supervised learning objective to derive meaningful sentence embeddings in an unsupervised manner. |
| Outcome: | The proposed method outperforms baselines on common semantic textual similarity tasks and downstream supervised tasks and achieves performance competitive with supervised methods on various tasks. |
Copied to clipboard
| Challenge: | Contextual large language model embeddings are often monolingual, do not scale, and struggle in multilingual settings. |
| Approach: | They propose a hierarchical approach to embed news articles and social media data using Matryoshka embeddings that can determine story similarity at varying levels of granularity based on which subset of dimensions is examined. |
| Outcome: | The proposed model achieves state-of-the-art performance on the SemEval 2022 task 8 dataset. |
Copied to clipboard
| Challenge: | Existing methods to discover facts from natural language text are based on relation extraction and open information extraction. |
| Approach: | They propose a task of generating a machine-readable representation of the most prominent information in a text document as a set of facts. |
| Outcome: | The proposed system outperforms baselines and text summarizers in a supervised evaluation of salience tasks. |
Copied to clipboard
| Challenge: | Existing approaches to intent detection rely on epoch wise clustering and classification based on labeled and unlabeled data. |
| Approach: | They propose an end-to-end deep contrastive clustering algorithm that jointly updates model parameters and cluster centers via supervised and self-supervised learning. |
| Outcome: | The proposed approach outperforms baselines on five public datasets and human-in-the-loop variant for practical deployment. |
Copied to clipboard
| Challenge: | End-to-end (E2E) spoken language understanding models are constrained by the cost of collecting speech-semantics pairs. |
| Approach: | They propose a model that learns E2E SLU without speech-semantics pairs . they propose cross-modal selective self-training (CMSST) to address imbalance and noise issues . |
| Outcome: | The proposed model learns E2E SLU without speech-semantics pairs . the proposed model requires the domains of speech-text and text-sensitization to match . |
Copied to clipboard
| Challenge: | Existing prompt engineering methods rely on randomly selected evaluation subsets, leading to suboptimal prompts. |
| Approach: | They propose an iterative evaluation data selection approach for effective prompt optimization using real time model performance. |
| Outcome: | The proposed approach improves effectiveness by 1.6% to 3.1% and stability by 50% to 55.5% on two datasets BIG-bench and LIAR and two models GPT-3.5 and GPT-4o-mini. |
Copied to clipboard
| Challenge: | Existing text embeddings are evaluated on a small set of datasets, not covering their possible applications to other tasks. |
| Approach: | They propose a benchmarking framework that evaluates 8 embedding tasks covering 58 datasets and 112 languages. |
| Outcome: | The proposed model is the most comprehensive benchmark of text embeddings to date. |
Copied to clipboard
| Challenge: | Pretrained language models often need to specialize to specific domains. |
| Approach: | They propose an approach that performs weight-space averaging of adapters trained on different domains. |
| Outcome: | The proposed approach improves performance to new domains without extra training. |
Copied to clipboard
| Challenge: | Short texts pose significant challenges for clustering due to semantic sparsity, limited context and fuzzy category boundaries. |
| Approach: | proposed framework incorporates neighborhood information at instance and cluster levels . a cluster-level framework introduces fuzzy neighborhood-aware weighting . |
| Outcome: | The proposed framework outperforms state-of-the-art models on short texts . it excludes neighbors from negative sample set to enhance inter-cluster separability . |
Copied to clipboard
| Challenge: | Using mixture-of-experts (MoE) to deal with language heterogeneity is a challenge in neural machine translation (NMT). |
| Approach: | They propose a lightweight MoE-based NMT model that is trained via an elaborate stage-wise training strategy. |
| Outcome: | The proposed model achieves stable improvements in translation tasks by introducing fewer extra parameters compared to baseline models. |
Copied to clipboard
| Challenge: | k-Nearest-Neighbor Machine Translation (kNN-MT) is a non-parametric solution for domain adaptation . previous studies have shown that kNN retrieval is at the expense of high latency . |
| Approach: | They propose to use clustering to improve retrieval efficiency by combining a non-parametric MT with an in-domain feature-based retrieval module. |
| Outcome: | The proposed method reduces translation latency by 57% while maintaining the most useful information of the original datastore. |
Copied to clipboard
| Challenge: | a number of machine learning models inherit and amplify the societal biases in data. |
| Approach: | a new bias detection technique based on clustering is proposed to detect local biases in data . authors propose to use LOGAN to analyze local bias in data. |
| Outcome: | The proposed technique detects bias in a local region and allows better analysis of model predictions. |
Copied to clipboard
| Challenge: | Existing methods for deep clustering optimization with shallow models have limited performance due to poor power of feature learning. |
| Approach: | They propose a general deep clustering optimization method that leverages information feedback to construct generalized labels to optimize the deep model. |
| Outcome: | The proposed method reduces the impact of noise on the clustering process by using correlation relationship between the samples. |
Copied to clipboard
| Challenge: | Recent approaches demonstrate that MLLMs can be adapted into competitive embedding models via large-scale contrastive learning. |
| Approach: | They propose a compressed pre-training phase which serves as a warm-up stage for contrastive learning. |
| Outcome: | The proposed model achieves state-of-the-art among MLLMs of comparable size on the MMEB, realizing optimization in both efficiency and effectiveness. |
Copied to clipboard
| Challenge: | Recent work on news comment summarization has focused on extractive methods within constraints. |
| Approach: | They propose an enhanced fast clustering algorithm that maintains a dynamic similarity threshold to ensure high density of each comment cluster being built. |
| Outcome: | The proposed method improves the baseline methods and the test suite on real-world news comments. |
Copied to clipboard
| Challenge: | Existing methods for learning to compare social media users fail to generalize to new users or even to previously known users. |
| Approach: | They propose a procedure to learn a mapping from short episodes of user activity to a vector space in which the distance between points captures the similarity of the corresponding users’ invariant features. |
| Outcome: | The proposed procedure can be applied to users not seen at training time and enables efficient comparisons of users in the resulting vector space. |
Copied to clipboard
| Challenge: | Existing text embedding benchmarks for financial domains are inadequately addressing the nuanced requirements of specialized domains like finance. |
| Approach: | They propose a finance-adapted embedding model that outperforms general-purpose models . they also introduce a new model, Fin-E5, which is also open-sourced . |
| Outcome: | The proposed framework outperforms general-purpose models on financial embedding tasks. |
Copied to clipboard
| Challenge: | Current image clustering methods neglect the use of generated textual descriptions. |
| Approach: | They propose to use image captioning and visual question-answering to cluster images . they propose a new approach to inject task- or domain knowledge into image clustering . |
| Outcome: | The proposed method outperforms existing methods on eight image clustering datasets. |
Copied to clipboard
| Challenge: | Data quality is vital for business decisions; poor data quality costs organizations an average of $12.9 million annually. |
| Approach: | They propose a framework that combines statistical inliner detection with LLM-driven rule and code generation. |
| Outcome: | The proposed framework produces semantically valid quality rules and validates them with retrieval-augmented generation (RAG) Extensive evaluations on benchmark datasets confirm the effectiveness of the proposed framework. |
Copied to clipboard
| Challenge: | a novel method for online news stream clustering is proposed . a user can scour the many news sources multiple times a day to find news articles . |
| Approach: | They propose a method for online news stream clustering that is a variant of the streaming K-means algorithm. |
| Outcome: | The proposed model achieves state-of-the-art on a standard stream clustering dataset of English documents. |
Copied to clipboard
| Challenge: | Current word embeddings in natural language processing do capture context and thus can be leveraged to enrich linguistic analyses. |
| Approach: | They propose a model which leverages pre-trained BERT to cluster contextualized representations of a word based on context in which it appears and labels of items it occurs in. |
| Outcome: | The proposed model can detect interpretable, finer-grained context patterns associated with (im)polite language. |
Copied to clipboard
| Challenge: | Existing approaches to linking entities ignore relationships between entities in biomedical knowledge bases. |
| Approach: | They propose a model which can link mentions of unseen entities using learned representations of entities. |
| Outcome: | The proposed model improves on the largest publicly available biomedical dataset by 3.0 points of accuracy and 2.3 points of reliability. |
Copied to clipboard
| Challenge: | Existing work on intent-related models fails to capture long-term dependencies in user behavior and fails to effectively utilize item relevance. |
| Approach: | They propose a sequential recommendation framework that combine temporal variability with position encoding that has extrapolation properties to encode sequences, thereby expanding the model’s view of user behavior. |
| Outcome: | The proposed model improves on three real datasets by 0.8% to 14.7% compared to baselines. |
Copied to clipboard
| Challenge: | a new framework to analyze how latent concepts are encoded in representations learned in pre-trained lan-guage models is proposed . conceptX uses clustering to discover the encoded concepts and align them with a large set of human-defined concepts. |
| Approach: | They propose a framework to analyze how latent concepts are encoded in representations learned within pre-trained lan-guage models. |
| Outcome: | The proposed framework explains encoded concepts by aligning with human-defined concepts. |
Copied to clipboard
| Challenge: | Existing studies have shown that data diversity affects the performance of LMs if we train a single LM over the entire dataset. |
| Approach: | They propose an autoencoding topic model with a mixture prior to perform clustering for the data. |
| Outcome: | The proposed model can learn knowledge from different samples while extracting cluster-specific features. |
Copied to clipboard
| Challenge: | Existing state-of-the-art (SOTA) SED models rely on graph neural networks (GNNs) Existing SED frameworks rely heavily on GNNs, which require complex graph construction and time-consuming training processes. |
| Approach: | They propose a framework that leverages the rich background knowledge of large language models to formalize and disambiguate short texts by completing abbreviations and summarizing informal expressions. |
| Outcome: | The proposed framework outperforms existing models on two challenging real-world datasets. |
Copied to clipboard
| Challenge: | Existing methods for text clustering use static pseudo-oracles, i.e., unidirectionally querying them for similarity assessment or data augmentation. |
| Approach: | They propose a training framework that enables bidirectional refinement between LLMs and embedding models by using task-aware prompts to guide the LLM in generating interpretations for the input texts. |
| Outcome: | Experiments on 14 benchmark datasets across 5 tasks demonstrate the effectiveness of the proposed training framework. |
Copied to clipboard
| Challenge: | Weak supervision is a problem in text classification, but it requires corpusspecific knowledge. |
| Approach: | They propose a framework for extremely weak supervision that can be used to train a text classifier. |
| Outcome: | The proposed framework outperforms seed-driven weakly supervised methods on 7 benchmark datasets. |
Copied to clipboard
| Challenge: | Existing methods for instruction tuning are limited due to the increasing volume of instruction datasets and the increased computational costs. |
| Approach: | They propose to extract a small and highly informative subset of training samples from a large dataset that achieves comparable performance to the full dataset. |
| Outcome: | The proposed algorithm outperforms other unsupervised methods and achieves comparable performance to the full dataset. |
Copied to clipboard
| Challenge: | a new approach to cognate detection is proposed to capture the remaining similarities between cognate word forms after thousands of years of divergence. |
| Approach: | They propose a method which uses information weighting and sound correspondence modeling to improve cognate detection. |
| Outcome: | The proposed approach improves on the measure of form similarity and distance-based cognate clustering. |
Copied to clipboard
| Challenge: | Scholars often need to go beyond textual analysis for establishing provenance of historical documents. |
| Approach: | They propose a deep and interpretable probabilistic generative model to analyze glyph shapes in printed Early Modern documents by generating a latent vector responsible for inking variations, jitter, noise and other unforeseen phenomena. |
| Outcome: | The proposed model outperforms interpretable clustering baselines and overly-flexible deep generative models on the task of completely unsupervised discovery of typefaces in mixed-fonts documents. |
Copied to clipboard
| Challenge: | Existing methods for Open Relation Extraction (OpenRE) use a two-stage pipeline, which learns relation representations and assignments in the first stage, then manually labels relation for each cluster. |
| Approach: | They propose a method that performs relation learning and relation labeling simultaneously without a significant increase in human effort. |
| Outcome: | The proposed method improves existing SOTA methods by 13.8% and 10.6% on two datasets. |
Copied to clipboard
| Challenge: | Topic models are useful tools for analyzing and interpreting the main underlying themes of large corpora of text. |
| Approach: | They propose a self-supervised neural topic model that learns a topic representation jointly from three co-occurring words and a document that the triple originates from. |
| Outcome: | The proposed model outperforms existing topic models in coherence metrics and document clustering accuracy. |
Copied to clipboard
| Challenge: | In argumentation, framing is used to emphasize a specific aspect of a topic while concealing others. |
| Approach: | They propose an unsupervised method for framing arguments into non-overlapping frames . authors propose a corpus of 12, 326 debate-portal arguments organized along the frames of debates' topics . |
| Outcome: | The proposed method outperforms baselines on the argumentation task by 0.28 points. |
Copied to clipboard
| Challenge: | Recent studies have shown great improvements in instruction-following capability through additional training for instruction- following tasks. |
| Approach: | They propose to use a Transformer-based causal language model to study instruction-following capabilities. |
| Outcome: | The proposed model learns task-specific information by clustering data within its hidden space, with this clustering process evolving dynamically during learning. |
Copied to clipboard
| Challenge: | Existing OpenRE methods assume unlabeled data is a mixture of known and novel instances. |
| Approach: | They propose a generalized OpenRE setting that considers unlabeled data as a mixture of known and novel instances. |
| Outcome: | The proposed framework outperforms baselines in relation classification and clustering on three benchmark datasets. |
Copied to clipboard
| Challenge: | Using hidden-state vectors of recurrent neural networks (RNNs) we examine the assumption that hidden- state vectors tend to form clusters of semantically similar vectors, which we dub the clustering hypothesis. |
| Approach: | They propose to use recurrent neural networks (RNNs) that model processes with internal states to test their hypothesis. |
| Outcome: | The proposed model is based on a set of RNNs that were trained to recognize regular languages and a context-free language. |
Copied to clipboard
| Challenge: | Infusing clustering with active learning with AL can overcome the bias issue of both AL and traditional annotation methods while exploiting AL’s annotation efficiency. |
| Approach: | They propose an algorithm that dynamically adjusts clustering and annotation efforts in response to an estimated classifier error-rate. |
| Outcome: | The proposed algorithm outperforms baseline AL approaches with pretrained transformers and traditional Support Vector Machines on eight datasets for emotion, hatespeech, dialog act, and book type detection tasks. |
Copied to clipboard
| Challenge: | a new study examines the degree of alignment between languages in multilingual embeddings . cross-lingual embeds are designed to encode linguistic concepts that bridge equivalent semantic meaning . a comprehensive approach is needed to address these questions. |
| Approach: | They employ clustering to uncover latent concepts within multilingual models . they introduce two metrics to quantify alignment and overlap of these concepts . |
| Outcome: | The proposed model can capture linguistic nuances across languages, but is not language-agnostic? the proposed model is able to capture nuances in multiple languages, the authors say. |
Copied to clipboard
| Challenge: | Existing resources cover only a small number of tasks, limiting its practical usefulness. |
| Approach: | They propose a zero-shot learning approach to script parsing which enables us to acquire script knowledge without domain-specific annotations. |
| Outcome: | The proposed model outperforms a previous model with scenario-specific supervision and achieves 68.1/74.4 average F1 for event / participant parsing. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) exhibit bias toward underrepresented groups, despite advances in active learning. |
| Approach: | They propose a clustering-based active learning framework enhanced with knowledge distillation that transforms the intermediate outputs of the learner model to yield more representative models without prior knowledge of underlying data distribution. |
| Outcome: | The proposed framework improves performance across data subgroups and lexical diversity, underscoring the model’s resilience to skewness in available data. |
Copied to clipboard
| Challenge: | Conditional language models can generate a diverse set of outputs, but for open-ended tasks, beam search is ill-suited to generating a set of diverse sequences. |
| Approach: | They propose a method where we over-sample candidates and use clustering to remove similar sequences to achieve high diversity without sacrificing quality. |
| Outcome: | The proposed method over-samples candidates and removes similar sequences to achieve high diversity without sacrificing quality. |
Copied to clipboard
| Challenge: | Existing representation learning models do not capture the intra-sentential and inter-sententential features of long-text. |
| Approach: | They propose a graph-based representation for document clustering that builds a Graph Autoencoder on a Keyword Correlation Graph. |
| Outcome: | The proposed graph autoencoder can achieve better clustering performance than existing features. |
Copied to clipboard
| Challenge: | Large pretrained language models (LMs) have been criticized for lack of grounding, i.e., connecting words to their meanings in the physical world. |
| Approach: | They compare vision-and-language (VL) models trained jointly on text and image or video data to find out how they compare to text-only counterparts. |
| Outcome: | The proposed model outperforms the text-only variants on a commonsense question answering task. |
Copied to clipboard
| Challenge: | Existing methods to extract product features from unstructured text still suffer from problems . e-commerce platforms are focusing on multi-scale values, which can be confusing . |
| Approach: | They propose a pre-training technique to automatically obtain attribute value pairs from product descriptions to aid e-commerce. |
| Outcome: | The proposed method improves on the existing token-level masking strategy and achieves state-of-the-art on four benchmarks. |
Copied to clipboard
| Challenge: | Existing studies have not explored the relationships between lyrics and dance motions . previous studies focused on synthesizing or retrieving dance motion from lyrics . |
| Approach: | They propose a method to detect parts of songs where meaningful relationships exist . they use clustering to transform lyrics and dance motions into symbols . |
| Outcome: | The proposed method outperforms existing methods on prose and non-dance dance motions. |
Copied to clipboard
| Challenge: | Existing methods for finding similar sentences require multiple inferences . a modern GPU requires 65 hours to find the most similar pair in 10,000 sentences . |
| Approach: | They propose a modification of the pretrained BERT network that uses siamese and triplet networks to derive semantically meaningful sentence embeddings. |
| Outcome: | The proposed method outperforms existing methods on sentence-pair regression tasks. |
Copied to clipboard
| Challenge: | Recent studies highlight the effectiveness of using in-context learning (ICL) to steer large language models in processing tabular data. |
| Approach: | They propose a method that uses clustering and evolutionary strategies to curate a representative sample set from training data. |
| Outcome: | The proposed method significantly improves fairness across various metrics, showing its efficacy in real-world scenarios. |
Copied to clipboard
| Challenge: | Generalized category discovery (GCD) is a crucial task in open-world computing, where new categories frequently emerge, necessitating models that can adapt and learn continually. |
| Approach: | They propose to integrate the feedback from LLMs into an active learning paradigm to simplify the labeling task and minimize the spread of inaccurate feedback. |
| Outcome: | The proposed approach significantly improves baseline models at a nominal average cost. |
Copied to clipboard
| Challenge: | Existing event extraction methods require predefined event types and their annotations to learn event extractors. |
| Approach: | They propose to represent each event type as a cluster of predicate sense, object head> pairs. |
| Outcome: | The proposed method can discover salient and high-quality event types on three datasets from different domains. |
Copied to clipboard
| Challenge: | Existing methods focus on transferring in-domain (IND) prior knowledge to out-of-domain data through pre-training and clustering. |
| Approach: | They propose a Pseudo-Label enhanced Prototypical Contrastive Learning model for uniformed intent discovery that integrates supervised and pseudo signals from IND and OOD data. |
| Outcome: | The proposed method has been proven effective in two different settings of discovering new intents. |
Copied to clipboard
| Challenge: | Supervised Semantic Differential (SSD) is a mixed quantitative–interpretive method that models how text meaning varies with continuous individual-difference variables . currently no systematic method exists for choosing the number of retained components, introducing avoidable researcher degrees of freedom in the analysis pipeline. |
| Approach: | They propose a PCA sweep procedure that treats dimensionality selection as a joint criterion over representation capacity, gradient interpretability, and stability across nearby values of K. |
| Outcome: | The proposed method is based on a corpus of short posts about artificial intelligence written by Prolific participants who also completed Admiration and Rivalry narcissism scales. |
Copied to clipboard
| Challenge: | Dialects are one of the main drivers of language variation, a major challenge for natural language processing tools. |
| Approach: | They use a corpus of 16.8M anonymous online posts to learn continuous document representations of cities. |
| Outcome: | The proposed method matches dialect areas at different granularities against an existing dialect map. |
Copied to clipboard
| Challenge: | TR-MTEB is the first large-scale, task-diverse benchmark for sentence embedding models for Turkish. |
| Approach: | a new benchmark evaluates sentence embedding models for Turkish . TR-MTEB covers six core tasks and 26 high-quality datasets . |
| Outcome: | The TR-MTEB benchmark covers six core tasks and includes 26 high-quality datasets . the models achieve competitive performance across most tasks and significantly improve on baseline models. |
Copied to clipboard
| Challenge: | a novel method for clustering news across languages is proposed . a key challenge in handling news streams is that they must be generated on the fly . |
| Approach: | They propose a method for clustering news across languages into monolingual and crosslingual clusters . they use real news datasets in multiple languages to find an ever growing number of cluster labels . |
| Outcome: | The proposed method produces state-of-the-art results on real news datasets in German, English and Spanish. |
Copied to clipboard
| Challenge: | Existing approaches to self-training are based on reject sampling and lack quality reasoning paths. |
| Approach: | They propose a framework for self-training using a generate-and-filter paradigm . they propose to identify diverse and informative samples from redundant data and exploit them more strategically. |
| Outcome: | The proposed framework exploits informative samples from redundant data and improves reasoning trajectory prospecting. |
Copied to clipboard
| Challenge: | Named Entity Recognition (NER) is a task of discovering information entities and identifying their corresponding categories. |
| Approach: | They propose a NER-specific framework to inject coarse-to-fine named entity knowledge into pre-trained models by using a remote supervision strategy. |
| Outcome: | The proposed framework achieves significant improvements against several pre-trained base-lines, demonstrating its effectiveness in label-few and low-resource scenarios. |
Copied to clipboard
| Challenge: | Existing methods to fine-tune pre-trained models for text classification are poor in practice. |
| Approach: | They propose to add an intermediate unsupervised classification task between pre-training and fine-tuning phases to boost performance of pre-trained models. |
| Outcome: | The proposed method improves performance on topical classification tasks when labeled data is scarce. |
Copied to clipboard
| Challenge: | a crowdsourcing platform is used to collect linguistic data and document language use, with a focus on regional variation in European French. |
| Approach: | They propose a crowdsourcing platform to collect linguistic data and document language use with a special focus on regional variation in European French. |
| Outcome: | The proposed platform collects linguistic data and documents language use with a special focus on regional variation in European French. |
Copied to clipboard
| Challenge: | Existing approaches to manage toxic speech on social platforms are limited . however, there is a need for more proactive moderation of abusive speech. |
| Approach: | They extend parallel text detoxification corpus to new languages to test the approach . they propose a method that combines toxic and non-toxic sentences into a more neutral form . |
| Outcome: | The proposed method integrates the descriptive features of toxic and non-toxic sentences into a more neutral or non- toxic form. |
Copied to clipboard
| Challenge: | Low-resource languages often suffer from a lack of high-coverage lexical resources. |
| Approach: | They propose a method to generate cognate tables by clustering words from existing lexical resources. |
| Outcome: | The proposed method outperforms baselines on the Romance and Turkic language families. |
Copied to clipboard
| Challenge: | Temporal Logic (STL) is a formal specification tool for cyber-physical systems . but it is difficult to transform ambiguous and complex data into STL, a paper argues . |
| Approach: | They propose a NL-STL dataset with 16,000 samples enriched with diverse patterns . they propose KGST framework to transform natural language into STL using a generate-then-refine process . |
| Outcome: | The proposed dataset outperforms baseline models in diversity and accuracy . the proposed dataset contains 16,000 samples enriched with diverse patterns . |
Copied to clipboard
| Challenge: | Existing methods to measure semantic change with contextual word embeddings (CWEs) are not suitable for highly imbalanced datasets and pose challenges for interpretation. |
| Approach: | They propose an interpretable, feature-level approach to analyzing language change using k-sparse autoencoders to trace the semantic evolution of the term "indigène(s)" between 1825 and 1950. |
| Outcome: | The proposed approach can learn interpretable features from over 210,000 CWEs generated using sentences from the French National Library. |
Copied to clipboard
| Challenge: | Neural Topic Models and Large Language Models (LLMs) primarily use contextual embeddings from LLMs, which are not optimal for clustering or topic generation. |
| Approach: | They propose a framework that leverages Encoder-Decoders to generate highly clusterable embeddings that could generate topics that exhibit enhanced clusterability and enhanced semantic coherence compared to existing methods. |
| Outcome: | The proposed framework is efficient to train and exhibits high adaptability, demonstrating its potential for a wide array of applications. |
Copied to clipboard
| Challenge: | Existing methods for clustering short texts are inadequate due to the limited amount of information provided by each text sample. |
| Approach: | They propose a Mutual Information Maximization Framework for Short Text Clustering which maximizes mutual information between representations on sequence and token levels. |
| Outcome: | The proposed framework outperforms the state-of-the-art method in terms of Accuracy or Normalized Mutual Information in most cases. |
Copied to clipboard
| Challenge: | Existing unsupervised clustering methods lack label knowledge, resulting in suboptimal performance. |
| Approach: | They propose to use LLM-driven labels to generate positive pairs from embedded data and an embedder to obviate the need for negative pairs. |
| Outcome: | The proposed framework surpasses state-of-the-art benchmarks on a range of datasets and generates interpretable labels for improved understanding of clustering results. |
Copied to clipboard
| Challenge: | a comprehensive benchmark for Persian text embeddings is built upon the Massive Text Embedding Benchmark (MTEB) 63 datasets are included in the benchmark, including a novel task of summary retrieval. |
| Approach: | They propose a benchmark for Persian (Farsi) text embeddings built upon the Massive Text Embedding Benchmark. |
| Outcome: | The proposed framework includes 63 datasets spanning seven different tasks . the evaluation datasets were rigorously evaluated by humans and automated systems . |
Copied to clipboard
| Challenge: | Existing methods for enhancing dialogue performance rely on summarizing behavior . e-commerce chatbots need to align their dialogue strategies with human behavior to achieve coherent, human-like conversations with customers. |
| Approach: | They propose a method to extract core patterns from dialogue data and integrate them into models by mining service thought processes using a multi-agent aPproach. |
| Outcome: | The proposed method outperforms manual methods and outperfies baselines on Taobao in China. |
Copied to clipboard
| Challenge: | Existing methods to fine-tune discriminative models address these challenges by focusing on in-domain intents. |
| Approach: | They evaluate ChatGPT on OOD intent discovery and generalized intent discovery tasks . they outline the strengths and weaknesses of ChatGPt and outline their results . |
| Outcome: | The proposed task aims to extend a closed intent classifier to open-world intent sets. |
Copied to clipboard
| Challenge: | Existing approaches to SAS use unsupervised clustering and have teachers label some items after clustering. |
| Approach: | They propose to use semi-supervised clustering to provide structured groups of answers in addition to a score. |
| Outcome: | The proposed method improves clustering performance from 0.504 kappa for unsupervised clustering to 0.566 kppa. |
Copied to clipboard
| Challenge: | Sentence embeddings are central to natural language processing, but their internal features are not interpretable and users lack fine-grained control for downstream tasks. |
| Approach: | They propose a formal framework to characterize the organization of features in sentence embeddings . they show how they can be composed to capture richer semantic structures . |
| Outcome: | The proposed method can be used to capture richer semantic structures. |
Copied to clipboard
| Challenge: | Existing inference-time optimization strategies address the shortsightedness of auto-regressive generation, but the vast search space leads to excessive exploration and insufficient exploitation. |
| Approach: | They propose a decoding strategy that approximates two distributions via foresight and clustering to provide an efficient estimation of step value. |
| Outcome: | The proposed decoding strategy outperforms strong baselines in performance and efficiency. |
Copied to clipboard
| Challenge: | a large dataset of news articles spanning 20 languages is lacking for keyword extraction. |
| Approach: | They propose a large-scale multi-lingual keyword extraction dataset for 11 of 20 languages . authors believe it will help advance the field of automatic keyword extraction . |
| Outcome: | The proposed dataset is the first for 11 of 20 languages and is based on 540K+ news articles from the BBC News network. |
Copied to clipboard
| Challenge: | a paradigm discovery problem is a task of learning an inflectional morphological system from unannotated sentences. |
| Approach: | They formalize the paradigm discovery problem and develop evaluation metrics for judging systems . they use word embeddings and string similarity to cluster forms by cell and by paradigm . |
| Outcome: | The proposed system suggests clustering by cell across different inflection classes is the most pressing challenge for future work. |
Copied to clipboard
| Challenge: | Existing clustering-based open relation extraction methods use pre-trained language models . embeddings from language models are high-dimensional and anisotropic, so there is a gap . |
| Approach: | They propose a framework that makes two LLMs work collaboratively to achieve clustering. |
| Outcome: | The proposed framework outperforms existing methods by 1.4%3.13% on different datasets. |
Copied to clipboard
| Challenge: | Existing methods for obtaining task-specific labels require prior knowledge of clustering categories and uncontrollable clustering centers. |
| Approach: | They propose a framework for supervised clustering using a discrete process and a robust Contrastive Learning module. |
| Outcome: | The proposed framework outperforms state-of-the-art models on a real-world dataset with just one label per class . the proposed framework is based on k-means clustering and a robust Contrastive Learning module . |
Copied to clipboard
| Challenge: | Language models exhibit a drop in performance on noisy data, which can cause classifiers to incorrectly change their predictions. |
| Approach: | They propose to use Prototype-Based Networks to classify examples based on their similarity to prototypical examples of a class (prototypes) they show that PBNs offer more robustness under both targeted and static adversarial attacks. |
| Outcome: | The proposed model is robust to noise and targets both targeted and static attacks. |
Copied to clipboard
| Challenge: | Aspect-based sentiment analysis (ABSA) is a task that involves classifying aspects of products or services described in user reviews. |
| Approach: | They propose a method for constructing a corpus that is automatically annotated with implicit aspects by combining explicit and unlabeled sentences. |
| Outcome: | The proposed method achieves a maximum accuracy of 82% on mobile phone reviews. |
Copied to clipboard
| Challenge: | Abstractive summarization models (LLMs) have demonstrated impressive performance in various tasks, but they are still suffering from factual inconsistency problem called hallucination. |
| Approach: | They propose to improve the faithfulness of large language models by impelling them to process the entire article more fairly and faithfully. |
| Outcome: | The proposed strategy improves the faithfulness of large language models in summarization while maintaining their fluency and informativeness. |
Copied to clipboard
| Challenge: | Experimental results show that DMsECN outperforms existing models for document clustering . |
| Approach: | They propose a multi-view document clustering model with a processor and hybrid module . they demonstrate that DMsECN outperforms existing models by creating a consensus structure from multiple clustering structures. |
| Outcome: | The proposed model outperforms existing models on four multi-view document clustering datasets. |
Copied to clipboard
| Challenge: | Existing approaches to consolidate textual inputs are difficult to implement . a recent study aims to capture content overlap by combining multiple textual elements . |
| Approach: | They propose to align predicate-argument relations across texts to represent content overlap . their setting exploits QA-SRL, utilizing question-answer pairs to capture predicates . |
| Outcome: | The proposed task captures content overlap beyond lexical similarity and complements cross-document coreference with proposition-level links, offering potential use for downstream tasks. |
Copied to clipboard
| Challenge: | Empirical evidence suggests that simulated internal beliefs or knowledge can be extracted from language models but such methods require labels, which in some domains may not be readily provided due to human biases or because humans simply do not know the correct label. |
| Approach: | They propose a method to minimize the impact of unrelated features in activation space by clustering and normalizing activations of contrast pairs before applying unsupervised probing techniques. |
| Outcome: | Empirical evidence suggests that simulated internal beliefs or knowledge can be extracted from language model activations without human labels. |
Copied to clipboard
| Challenge: | Existing work on text summarization approaches are approaching or exceeding human excellence . |
| Approach: | They propose a framework that optimizes the balance between information volume and entropy in input texts. |
| Outcome: | The proposed framework optimizes information volume and entropy in input texts, achieving notable improvements in localized contexts. |
Copied to clipboard
| Challenge: | Existing methods for intent clustering rely on labeled examples or unsupervised fine-tuning to optimize results for each new dataset. |
| Approach: | They propose a method that uses an embedder to derive an embedding for each utterance and then pool them with the seed to improve the embeddable results. |
| Outcome: | The proposed method outperforms embedding methods and is comparable to state-of-the-art methods. |
Copied to clipboard
| Challenge: | Named Entity Recognition (NER) tasks are becoming more challenging due to the introduction of complex tagsets, which often leads to the failure of existing NER systems in accurately recognizing these entities. |
| Approach: | They propose a novel attack which relies on disentanglement and word attribution techniques to learn an embedding and identifying important words across both components. |
| Outcome: | The proposed approach improves the F1 score over the original LLM model by 8% and 18% on CoNLL-2003 and Ontonotes 5.0 datasets respectively. |
Copied to clipboard
| Challenge: | Similar single-label XWS settings cannot be easily adapted for multi-l label classification. |
| Approach: | They propose a novel method for open-world multi-label text classification under extremely weak supervision where the user provides a brief description without any labels or ground-truth label space. |
| Outcome: | The proposed method exhibits a remarkable increase in ground-truth label space coverage on various datasets. |
Copied to clipboard
| Challenge: | Paraphrasing is an important aspect of natural-language generation that can produce more variety in the way specific content is presented. |
| Approach: | They propose to use contextual paraphrasing to capture the meaning of a sentence while performing dialogue act clustering. |
| Outcome: | The proposed task combines paraphrases with dialogue act clustering to capture such contextual paraphrasing. |
Copied to clipboard
| Challenge: | Extensive experiments on 14 datasets show that ClusterLLM consistently improves clustering quality, at an average cost of $0.6 per dataset. |
| Approach: | They propose a text clustering framework that leverages feedback from an instruction-tuned large language model, such as ChatGPT. |
| Outcome: | Extensive experiments on 14 datasets show that ClusterLLM consistently improves clustering quality, at an average cost of $0.6 per dataset. |
Copied to clipboard
| Challenge: | Existing intent discovery methods focus on transferring prior knowledge of known intents to unknown ones. |
| Approach: | They propose to use frame knowledge as conceptual semantic guidance to bridge the gap between known intents representation learning and unknown intents clustering. |
| Outcome: | The proposed method outperforms solid baselines on two benchmark datasets. |
Copied to clipboard
| Challenge: | Experimental results show that topic modeling is competitive compared to closed-source methods. |
| Approach: | They propose a semi-supervised topic modeling method that combines LLMs with clustering to improve topic generation and distribution. |
| Outcome: | The proposed method outperforms state-of-the-art methods that utilize GPT-4 on topic alignment and exhibits competitive performance compared to Neural Topic Models on topic quality. |
Copied to clipboard
| Challenge: | Large language models (LLMs) have shown impressive performance on downstream tasks, but if they cannot be fully described in prompts, they could fail to perform the task. |
| Approach: | They propose a method to contextualize a task toward a large language model (LLM) they use open-ended zero-shot inference from the entire dataset to aggregate the inference results and incorporate the aggregated meta-information for the actual task. |
| Outcome: | The proposed method improves text clustering tasks and improves on several datasets. |
Copied to clipboard
| Challenge: | Gradient-based data influence approximation is not feasible in practice. |
| Approach: | They propose a gradient-based data selection framework with clustering and a modified Upper Confidence Bound algorithm to solve this problem. |
| Outcome: | The proposed framework can achieve comparable results to the original gradient-based data selection methods while reducing computational consumption. |
Copied to clipboard
| Challenge: | Text embeddings are a fundamental component in many NLP tasks, but their interpretation and explanation remain challenging. |
| Approach: | They propose a framework for interpretable text embeddings and text similarity explanation . they characterize the main ideas, approaches, and trade-offs and discuss lessons learned . |
| Outcome: | The proposed methods are compared with existing models and compare them with existing ones. |
Copied to clipboard
| Challenge: | Existing frameworks for QA with large language models are difficult to implement due to noise, limited context length and latency. |
| Approach: | They propose a model-agnostic framework to address problems in QA with large language models. |
| Outcome: | The proposed framework reduces noise in the ASR output and the limited context length of LLMs and improves performance on the widely used Spoken-SQuAD dataset. |
Copied to clipboard
| Challenge: | Traditional topic modeling treats each document as a single, coherent unit of topic. |
| Approach: | They propose a paradigm that redefines topic assignment at the level of segments . they propose 'segment intrusion task' to extend word intrusion to the span level . |
| Outcome: | The proposed paradigm improves topic purity, interpretability and applicability to multi-theme corpora. |
Copied to clipboard
| Challenge: | Document clustering does not inherently ensure thematic consistency. |
| Approach: | They propose a framework that constructs a similarity graph over document embeddings and applies iterative graph-based clustering algorithms to partition the corpus into initial clusters. |
| Outcome: | The proposed framework constructs a similarity graph over document embeddings and applies iterative graph-based clustering algorithms to partition the corpus into initial clusters. |
Copied to clipboard
| Challenge: | Existing methods for scaling test-time computation rely on external models that introduce substantial computational overhead and fail to capture context-aware semantics. |
| Approach: | They propose a method that leverages the generator LLM’s internal hidden states for clustering, eliminating the need for external models. |
| Outcome: | The proposed method improves the computational efficiency of test-time scaling while maintaining or exceeding the performance of existing methods. |
Copied to clipboard
| Challenge: | Prompt-based text embedding models generate task-specific embeddables but have thousands of dimensions . dimensionality reductions for embedded text can result in performance degradations of only the first 25% of the dimensions resulting in a very small degradation . |
| Approach: | They investigate how post-hoc dimensionality reduction affects performance of various tasks . they find that embeddings for classification and clustering exhibit lower intrinsic dimensionalities . |
| Outcome: | The proposed model generates task-specific embeddings upon receiving tailored prompts, but has thousands of dimensions and high storage costs. |
Copied to clipboard
| Challenge: | Sticky tokens, when repeatedly inserted into sentences, pull sentence similarity toward a certain value, disrupting the normal distribution of embedding distances and degrading downstream performance. |
| Approach: | They propose a method to detect “sticky tokens” by sentence and token filtering and apply it to 40 checkpoints across 14 model families. |
| Outcome: | The proposed method detects 868 sticky tokens across 14 models and shows that their presence does not correlate with model size or vocabulary size. |
Copied to clipboard
| Challenge: | In text embeddings from PLMs are essential for many NLP applications, but performance degrades on longer texts. |
| Approach: | They propose a method which mitigates the phenomenon of Length Collapse . they propose TempScale to ensure more consistent embeddings across different text lengths . |
| Outcome: | The proposed method improves performance on MTEB and LongEmbed by 0.94% on short and 1.10% on long texts. |
Copied to clipboard
| Challenge: | Existing tools for lemmatization in morphologically rich languages with ambiguous orthography face inconsistent standards and limited genre coverage. |
| Approach: | They propose two new approaches that frame lemmatization as classification into a Lemma-POS-Gloss tagset, leveraging machine translation and semantic clustering. |
| Outcome: | The proposed models perform better than existing models and are more interpretable, the authors show. |
Copied to clipboard
| Challenge: | Existing methods for integrating knowledge graphs with large language models lack continuous learning capabilities. |
| Approach: | They propose an agent framework with a dynamic, evolvable memory mechanism specifically designed for KG reasoning. |
| Outcome: | EvoMemKG achieves state-of-the-art performance without training or tools . it achieves improvements of up to 20% over baseline on multi-hop queries . |
Copied to clipboard
| Challenge: | Misinformation evolves as it spreads, shifting in language, framing, and moral emphasis to adapt to new audiences. |
| Approach: | They propose a multi-round, persona-conditioned framework that simulates how claims are iteratively reinterpreted by agents with distinct ideological perspectives. |
| Outcome: | The proposed framework generates persona-specific claims across multiple rounds . it is based on an uncensored large language model and is scalable to multiple tasks . |
Copied to clipboard
| Challenge: | Retrieval-augmented generation (RAG) enables large language models to incorporate external knowledge at inference. |
| Approach: | They propose a hierarchical framework that organizes triplets into subtopics and topics to enhance connectivity and integrate dispersed information. |
| Outcome: | Experiments on abstractive and specific QA benchmarks show that TH-RAG outperforms strong baselines in accuracy and robustness while remaining efficient. |
Copied to clipboard
| Challenge: | Text embeddings are used in many NLP tasks, including document clustering, semantic search, question answering, and classification. |
| Approach: | They introduce the Polish Massive Text Embedding Benchmark (PL-MTEB) it is a comprehensive benchmark for text embeddings in the Polish language. |
| Outcome: | The proposed model is based on 30 different NLP tasks in the Polish language. |
Copied to clipboard
| Challenge: | Existing QA benchmarks do not explicitly support document-grounded related insight generation . Existing document-based QA efforts focus on answering fact-based questions . |
| Approach: | They propose a task to generate additional insights from a document collection that improves, extends or rethinks an initial answer to an open-ended question. |
| Outcome: | The proposed task improves, extends, or rethinks an answer to an open-ended question. |