Challenge: ambiguity is a major obstacle to providing services based on sentence classification . authors use similarity in a semantic space to detect ambiguities in training data and scenarios .
Approach: They use similarity in a semantic space to detect ambiguities in service scenarios and training data.
Outcome: The proposed approach can detect ambiguities and debug services.

Similar Papers

Probing for Semantic Classes: Diagnosing the Meaning Content of Word Embeddings (P19-1)

Copied to clipboard

Challenge: Empirical analysis of word embeddings of ambiguous words is limited by the small size of manually annotated resources and by the fact that word senses are treated as unrelated individual concepts.
Approach: They present a large dataset based on manual Wikipedia annotations and word senses, where word sense from different words are related by semantic classes.
Outcome: The proposed method can predict whether a word is single-sense or multi-sensor, if the sense is frequent, and it can predict rare senses.
Sentence Meta-Embeddings for Unsupervised Semantic Textual Similarity (2020.acl-main)

Copied to clipboard

Challenge: Existing word embeddings combine complementary strengths of their components to achieve unsupervised semantic similarity (STS).
Approach: They propose to ensemble pre-trained sentence encoders into sentence meta-embeddings to achieve unsupervised Semantic Textual Similarity (STS) they adapt dimensionality reduction, generalized Canonical Correlation Analysis and cross-view auto-encoders to their work.
Outcome: The proposed method achieves 3.7% to 6.4% Pearson’s r over single-source word embeddings on the STS Benchmark and on the StS12-STS16 datasets.
Uncertainty-Aware Contrastive Sentence Embedding With Local Context Representation for Text Classification (2026.findings-acl)

Copied to clipboard

Challenge: Existing models for text classification are based on encoder-only transformers and generative pre-trained transformers.
Approach: They propose an uncertainty-aware contrastive sentence embedding approach that addresses language ambiguity and inter-class separability for a text classification task.
Outcome: The proposed approach improves classification accuracy on public datasets compared with state-of-the-art methods.
ALIGN-SIM: A Task-Free Test Bed for Evaluating and Interpreting Sentence Embeddings through Semantic Similarity Alignment (2024.findings-emnlp)

Copied to clipboard

Challenge: Sentence embeddings play a pivotal role in a wide range of NLP tasks . evaluating and interpreting these dense vectors remains an open challenge to date .
Approach: They propose a task-free test bed for evaluating and interpreting sentence embeddings . they examined five classical and eight LLM-induced sentence embedders based on semantic similarity alignment criteria .
Outcome: The proposed test bed consists of five semantic similarity alignment criteria . it shows that none of the embeddings aligned with the criteria compared to other benchmarks .
What you can cram into a single $&!#* vector: Probing sentence embeddings for linguistic properties (P18-1)

Copied to clipboard

Challenge: a lack of understanding of the properties of sentence embeddings is limiting the use of the techniques.
Approach: They propose 10 probing tasks designed to capture simple linguistic features of sentences . they use three different encoders to train embeddings in eight different ways .
Outcome: The proposed tasks capture key linguistic features of sentences, but they are difficult to infer from them.
Disentangling Meaning and Language Components in Diverse Multilingual Sentence Embeddings (2026.acl-srw)

Copied to clipboard

Challenge: Existing studies have reported language specificity in multilingual sentence embeddings, resulting in language-specific subspaces.
Approach: They propose to disentangle multilingual sentence embeddings into language-dependent and language-agnostic components to improve cross-lingual similarity estimation.
Outcome: The proposed methods improve cross-lingual similarity estimation across multiple embeddings.
Detecting Ambiguous Utterances in an Intelligent Assistant (2024.emnlp-industry)

Copied to clipboard

Challenge: ambiguous utterances can be interpreted as either chat or task intents in intelligent assistants . ambiguity of intent is particularly noticeable in intelligent devices where task-oriented and non-task-oriented utterrances are mixed and most utterations are short due to characteristics of devices.
Approach: They propose to feed sentence embeddings developed from microblogs and search logs with a self-attention mechanism to detect ambiguous utterances robustly.
Outcome: The proposed model outperforms baselines and a strong LLM-based model.
CASE – Condition-Aware Sentence Embeddings for Conditional Semantic Textual Similarity Measurement (2026.eacl-long)

Copied to clipboard

Challenge: Recent approaches use semantic similarity to improve the quality of sentence embeddings, but it is difficult to measure the similarity between sentences.
Approach: They propose a condition-aware sentence embedding method that uses an LLM encoder to create an embeddable sentence under a given condition.
Outcome: The proposed method improves the performance of LLM-based embeddings and the isotropy of the embeddable space despite requiring a small number of dimensions.
Task-oriented Word Embedding for Text Classification (C18-1)

Copied to clipboard

Challenge: Existing word embeddings only consider contextual information, which is suboptimal when used in various tasks due to a lack of task-specific features.
Approach: They propose a task-oriented word embedding method that regularizes the distribution of words to enable a clear classification boundary.
Outcome: The proposed method outperforms the state-of-the-art methods on a text classification task.
Static Word Embeddings for Sentence Semantic Representation (2025.emnlp-main)

Copied to clipboard

Challenge: Existing methods to learn fixed-length embeddings for sentence semantics require large computational cost, making it difficult to process billions of sentences cost-efficiently or deploy models on resource-constrained devices such as smartphones.
Approach: They propose to extract word embeddings from a pre-trained Sentence Transformer and improve them with sentence-level principal component analysis followed by knowledge distillation or contrastive learning.
Outcome: The proposed model outperforms existing models on sentence semantic tasks and surpasses a basic Sentence Transformer model (SimCSE) on a text embedding benchmark.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations