Empirical Linguistic Study of Sentence Embeddings (P19-1)

Copied to clipboard

Challenge: a new method of analysing sentence embeddings shows that linguistic information is retained in the vector representations of sentences.
Approach: They propose a method of analysing the content of sentence embeddings based on probing tasks and contrasting languages.
Outcome: The proposed method is based on probing tasks and classification datasets for two contrasting languages.

Similar Papers

What you can cram into a single $&!#* vector: Probing sentence embeddings for linguistic properties (P18-1)

Copied to clipboard

Challenge: a lack of understanding of the properties of sentence embeddings is limiting the use of the techniques.
Approach: They propose 10 probing tasks designed to capture simple linguistic features of sentences . they use three different encoders to train embeddings in eight different ways .
Outcome: The proposed tasks capture key linguistic features of sentences, but they are difficult to infer from them.
Exploring Semantic Properties of Sentence Embeddings (P18-2)

Copied to clipboard

Challenge: Neural vector representations are ubiquitous throughout all subfields of natural language processing.
Approach: They propose a framework that generates triplets of sentences to explore how changes in the syntactic structure or semantics of a given sentence affect their similarity.
Outcome: The proposed framework generates triplets of sentences to explore how changes in the syntactic structure or semantics of a given sentence affect the similarities obtained between their embeddings.
Evaluation of Sentence Representations in Polish (2020.lrec-1)

Copied to clipboard

Challenge: Existing methods for learning sentence representations have been limited in low-resource languages such as Polish .
Approach: They propose two new Polish datasets for evaluating sentence embeddings and evaluate eight different methods including Polish and multilingual models.
Outcome: The proposed methods show strengths and weaknesses in Polish and multilingual models.
Are the Best Multilingual Document Embeddings simply Based on Sentence Embeddings? (2023.findings-eacl)

Copied to clipboard

Challenge: obtaining document embeddings at document level is challenging due to computational requirements and lack of appropriate data.
Approach: They compare methods to produce document-level representations from sentences based on LASER, LaBSE, and Sentence BERT pre-trained multilingual models.
Outcome: The proposed methods produce document-level representations from sentences in 8 languages . the results show that a clever combination of sentence embeddings is usually better than encoding the full document as a single unit.
Disentangling Meaning and Language Components in Diverse Multilingual Sentence Embeddings (2026.acl-srw)

Copied to clipboard

Challenge: Existing studies have reported language specificity in multilingual sentence embeddings, resulting in language-specific subspaces.
Approach: They propose to disentangle multilingual sentence embeddings into language-dependent and language-agnostic components to improve cross-lingual similarity estimation.
Outcome: The proposed methods improve cross-lingual similarity estimation across multiple embeddings.
Probing Multimodal Embeddings for Linguistic Properties: the Visual-Semantic Case (2020.coling-main)

Copied to clipboard

Challenge: Semantic embeddings have advanced the state of the art for natural language processing tasks . but their inner workings are poorly understood and there is a shortage of analysis tools .
Approach: They propose to extend visual-semantic embeddings to multimodal domains by defining probing tasks for embeddable image-caption pairs and testing them with classifiers.
Outcome: The proposed probing tasks show up to 16% more accurate on visual-semantic embeddings compared to unimodal embedders . the proposed extensions to multimodal domains have been lauded as promising in natural language processing .
Sentence Analogies: Linguistic Regularities in Sentence Embeddings (2020.coling-main)

Copied to clipboard

Challenge: Word vectors are often evaluated by assessing to what degree they exhibit regularities with regard to relationships considered in word analogies.
Approach: They propose a number of schemes to induce evaluation data based on lexical analogy data as well as semantic relationships between sentences.
Outcome: The proposed models reflect regularities in lexical analogies and semantic relationships between sentences.
Analyzing the Surprising Variability in Word Embedding Stability Across Languages (2021.emnlp-main)

Copied to clipboard

Challenge: Word embeddings are powerful representations that form the foundation of many natural language processing architectures.
Approach: They explore word embedding stability in a wide range of languages to gain insight into their stability.
Outcome: The proposed results provide insights into word embedding stability in English and other languages.
Static Word Embeddings for Sentence Semantic Representation (2025.emnlp-main)

Copied to clipboard

Challenge: Existing methods to learn fixed-length embeddings for sentence semantics require large computational cost, making it difficult to process billions of sentences cost-efficiently or deploy models on resource-constrained devices such as smartphones.
Approach: They propose to extract word embeddings from a pre-trained Sentence Transformer and improve them with sentence-level principal component analysis followed by knowledge distillation or contrastive learning.
Outcome: The proposed model outperforms existing models on sentence semantic tasks and surpasses a basic Sentence Transformer model (SimCSE) on a text embedding benchmark.
Probing for Semantic Classes: Diagnosing the Meaning Content of Word Embeddings (P19-1)

Copied to clipboard

Challenge: Empirical analysis of word embeddings of ambiguous words is limited by the small size of manually annotated resources and by the fact that word senses are treated as unrelated individual concepts.
Approach: They present a large dataset based on manual Wikipedia annotations and word senses, where word sense from different words are related by semantic classes.
Outcome: The proposed method can predict whether a word is single-sense or multi-sensor, if the sense is frequent, and it can predict rare senses.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations