Challenge: Normative studies on modality for English words are relatively common . however, they are limited to a relatively small number of languages and require costly ratings.
Approach: They aim to learn a mapping between word embeddings and modality norms by training on a high-resource language and testing on . monolingual and crosslingual word embeds are used to predict modality association scores .
Outcome: The proposed model predicts modality associations even when trained on an English resource and tested on a completely unseen language.

Similar Papers

Joint Training for Learning Cross-lingual Embeddings with Sub-word Information without Parallel Corpora (2020.starsem-1)

Copied to clipboard

Challenge: Existing methods for learning cross-lingual word embeddings incorporate sub-word information during training.
Approach: They propose a method that incorporates sub-word information during training to learn cross-lingual word embeddings from monolingual data and a bilingual lexicon.
Outcome: The proposed method improves on bilingual lexicon induction, monolingual word similarity, and document classification using low-resource languages.
Inducing Language-Agnostic Multilingual Representations (2021.starsem-1)

Copied to clipboard

Challenge: Cross-lingual representations have the potential to make NLP techniques available to the vast majority of languages in the world, but they currently require large pretraining corpora or access to typologically similar languages.
Approach: They propose to remove language identity signals from multilingual embeddings by re-aligning vector spaces of target languages to a pivot source language and removing language-specific means and variances.
Outcome: The proposed approaches reduce cross-lingual transfer gap by 8.9 points (m-BERT) and 18.2 points (XLM-R) on average across all tasks and languages.
Evaluating a Joint Training Approach for Learning Cross-lingual Embeddings with Sub-word Information without Parallel Corpora on Lower-resource Languages (2021.starsem-1)

Copied to clipboard

Challenge: Cross-lingual word embeddings provide a way for information to be transferred between languages.
Approach: They propose a joint training approach that incorporates sub-word information during training to learn cross-lingual embeddings.
Outcome: The proposed method improves bilingual lexicon induction, especially for out-of-vocabulary words (OOVs) it is able to represent out- of-vocal words (OVs) and is more isomorphic than previous methods.
Assessing Polyseme Sense Similarity through Co-predication Acceptability and Contextualised Embedding Distance (2020.starsem-1)

Copied to clipboard

Challenge: Co-predication is a commonly used linguistic test to tell apart shifts in polysemic sense from changes in homonymic meaning.
Approach: They examine how co-predication acceptability relates to explicit ratings of polyseme word sense similarity and how well they can be predicted through the distance between target words’ contextualised word embeddings.
Outcome: The proposed measures can be predicted through the distance between target words’ contextualised word embeddings.
Representation of Lexical Stylistic Features in Language Models’ Embedding Space (2023.starsem-1)

Copied to clipboard

Challenge: lexical stylistic notions such as complexity, formality, and figurativeness can be identified in pretrained Language Models . static embeddings encode these features more accurately at the level of words and phrases whereas contextualized LMs perform better on sentences.
Approach: They propose to derive a vector representation for stylistic notions from seed pairs . they find that static embeddings encode stylistic features more accurately .
Outcome: The proposed representations can be used to characterize new texts in terms of these dimensions using a small number of seed pairs.
Comparison and Combination of Sentence Embeddings Derived from Different Supervision Signals (2022.starsem-1)

Copied to clipboard

Challenge: Existing methods to derive sentence embeddings have not been well understood what properties are captured in the resulting sentences depending on the supervision signals.
Approach: They propose to combine two types of sentence embedding methods with similar architectures and tasks to investigate their properties.
Outcome: The proposed methods perform better on unsupervised and downstream tasks than the proposed methods on untrained STS tasks and probing tasks.
Leverage Points in Modality Shifts: Comparing Language-only and Multimodal Word Representations (2023.starsem-1)

Copied to clipboard

Challenge: a recent study of the effect of visual grounding on language representations has given a new life to the debate around extractability and quality of semantic information in representations trained solely on textual input.
Approach: They compare word embeddings from vision-and-language models to text-only models . they identify meaning properties and relations that characterize words whose embeddements are most affected by visual grounding .
Outcome: The proposed model differs from text-only models on semantic representations of language . the study is the first large-scale study of the effect of visual grounding on language representations .
When Polysemy Matters: Modeling Semantic Categorization with Word Embeddings (2022.starsem-1)

Copied to clipboard

Challenge: Recent work using word embeddings to model semantic categorization has shown that static models outperform contextual models.
Approach: They consider polysemy as a possible confounding factor in categorization decisions . they compare sense-level embeddings with previously studied static embedds .
Outcome: The proposed model outperforms static models on coarse- and fine-grained categorization tasks.
Disambiguating Emotional Connotations of Words Using Contextualized Word Representations (2024.starsem-1)

Copied to clipboard

Challenge: BERT, RoBERTa, XLNet, and GPT-2 models effectively discern emotional connotations of words, demonstrating superior performance and greater resilience against biases.
Approach: They propose to use contextualized word representations to examine how words can be used to distinguish emotional connotations across contexts.
Outcome: The proposed models show that they can distinguish emotional connotations of words in different contexts.
„Mann“ is to “Donna” as「国王」is to « Reine » Adapting the Analogy Task for Multilingual and Contextual Embeddings (2023.starsem-1)

Copied to clipboard

Challenge: a lack of comparable multilingual benchmarks and a consensual evaluation protocol for contextual models remains an open question.
Approach: They propose a multilingual analogy dataset and evaluate human and contextual embedding performance.
Outcome: The proposed dataset evaluates human and contextual embedding models on the analogy task.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations