Challenge: Existing work on euphemism disambiguation tasks has focused on transformers . euphorias are expressions that soften the message they convey, therefore dictionary-based approaches are ineffective .
Approach: They propose to annotate PETs for vagueness and use transformers to classify PETs . they perform euphemism disambiguation experiments in three different languages .
Outcome: The proposed models perform well in English euphemism disambiguation task . preliminary results will be used to launch future work .

Similar Papers

MASSIVE Multilingual Abstract Meaning Representation: A Dataset and Baselines for Hallucination Detection (2024.starsem-1)

Copied to clipboard

Challenge: Abstract Meaning Representation (AMR) is a semantic formalism that captures the core meaning of an utterance.
Approach: They propose to use AMR to map meanings of 1,685 utterances to 50+ languages to build a dataset 20 times larger than existing resources.
Outcome: The proposed dataset covers more languages, has more utterances, and has localized or translated entities for each language.
Inducing Language-Agnostic Multilingual Representations (2021.starsem-1)

Copied to clipboard

Challenge: Cross-lingual representations have the potential to make NLP techniques available to the vast majority of languages in the world, but they currently require large pretraining corpora or access to typologically similar languages.
Approach: They propose to remove language identity signals from multilingual embeddings by re-aligning vector spaces of target languages to a pivot source language and removing language-specific means and variances.
Outcome: The proposed approaches reduce cross-lingual transfer gap by 8.9 points (m-BERT) and 18.2 points (XLM-R) on average across all tasks and languages.
VOLIMET: A Parallel Corpus of Literal and Metaphorical Verb-Object Pairs for English–German and English–French (2024.starsem-1)

Copied to clipboard

Challenge: Metaphorical language is a complex interplay of cultural and linguistic elements that characterizes metaphorical language . a corpus of parallel sentences containing gold standard alignments of metaphorical verb-object pairs and literal paraphrases is presented .
Approach: They propose to analyze metaphorical verb-object pairs and literal paraphrases in parallel sentences from English to German and French.
Outcome: The proposed corpus of 2,916 parallel sentences reveals monolingual patterns for metaphorical vs. literal uses in English . cross-lingually, the results show a rich variability in translations as well as different behaviors for the two target languages .
Compound or Term Features? Analyzing Salience in Predicting the Difficulty of German Noun Compounds across Domains (2021.starsem-1)

Copied to clipboard

Challenge: Using domain-specific vocabulary, it is important to analyse domain-related characteristics to improve the communication between lay people and experts.
Approach: They focus on the interaction of compound-based lexical features (such as frequency and productivity) and terminology-based features (contrasting domain-specific and general language) across word representations and classifiers.
Outcome: The proposed model shows that the interaction of compound-based lexical features and terminology-based features across word representations and classifiers is important for a broad binary distinction into ‘easy’ vs. ‘difficult’ general-language compound frequency is sufficient, but for . a more fine-grained four-class distinction it is crucial to include contrastive termhood features and compound and constituent features.
Token Sequence Labeling vs. Clause Classification for English Emotion Stimulus Detection (2020.starsem-1)

Copied to clipboard

Challenge: Emotion stimulus detection is the task of finding the cause of an emotion in a textual description.
Approach: They propose to evaluate whether clause classification or token sequence labeling is better for emotion stimulus detection in English.
Outcome: The proposed framework compares clause classification and token sequence labeling on four English datasets.
Assessing Polyseme Sense Similarity through Co-predication Acceptability and Contextualised Embedding Distance (2020.starsem-1)

Copied to clipboard

Challenge: Co-predication is a commonly used linguistic test to tell apart shifts in polysemic sense from changes in homonymic meaning.
Approach: They examine how co-predication acceptability relates to explicit ratings of polyseme word sense similarity and how well they can be predicted through the distance between target words’ contextualised word embeddings.
Outcome: The proposed measures can be predicted through the distance between target words’ contextualised word embeddings.
DRS Parsing as Sequence Labeling (2022.starsem-1)

Copied to clipboard

Challenge: a new semantic parser for English, German, Italian, and Dutch discourse representation structures is developed . we present a system that maps tokens to finite set of meaning fragments and is more transparent . a comprehensive error analysis highlights areas for future work on semantic parses .
Approach: They propose a fully trainable semantic parser for English, German, Italian, and Dutch discourse representation structures that maps each token to one of a finite set of meaning fragments.
Outcome: The proposed system is more transparent and useful for human-in-the-loop annotations.
Leveraging Three Types of Embeddings from Masked Language Models in Idiom Token Classification (2022.starsem-1)

Copied to clipboard

Challenge: Recent research shows that contextualized word embeddings can give promising results for idiom token classification.
Approach: They propose to leverage contextualized word embeddings from masked language models to improve idiom token classification.
Outcome: The proposed method improves idiom token classification for English and Japanese datasets.
How Are Idioms Processed Inside Transformer Language Models? (2023.starsem-1)

Copied to clipboard

Challenge: idioms are prevalent in natural language, but how do they be processed?
Approach: They analyze the embeddings of idiomatic and literal expressions across all layers of the networks at both the sentence and word levels.
Outcome: The proposed models represent idioms distinctively compared to literal language, the study finds .
Multilingual Extraction and Categorization of Lexical Collocations with Graph-aware Transformers (2022.starsem-1)

Copied to clipboard

Challenge: lexical collocations exhibit varying degrees of frozenness due to their varying degree of frozenncy.
Approach: They propose a sequence tagging BERT-based model enhanced with a graph-aware transformer architecture and evaluate the task of collocation recognition in context.
Outcome: The proposed model encoding syntactic dependencies is useful, and provides insights on differences in collocation typification in English, Spanish and French.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations