Towards a Variability Measure for Multiword Expressions (N18-2)

Copied to clipboard

Challenge: Multiword expressions (MWEs) are groups of words whose meaning does not derive from the meaning of their components and from their syntactic structure in a regular way.
Approach: They propose to use a language-independent measure of variability dedicated to verbal MWEs based on syntactic and discontinuity-related clues to assess its relevance with respect to a linguistic benchmark.
Outcome: The proposed measure is useful for VMWE classification and variant identification on a French corpus.

Similar Papers

Verbal Multiword Expression Identification: Do We Need a Sledgehammer to Crack a Nut? (2020.coling-main)

Copied to clipboard

Challenge: Multiword expressions (MWEs) are word combinations idiosyncratic with respect to syntax or semantics.
Approach: They propose to use a language-independent system to identify previously seen VMWEs by combining filters to obtain the best averaged F-score over 11 languages and the best score for both seen and unseen VMwes.
Outcome: The proposed system obtains the best averaged F-score over 11 languages and even the best score for both seen and unseen VMWEs due to the high proportion of seen VMwes in texts.
If you’ve seen some, you’ve seen them all: Identifying variants of multiword expressions (C18-1)

Copied to clipboard

Challenge: Multiword expressions (VMWEs) show idiosyncratic variability, which is challenging for NLP applications.
Approach: They propose to use a model to identify variants of previously seen VMWEs by comparing VMWAs with morpho-syntactic variations.
Outcome: The proposed approach outperforms a baseline by 4 percent points of F-measure on a French corpus.
CoAM: Corpus of All-Type Multiword Expressions (2025.acl-long)

Copied to clipboard

Challenge: Existing datasets for multiword expressions are inconsistently annotated, limited to a single type of MWE, or limited in size.
Approach: They propose to use a new interface to generate MWE annotations for the first time in a dataset of MWE identification.
Outcome: The proposed model outperforms existing models on the DiMSUM dataset.
Construction of Large-scale English Verbal Multiword Expression Annotated Corpus (L18-1)

Copied to clipboard

Challenge: In this paper, we focus on verbal MWEs, whose accurate recognition is challenging because they could be discontinuous.
Approach: They conduct large-scale annotations of VMWEs on the Wall Street Journal portion of Ontonotes . they first construct a VMwe dictionary based on the english-language Wiktionary .
Outcome: The proposed resource annotates 7,833 VMWE instances belonging to various categories . the authors hope the results will help to develop models for MWE recognition and dependency parsing .
Evaluating the Impact of Verbal Multiword Expressions on Machine Translation (2026.acl-long)

Copied to clipboard

Challenge: Verbal multiword expressions (VMWEs) are difficult for machine translation because their meanings are often not recoverable from their component words.
Approach: They analyze the impact of verbal idioms, verb-particle constructions, and light verb constructions on machine translation quality from English to multiple languages.
Outcome: The proposed system improves translation quality by focusing on verb idioms, verb-particle constructions and light verb constructions.
Vector Spaces for Quantifying Disparity of Multiword Expressions in Annotated Text (2024.acl-srw)

Copied to clipboard

Challenge: We show that multiword expressions are a good study for linguistic diversity due to theiridiosyncratic nature.
Approach: They train static MWE-aware word embeddings for verbal MWEs in 14 languages . they find that the disparity measure aggregatingthem at a global scale correlates with the number of types .
Outcome: The proposed method is based on a set of vector spaces for VMWEs in 14 languages.
Representing Multiword Term Variation in a Terminological Knowledge Base: a Corpus-Based Study (2020.lrec-1)

Copied to clipboard

Challenge: Multiword terms are the most frequent type of lexical units in scientific and technical communication. rendering them in another language is not easy due to their cognitive complexity, proliferation of different forms, and their unsystematic representation in terminographic resources.
Approach: They evaluated Spanish translation variants of multiword terms in three parallel corpora, two comparable corporales and two terminological resources.
Outcome: The results show that multiword terms exhibit a significant degree of term variation . the proposed model is based on a set of criteria for determining which variants should be selected .
Detecting Multiword Expression Type Helps Lexical Complexity Assessment (2020.lrec-1)

Copied to clipboard

Challenge: Multiword expressions (MWEs) represent lexemes that should be treated as single lexical units due to their idiosyncratic nature.
Approach: They re-annotate a complex word identification shared task 2018 dataset . they find that a lexical complexity assessment system benefits from the information .
Outcome: The proposed dataset provides valuable information for the text simplification community.
Verbal Multiword Expressions for Identification of Metaphor (2020.acl-main)

Copied to clipboard

Challenge: Metaphor is a linguistic device in which a concept is expressed by mentioning another . Verbal MWEs are examples of non-literal language in which multiple words form a single unit of meaning.
Approach: They propose to analyze the interplay between metaphor and multiword expressions processing by informing the model of the presence of MWEs.
Outcome: The proposed architecture reach state-of-the-art on two established metaphor datasets.
MultiMWE: Building a Multi-lingual Multi-Word Expression (MWE) Parallel Corpora (2020.lrec-1)

Copied to clipboard

Challenge: Existing bilingual or multi-lingual MWE corpora are limited for multilingual use . only 871 pairs of English-German MWEs are available for research .
Approach: They present a collection of bilingual and multi-lingual MWEs extracted from parallel corpora.
Outcome: The available bilingual or multi-lingual MWE corpus is very limited . the collection is a small collection of 871 pairs of English-German MWEs .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations