Papers by Agata Savary

8 papers
Cross-type French Multiword Expression Identification with Pre-trained Masked Language Models (2024.lrec-main)

Copied to clipboard

Challenge: Multiword expressions (MWEs) have linguistic features that distinguish them from regular word groupings.
Approach: They propose a combination of two systems that learn verbal multiword expressions and non-verbal MWEs to improve performance on a cross-type dataset .
Outcome: The proposed system improves the F1 score on a french treebank with VMWEs and nVMWES training data.
Evaluating Diversity of Multiword Expressions in Annotated Text (2022.coling-1)

Copied to clipboard

Challenge: Using the extensive formalization and measures of diversity developed in ecology, we evaluate the variety and balance of multiword expression annotation produced by automatic annotation systems.
Approach: They propose to use the formalization and measures of diversity developed in ecology to evaluate the variety and balance of multiword expression annotation produced by automatic annotation systems.
Outcome: The proposed measures validate or invalidate their pertinence for multiword expressions in annotated texts.
Annotating Verbal Multiword Expressions in Arabic: Assessing the Validity of a Multilingual Annotation Procedure (2022.lrec-1)

Copied to clipboard

Challenge: a subset of 1,062 sentences from the Prague Arabic Dependency Treebank PADT were selected and annotated by two Arabic native speakers independently.
Approach: They propose to use Arabic as an annotation framework to extend PARSEME to modern standard Arabic by measuring inter-annotator agreement.
Outcome: The proposed framework is based on a subset of 1,062 sentences from the Prague Arabic Dependency Treebank PADT and is already exceeding the smallest corpus of the PARSEME suite.
Verbal Multiword Expression Identification: Do We Need a Sledgehammer to Crack a Nut? (2020.coling-main)

Copied to clipboard

Challenge: Multiword expressions (MWEs) are word combinations idiosyncratic with respect to syntax or semantics.
Approach: They propose to use a language-independent system to identify previously seen VMWEs by combining filters to obtain the best averaged F-score over 11 languages and the best score for both seen and unseen VMwes.
Outcome: The proposed system obtains the best averaged F-score over 11 languages and even the best score for both seen and unseen VMWEs due to the high proportion of seen VMwes in texts.
A Diversity Diet for a Healthier Model: A Case Study of French ModernBERT (2026.findings-acl)

Copied to clipboard

Challenge: Diversity has been gaining interest in the NLP community in recent years.
Approach: They propose to use diversity-driven sampling to pre-train models on French with a fixed compute budget.
Outcome: The diversity-driven sampling reduces the pre-training dataset by 94% and the pretraining time by 73% while maintaining comparable performance.
If you’ve seen some, you’ve seen them all: Identifying variants of multiword expressions (C18-1)

Copied to clipboard

Challenge: Multiword expressions (VMWEs) show idiosyncratic variability, which is challenging for NLP applications.
Approach: They propose to use a model to identify variants of previously seen VMWEs by comparing VMWAs with morpho-syntactic variations.
Outcome: The proposed approach outperforms a baseline by 4 percent points of F-measure on a French corpus.
Vector Spaces for Quantifying Disparity of Multiword Expressions in Annotated Text (2024.acl-srw)

Copied to clipboard

Challenge: We show that multiword expressions are a good study for linguistic diversity due to theiridiosyncratic nature.
Approach: They train static MWE-aware word embeddings for verbal MWEs in 14 languages . they find that the disparity measure aggregatingthem at a global scale correlates with the number of types .
Outcome: The proposed method is based on a set of vector spaces for VMWEs in 14 languages.
Towards a Variability Measure for Multiword Expressions (N18-2)

Copied to clipboard

Challenge: Multiword expressions (MWEs) are groups of words whose meaning does not derive from the meaning of their components and from their syntactic structure in a regular way.
Approach: They propose to use a language-independent measure of variability dedicated to verbal MWEs based on syntactic and discontinuity-related clues to assess its relevance with respect to a linguistic benchmark.
Outcome: The proposed measure is useful for VMWE classification and variant identification on a French corpus.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations