Papers by Agata Savary
Cross-type French Multiword Expression Identification with Pre-trained Masked Language Models (2024.lrec-main)
Copied to clipboard
| Challenge: | Multiword expressions (MWEs) have linguistic features that distinguish them from regular word groupings. |
| Approach: | They propose a combination of two systems that learn verbal multiword expressions and non-verbal MWEs to improve performance on a cross-type dataset . |
| Outcome: | The proposed system improves the F1 score on a french treebank with VMWEs and nVMWES training data. |
Evaluating Diversity of Multiword Expressions in Annotated Text (2022.coling-1)
Copied to clipboard
| Challenge: | Using the extensive formalization and measures of diversity developed in ecology, we evaluate the variety and balance of multiword expression annotation produced by automatic annotation systems. |
| Approach: | They propose to use the formalization and measures of diversity developed in ecology to evaluate the variety and balance of multiword expression annotation produced by automatic annotation systems. |
| Outcome: | The proposed measures validate or invalidate their pertinence for multiword expressions in annotated texts. |
Annotating Verbal Multiword Expressions in Arabic: Assessing the Validity of a Multilingual Annotation Procedure (2022.lrec-1)
Copied to clipboard
Najet Hadj Mohamed, Cherifa Ben Khelil, Agata Savary, Iskandar Keskes, Jean-Yves Antoine, Lamia Hadrich-Belguith
| Challenge: | a subset of 1,062 sentences from the Prague Arabic Dependency Treebank PADT were selected and annotated by two Arabic native speakers independently. |
| Approach: | They propose to use Arabic as an annotation framework to extend PARSEME to modern standard Arabic by measuring inter-annotator agreement. |
| Outcome: | The proposed framework is based on a subset of 1,062 sentences from the Prague Arabic Dependency Treebank PADT and is already exceeding the smallest corpus of the PARSEME suite. |
Verbal Multiword Expression Identification: Do We Need a Sledgehammer to Crack a Nut? (2020.coling-main)
Copied to clipboard
| Challenge: | Multiword expressions (MWEs) are word combinations idiosyncratic with respect to syntax or semantics. |
| Approach: | They propose to use a language-independent system to identify previously seen VMWEs by combining filters to obtain the best averaged F-score over 11 languages and the best score for both seen and unseen VMwes. |
| Outcome: | The proposed system obtains the best averaged F-score over 11 languages and even the best score for both seen and unseen VMWEs due to the high proportion of seen VMwes in texts. |
A Diversity Diet for a Healthier Model: A Case Study of French ModernBERT (2026.findings-acl)
Copied to clipboard
| Challenge: | Diversity has been gaining interest in the NLP community in recent years. |
| Approach: | They propose to use diversity-driven sampling to pre-train models on French with a fixed compute budget. |
| Outcome: | The diversity-driven sampling reduces the pre-training dataset by 94% and the pretraining time by 73% while maintaining comparable performance. |
If you’ve seen some, you’ve seen them all: Identifying variants of multiword expressions (C18-1)
Copied to clipboard
| Challenge: | Multiword expressions (VMWEs) show idiosyncratic variability, which is challenging for NLP applications. |
| Approach: | They propose to use a model to identify variants of previously seen VMWEs by comparing VMWAs with morpho-syntactic variations. |
| Outcome: | The proposed approach outperforms a baseline by 4 percent points of F-measure on a French corpus. |
Vector Spaces for Quantifying Disparity of Multiword Expressions in Annotated Text (2024.acl-srw)
Copied to clipboard
| Challenge: | We show that multiword expressions are a good study for linguistic diversity due to theiridiosyncratic nature. |
| Approach: | They train static MWE-aware word embeddings for verbal MWEs in 14 languages . they find that the disparity measure aggregatingthem at a global scale correlates with the number of types . |
| Outcome: | The proposed method is based on a set of vector spaces for VMWEs in 14 languages. |
Towards a Variability Measure for Multiword Expressions (N18-2)
Copied to clipboard
| Challenge: | Multiword expressions (MWEs) are groups of words whose meaning does not derive from the meaning of their components and from their syntactic structure in a regular way. |
| Approach: | They propose to use a language-independent measure of variability dedicated to verbal MWEs based on syntactic and discontinuity-related clues to assess its relevance with respect to a linguistic benchmark. |
| Outcome: | The proposed measure is useful for VMWE classification and variant identification on a French corpus. |