Papers by Mathieu Constant
„Mann“ is to “Donna” as「国王」is to « Reine » Adapting the Analogy Task for Multilingual and Contextual Embeddings (2023.starsem-1)
Copied to clipboard
| Challenge: | a lack of comparable multilingual benchmarks and a consensual evaluation protocol for contextual models remains an open question. |
| Approach: | They propose a multilingual analogy dataset and evaluate human and contextual embedding performance. |
| Outcome: | The proposed dataset evaluates human and contextual embedding models on the analogy task. |
Beyond Model Performance: Can Link Prediction Enrich French Lexical Graphs? (2024.lrec-main)
Copied to clipboard
| Challenge: | lexical resources are essential for the development of NLP systems, but with advances in language models and deep learning, they are increasingly being replaced by web-derived text. |
| Approach: | They propose a resource-centric study of link prediction approaches over French lexical-semantic graphs. |
| Outcome: | The proposed method is more accurate and reliable than previous methods. |
Rigor Mortis: Annotating MWEs with a Gamified Platform (2020.lrec-1)
Copied to clipboard
| Challenge: | gamification of the platform should be improved, in order to attract and retain more players. |
| Approach: | They propose to use a gamified crowdsourcing platform to evaluate the intuition of speakers and then train them to annotate multi-word expressions in French corpora. |
| Outcome: | The proposed platform evaluates the speakers' intuition and trains them to annotate multi-word expressions in French corpora. |
How to Dissect a Muppet: The Structure of Transformer Embedding Spaces (2022.tacl-1)
Copied to clipboard
| Challenge: | Pretrained embeddings based on the Transformer architecture have taken the NLP community by storm . a novel decomposition of Transformer output embeddables is demonstrated . |
| Approach: | They propose to decompose Transformer output embeddings into a sum of vector factors . they show multi-head attentions and feed-forwards are not equally useful in downstream applications . |
| Outcome: | The proposed method outperforms recurrent architectures on a wide variety of tasks. |
Complex Word Identification: A Comparative Study between ChatGPT and a Dedicated Model for This Task (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing methods to assess lexical complexity are used to evaluate the difficulty of vocabulary for language learners. |
| Approach: | They propose to use pre-trained language models to assess the complexity of a word based on its context. |
| Outcome: | The proposed method outperforms the best systems in SemEval-2021. |