Papers by Didier Schwab
Providing Semantic Knowledge to a Set of Pictograms for People with Disabilities: a Set of Links between WordNet and Arasaac: Arasaac-WN (2020.lrec-1)
Copied to clipboard
Didier Schwab, Pauline Trial, Céline Vaschalde, Loïc Vial, Emmanuelle Esperanca-Rodier, Benjamin Lecouteux
| Challenge: | Pictograms are a tool that is increasingly used by people with cognitive or communication disabilities. |
| Approach: | They propose a database that links WordNet and Arasaac to link pictograms to semantic knowledge. |
| Outcome: | The proposed database links pictograms with WordNet and Arasaac to create language-independent prototypes. |
FlauBERT: Unsupervised Language Model Pre-training for French (2020.lrec-1)
Copied to clipboard
Hang Le, Loïc Vial, Jibril Frej, Vincent Segonne, Maximin Coavoux, Benjamin Lecouteux, Alexandre Allauzen, Benoit Crabbé, Laurent Besacier, Didier Schwab
| Challenge: | Language models are a key step to achieve state-of-the-art results in many different Natural Language Processing (NLP) tasks. |
| Approach: | They propose to use a language model that is pre-trained on a large and heterogeneous French corpus to train continuous word representations. |
| Outcome: | The proposed model outperforms existing models on a large and heterogeneous French corpus. |
Do Multilingual Neural Machine Translation Models Contain Language Pair Specific Attention Heads? (2021.findings-acl)
Copied to clipboard
| Challenge: | Recent studies on multilingual representations focus on whether there is an emergence of language-independent representations or whether multilingual models partition their weights among different languages. |
| Approach: | They analyze encoder self-attention and encoder-decoder attention heads in a multilingual neural translation model. |
| Outcome: | The proposed model is based on a multilingual neural translation model with a language-independent representation. |
Should Cross-Lingual AMR Parsing go Meta? An Empirical Assessment of Meta-Learning and Joint Learning AMR Parsing (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Cross-lingual AMR parsing is a task of predicting AMR graphs in a target language when training data is available only in . et al. (2018) evaluated meta-learning for cross-lingual parse in Croatian, Farsi, Korean, Chinese, and French. |
| Approach: | They propose to use meta-learning to tackle cross-lingual AMR parsing in a target language . they evaluate their models in k-shot scenarios and compare them to classical joint learning . |
| Outcome: | The proposed model performs better in 0-shot evaluation for Croatian, Farsi, Korean, Chinese, and French. |
A Multimodal French Corpus of Aligned Speech, Text, and Pictogram Sequences for Speech-to-Pictogram Machine Translation (2024.lrec-main)
Copied to clipboard
Cécile Macaire, Chloé Dion, Jordan Arrigo, Claire Lemaire, Emmanuelle Esperança-Rodier, Benjamin Lecouteux, Didier Schwab
| Challenge: | Existing algorithms for the automatic translation of spoken language into pictogram units are lacking for language impairments. |
| Approach: | They propose to use a French dataset that contains 230 hours of speech resources to create a rule-based pictogram grammar with a restricted vocabulary and a discussion of strategic decisions. |
| Outcome: | The proposed model is validated through multiple post-editing phases by expert annotators and is freely available under a non-commercial licence. |
WIKIR: A Python Toolkit for Building a Large-scale Wikipedia-based English Information Retrieval Dataset (2020.lrec-1)
Copied to clipboard
| Challenge: | ad-hoc information retrieval methods usually require large amounts of annotated data to be effective. |
| Approach: | They propose an open-source toolkit to automatically build large-scale English information retrieval datasets based on Wikipedia. |
| Outcome: | The proposed toolkit builds large-scale English information retrieval datasets based on Wikipedia with 59,252 queries and 2,617,003 pairs. |
Limitations of Human Identification of Automatically Generated Text (2024.lrec-main)
Copied to clipboard
Nadège Alavoine, Maximin Coavoux, Emmanuelle Esperança-Rodier, Romane Gallienne, Carlos Gonzalez Gallardo, Jérôme Goulian, Jose G. Moreno, Aurélie Névéol, Didier Schwab, Vincent Segonne, Johanna Simoens
| Challenge: | Neural text generation tools such as ChatGPT are gaining popularity . human annotations are considered gold standard labels for multiple tasks . |
| Approach: | They propose a new corpus in French and English for recognising automatically generated texts . they propose 'incontext' setup which makes explicit the interaction between two parties . |
| Outcome: | The proposed model generates fluent text, which requires much closer reading than the current model. |
Lightweight Adapter Tuning for Multilingual Speech Translation (2021.acl-short)
Copied to clipboard
| Challenge: | Adapter tuning is an efficient alternative to fine-tuning in NLP . a multilingual model could be outperformed by its bilingual counterparts . |
| Approach: | They propose to use adapter tuning to optimize for multilingual speech translation . they use pre-trained models to freeze pre-train parameters and inject lightweight modules . |
| Outcome: | The proposed adapters can specialize to specific language pairs with low extra cost . the proposed models outperform bilingual models on high-resource language pairs . |
UFSAC: Unification of Sense Annotated Corpora and Tools (L18-1)
Copied to clipboard
| Challenge: | a dozen sense annotated English corpora are used in Word Sense Disambiguation (WSD) a new format of corpus is proposed that can be used for training or testing a disambiguation system . |
| Approach: | They propose a format of corpus that can be used for training or testing a disambiguation system . they provide the source code and a complete Java API for manipulating corpora in this format . |
| Outcome: | The proposed format of corpus can be used for training or testing a disambiguation system . the source code and a complete Java API are provided for building the corpus . |
Automatic Speech Recognition and Query By Example for Creole Languages Documentation (2022.findings-acl)
Copied to clipboard
| Challenge: | CREAM project aims to provide linguists with new methods for language documentation based on automatic speech recognition and keyword-spotting. |
| Approach: | They propose to use one hour of annotated data to design an automatic speech recognition system for two Creole languages. |
| Outcome: | The proposed model is based on an hour of annotated data and is usable by linguists. |
What Matters to an LLM? Behavioral and Computational Evidences from Summarization (2026.findings-eacl)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are increasingly entrusted with the management of information. |
| Approach: | They combine behavioral and computational analyses to find out what LLMs prioritize . they generate length-controlled summaries and derive empirical importance distributions . |
| Outcome: | The proposed model converges on consistent importance patterns and clusters more by family than by size. |
Jargon: A Suite of Language Models and Evaluation Tasks for French Specialized Domains (2024.lrec-main)
Copied to clipboard
Vincent Segonne, Aidan Mannion, Laura Cristina Alonzo Canul, Alexandre Daniel Audibert, Xingyu Liu, Cécile Macaire, Adrien Pupier, Yongxin Zhou, Mathilde Aguiar, Felix E. Herron, Magali Norré, Massih R Amini, Pierrette Bouillon, Iris Eshkol-Taravella, Emmanuelle Esperança-Rodier, Thomas François, Lorraine Goeuriot, Jérôme Goulian, Mathieu Lafourcade, Benjamin Lecouteux, François Portet, Fabien Ringeval, Vincent Vandeghinste, Maximin Coavoux, Marco Dinarelli, Didier Schwab
| Challenge: | Pretrained language models are the de facto backbone of most state-of-the-art NLP systems. |
| Approach: | They propose a family of domain-specific pretrained PLMs for French focusing on three important domains: transcribed speech, medicine, and law. |
| Outcome: | The proposed models perform better on transcribed speech, medicine, and law domains than state-of-the-art models on a diverse set of tasks and datasets. |
Dual-decoder Transformer for Joint Automatic Speech Recognition and Multilingual Speech Translation (2020.coling-main)
Copied to clipboard
| Challenge: | Existing models for automatic speech recognition and multilingual speech translation are on par with cascade counterparts. |
| Approach: | They propose a dual-decoder Transformer architecture that performs automatic speech recognition and multilingual speech translation. |
| Outcome: | The proposed models outperform the previously-reported highest translation performance in multilingual settings and bilingual one-to-one results. |