Papers by Benjamin Lecouteux
Providing Semantic Knowledge to a Set of Pictograms for People with Disabilities: a Set of Links between WordNet and Arasaac: Arasaac-WN (2020.lrec-1)
Copied to clipboard
Didier Schwab, Pauline Trial, Céline Vaschalde, Loïc Vial, Emmanuelle Esperanca-Rodier, Benjamin Lecouteux
| Challenge: | Pictograms are a tool that is increasingly used by people with cognitive or communication disabilities. |
| Approach: | They propose a database that links WordNet and Arasaac to link pictograms to semantic knowledge. |
| Outcome: | The proposed database links pictograms with WordNet and Arasaac to create language-independent prototypes. |
FlauBERT: Unsupervised Language Model Pre-training for French (2020.lrec-1)
Copied to clipboard
Hang Le, Loïc Vial, Jibril Frej, Vincent Segonne, Maximin Coavoux, Benjamin Lecouteux, Alexandre Allauzen, Benoit Crabbé, Laurent Besacier, Didier Schwab
| Challenge: | Language models are a key step to achieve state-of-the-art results in many different Natural Language Processing (NLP) tasks. |
| Approach: | They propose to use a language model that is pre-trained on a large and heterogeneous French corpus to train continuous word representations. |
| Outcome: | The proposed model outperforms existing models on a large and heterogeneous French corpus. |
Growing Trees on Sounds: Assessing Strategies for End-to-End Dependency Parsing of Speech (2024.acl-short)
Copied to clipboard
| Challenge: | Direct dependency parsing of the speech signal is proposed as a way of incorporating prosodic information into the parser and bypassing the limitations of a pipeline approach. |
| Approach: | They propose to use graph-based parsing and sequence labeling based parses to integrate prosodic information into the parser and bypass limitations of pipeline approaches. |
| Outcome: | The proposed graph based approach outperforms a pipeline approach on a large treebank of spoken french, despite having 30% fewer parameters. |
A Multimodal French Corpus of Aligned Speech, Text, and Pictogram Sequences for Speech-to-Pictogram Machine Translation (2024.lrec-main)
Copied to clipboard
Cécile Macaire, Chloé Dion, Jordan Arrigo, Claire Lemaire, Emmanuelle Esperança-Rodier, Benjamin Lecouteux, Didier Schwab
| Challenge: | Existing algorithms for the automatic translation of spoken language into pictogram units are lacking for language impairments. |
| Approach: | They propose to use a French dataset that contains 230 hours of speech resources to create a rule-based pictogram grammar with a restricted vocabulary and a discussion of strategic decisions. |
| Outcome: | The proposed model is validated through multiple post-editing phases by expert annotators and is freely available under a non-commercial licence. |
UFSAC: Unification of Sense Annotated Corpora and Tools (L18-1)
Copied to clipboard
| Challenge: | a dozen sense annotated English corpora are used in Word Sense Disambiguation (WSD) a new format of corpus is proposed that can be used for training or testing a disambiguation system . |
| Approach: | They propose a format of corpus that can be used for training or testing a disambiguation system . they provide the source code and a complete Java API for manipulating corpora in this format . |
| Outcome: | The proposed format of corpus can be used for training or testing a disambiguation system . the source code and a complete Java API are provided for building the corpus . |
Automatic Speech Recognition and Query By Example for Creole Languages Documentation (2022.findings-acl)
Copied to clipboard
| Challenge: | CREAM project aims to provide linguists with new methods for language documentation based on automatic speech recognition and keyword-spotting. |
| Approach: | They propose to use one hour of annotated data to design an automatic speech recognition system for two Creole languages. |
| Outcome: | The proposed model is based on an hour of annotated data and is usable by linguists. |
Jargon: A Suite of Language Models and Evaluation Tasks for French Specialized Domains (2024.lrec-main)
Copied to clipboard
Vincent Segonne, Aidan Mannion, Laura Cristina Alonzo Canul, Alexandre Daniel Audibert, Xingyu Liu, Cécile Macaire, Adrien Pupier, Yongxin Zhou, Mathilde Aguiar, Felix E. Herron, Magali Norré, Massih R Amini, Pierrette Bouillon, Iris Eshkol-Taravella, Emmanuelle Esperança-Rodier, Thomas François, Lorraine Goeuriot, Jérôme Goulian, Mathieu Lafourcade, Benjamin Lecouteux, François Portet, Fabien Ringeval, Vincent Vandeghinste, Maximin Coavoux, Marco Dinarelli, Didier Schwab
| Challenge: | Pretrained language models are the de facto backbone of most state-of-the-art NLP systems. |
| Approach: | They propose a family of domain-specific pretrained PLMs for French focusing on three important domains: transcribed speech, medicine, and law. |
| Outcome: | The proposed models perform better on transcribed speech, medicine, and law domains than state-of-the-art models on a diverse set of tasks and datasets. |
What Has LeBenchmark Learnt about French Syntax? (2024.lrec-main)
Copied to clipboard
| Challenge: | Pretrained acoustic models are increasingly used for downstream speech tasks such as automatic speech recognition, speech translation, spoken language understanding or speech parsing. |
| Approach: | They propose to probing a pretrained acoustic model for French for syntactic information using the Orféo treebank. |
| Outcome: | The proposed model is trained on 7k hours of spoken French and obtained reasonable results on tasks that require higher level linguistic knowledge. |