Papers by Benjamin Lecouteux

8 papers
Providing Semantic Knowledge to a Set of Pictograms for People with Disabilities: a Set of Links between WordNet and Arasaac: Arasaac-WN (2020.lrec-1)

Copied to clipboard

Challenge: Pictograms are a tool that is increasingly used by people with cognitive or communication disabilities.
Approach: They propose a database that links WordNet and Arasaac to link pictograms to semantic knowledge.
Outcome: The proposed database links pictograms with WordNet and Arasaac to create language-independent prototypes.
FlauBERT: Unsupervised Language Model Pre-training for French (2020.lrec-1)

Copied to clipboard

Challenge: Language models are a key step to achieve state-of-the-art results in many different Natural Language Processing (NLP) tasks.
Approach: They propose to use a language model that is pre-trained on a large and heterogeneous French corpus to train continuous word representations.
Outcome: The proposed model outperforms existing models on a large and heterogeneous French corpus.
Growing Trees on Sounds: Assessing Strategies for End-to-End Dependency Parsing of Speech (2024.acl-short)

Copied to clipboard

Challenge: Direct dependency parsing of the speech signal is proposed as a way of incorporating prosodic information into the parser and bypassing the limitations of a pipeline approach.
Approach: They propose to use graph-based parsing and sequence labeling based parses to integrate prosodic information into the parser and bypass limitations of pipeline approaches.
Outcome: The proposed graph based approach outperforms a pipeline approach on a large treebank of spoken french, despite having 30% fewer parameters.
A Multimodal French Corpus of Aligned Speech, Text, and Pictogram Sequences for Speech-to-Pictogram Machine Translation (2024.lrec-main)

Copied to clipboard

Challenge: Existing algorithms for the automatic translation of spoken language into pictogram units are lacking for language impairments.
Approach: They propose to use a French dataset that contains 230 hours of speech resources to create a rule-based pictogram grammar with a restricted vocabulary and a discussion of strategic decisions.
Outcome: The proposed model is validated through multiple post-editing phases by expert annotators and is freely available under a non-commercial licence.
UFSAC: Unification of Sense Annotated Corpora and Tools (L18-1)

Copied to clipboard

Challenge: a dozen sense annotated English corpora are used in Word Sense Disambiguation (WSD) a new format of corpus is proposed that can be used for training or testing a disambiguation system .
Approach: They propose a format of corpus that can be used for training or testing a disambiguation system . they provide the source code and a complete Java API for manipulating corpora in this format .
Outcome: The proposed format of corpus can be used for training or testing a disambiguation system . the source code and a complete Java API are provided for building the corpus .
Automatic Speech Recognition and Query By Example for Creole Languages Documentation (2022.findings-acl)

Copied to clipboard

Challenge: CREAM project aims to provide linguists with new methods for language documentation based on automatic speech recognition and keyword-spotting.
Approach: They propose to use one hour of annotated data to design an automatic speech recognition system for two Creole languages.
Outcome: The proposed model is based on an hour of annotated data and is usable by linguists.
Jargon: A Suite of Language Models and Evaluation Tasks for French Specialized Domains (2024.lrec-main)

Copied to clipboard

Challenge: Pretrained language models are the de facto backbone of most state-of-the-art NLP systems.
Approach: They propose a family of domain-specific pretrained PLMs for French focusing on three important domains: transcribed speech, medicine, and law.
Outcome: The proposed models perform better on transcribed speech, medicine, and law domains than state-of-the-art models on a diverse set of tasks and datasets.
What Has LeBenchmark Learnt about French Syntax? (2024.lrec-main)

Copied to clipboard

Challenge: Pretrained acoustic models are increasingly used for downstream speech tasks such as automatic speech recognition, speech translation, spoken language understanding or speech parsing.
Approach: They propose to probing a pretrained acoustic model for French for syntactic information using the Orféo treebank.
Outcome: The proposed model is trained on 7k hours of spoken French and obtained reasonable results on tasks that require higher level linguistic knowledge.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations