Papers by Didier Schwab

13 papers
Providing Semantic Knowledge to a Set of Pictograms for People with Disabilities: a Set of Links between WordNet and Arasaac: Arasaac-WN (2020.lrec-1)

Copied to clipboard

Challenge: Pictograms are a tool that is increasingly used by people with cognitive or communication disabilities.
Approach: They propose a database that links WordNet and Arasaac to link pictograms to semantic knowledge.
Outcome: The proposed database links pictograms with WordNet and Arasaac to create language-independent prototypes.
FlauBERT: Unsupervised Language Model Pre-training for French (2020.lrec-1)

Copied to clipboard

Challenge: Language models are a key step to achieve state-of-the-art results in many different Natural Language Processing (NLP) tasks.
Approach: They propose to use a language model that is pre-trained on a large and heterogeneous French corpus to train continuous word representations.
Outcome: The proposed model outperforms existing models on a large and heterogeneous French corpus.
Do Multilingual Neural Machine Translation Models Contain Language Pair Specific Attention Heads? (2021.findings-acl)

Copied to clipboard

Challenge: Recent studies on multilingual representations focus on whether there is an emergence of language-independent representations or whether multilingual models partition their weights among different languages.
Approach: They analyze encoder self-attention and encoder-decoder attention heads in a multilingual neural translation model.
Outcome: The proposed model is based on a multilingual neural translation model with a language-independent representation.
Should Cross-Lingual AMR Parsing go Meta? An Empirical Assessment of Meta-Learning and Joint Learning AMR Parsing (2024.findings-emnlp)

Copied to clipboard

Challenge: Cross-lingual AMR parsing is a task of predicting AMR graphs in a target language when training data is available only in . et al. (2018) evaluated meta-learning for cross-lingual parse in Croatian, Farsi, Korean, Chinese, and French.
Approach: They propose to use meta-learning to tackle cross-lingual AMR parsing in a target language . they evaluate their models in k-shot scenarios and compare them to classical joint learning .
Outcome: The proposed model performs better in 0-shot evaluation for Croatian, Farsi, Korean, Chinese, and French.
A Multimodal French Corpus of Aligned Speech, Text, and Pictogram Sequences for Speech-to-Pictogram Machine Translation (2024.lrec-main)

Copied to clipboard

Challenge: Existing algorithms for the automatic translation of spoken language into pictogram units are lacking for language impairments.
Approach: They propose to use a French dataset that contains 230 hours of speech resources to create a rule-based pictogram grammar with a restricted vocabulary and a discussion of strategic decisions.
Outcome: The proposed model is validated through multiple post-editing phases by expert annotators and is freely available under a non-commercial licence.
WIKIR: A Python Toolkit for Building a Large-scale Wikipedia-based English Information Retrieval Dataset (2020.lrec-1)

Copied to clipboard

Challenge: ad-hoc information retrieval methods usually require large amounts of annotated data to be effective.
Approach: They propose an open-source toolkit to automatically build large-scale English information retrieval datasets based on Wikipedia.
Outcome: The proposed toolkit builds large-scale English information retrieval datasets based on Wikipedia with 59,252 queries and 2,617,003 pairs.
Limitations of Human Identification of Automatically Generated Text (2024.lrec-main)

Copied to clipboard

Challenge: Neural text generation tools such as ChatGPT are gaining popularity . human annotations are considered gold standard labels for multiple tasks .
Approach: They propose a new corpus in French and English for recognising automatically generated texts . they propose 'incontext' setup which makes explicit the interaction between two parties .
Outcome: The proposed model generates fluent text, which requires much closer reading than the current model.
Lightweight Adapter Tuning for Multilingual Speech Translation (2021.acl-short)

Copied to clipboard

Challenge: Adapter tuning is an efficient alternative to fine-tuning in NLP . a multilingual model could be outperformed by its bilingual counterparts .
Approach: They propose to use adapter tuning to optimize for multilingual speech translation . they use pre-trained models to freeze pre-train parameters and inject lightweight modules .
Outcome: The proposed adapters can specialize to specific language pairs with low extra cost . the proposed models outperform bilingual models on high-resource language pairs .
UFSAC: Unification of Sense Annotated Corpora and Tools (L18-1)

Copied to clipboard

Challenge: a dozen sense annotated English corpora are used in Word Sense Disambiguation (WSD) a new format of corpus is proposed that can be used for training or testing a disambiguation system .
Approach: They propose a format of corpus that can be used for training or testing a disambiguation system . they provide the source code and a complete Java API for manipulating corpora in this format .
Outcome: The proposed format of corpus can be used for training or testing a disambiguation system . the source code and a complete Java API are provided for building the corpus .
Automatic Speech Recognition and Query By Example for Creole Languages Documentation (2022.findings-acl)

Copied to clipboard

Challenge: CREAM project aims to provide linguists with new methods for language documentation based on automatic speech recognition and keyword-spotting.
Approach: They propose to use one hour of annotated data to design an automatic speech recognition system for two Creole languages.
Outcome: The proposed model is based on an hour of annotated data and is usable by linguists.
What Matters to an LLM? Behavioral and Computational Evidences from Summarization (2026.findings-eacl)

Copied to clipboard

Challenge: Large Language Models (LLMs) are increasingly entrusted with the management of information.
Approach: They combine behavioral and computational analyses to find out what LLMs prioritize . they generate length-controlled summaries and derive empirical importance distributions .
Outcome: The proposed model converges on consistent importance patterns and clusters more by family than by size.
Jargon: A Suite of Language Models and Evaluation Tasks for French Specialized Domains (2024.lrec-main)

Copied to clipboard

Challenge: Pretrained language models are the de facto backbone of most state-of-the-art NLP systems.
Approach: They propose a family of domain-specific pretrained PLMs for French focusing on three important domains: transcribed speech, medicine, and law.
Outcome: The proposed models perform better on transcribed speech, medicine, and law domains than state-of-the-art models on a diverse set of tasks and datasets.
Dual-decoder Transformer for Joint Automatic Speech Recognition and Multilingual Speech Translation (2020.coling-main)

Copied to clipboard

Challenge: Existing models for automatic speech recognition and multilingual speech translation are on par with cascade counterparts.
Approach: They propose a dual-decoder Transformer architecture that performs automatic speech recognition and multilingual speech translation.
Outcome: The proposed models outperform the previously-reported highest translation performance in multilingual settings and bilingual one-to-one results.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations