Papers by Nathalie Camelin
A Multimodal Educational Corpus of Oral Courses: Annotation, Analysis and Case Study (2020.lrec-1)
Copied to clipboard
Salima Mdhaffar, Yannick Estève, Antoine Laurent, Nicolas Hernandez, Richard Dufour, Delphine Charlet, Geraldine Damnati, Solen Quiniou, Nathalie Camelin
| Challenge: | a corpus of spontaneous speech is being developed for educational use . the dataset will be freely available to the research community . |
| Approach: | They propose to use a French speech educational corpus to explore synchronous speech transcription and application in teaching situations. |
| Outcome: | The proposed corpus includes 10 hours of lectures, manually transcribed and segmented . the dataset will be freely available to the research community . |
The Spoken Language Understanding MEDIA Benchmark Dataset in the Era of Deep Learning: data updates, training and evaluation tools (2022.lrec-1)
Copied to clipboard
Gaëlle Laperrière, Valentin Pelloin, Antoine Caubrière, Salima Mdhaffar, Nathalie Camelin, Sahar Ghannay, Bassam Jabaian, Yannick Estève
| Challenge: | a growing number of studies address the spoken language understanding domain through a simple task like speech intent detection. |
| Approach: | They focus on the french MEDIA SLU dataset, which is distributed since 2005 . they propose a recipe for its use, including data preparation, training and evaluation scripts . |
| Outcome: | The MEDIA SLU dataset is used as a benchmark dataset for a large number of research projects. |
Impact Analysis of the Use of Speech and Language Models Pretrained by Self-Supersivion for Spoken Language Understanding (2022.lrec-1)
Copied to clipboard
Salima Mdhaffar, Valentin Pelloin, Antoine Caubrière, Gaëlle Laperriere, Sahar Ghannay, Bassam Jabaian, Nathalie Camelin, Yannick Estève
| Challenge: | Pretrained models have been introduced for both acoustic and language modeling. |
| Approach: | They present an error analysis of pretrained models using a french MEDIA benchmark dataset. |
| Outcome: | The proposed models have been able to improve on the french MEDIA benchmark dataset, which is one of the most challenging among all benchmarks accessible to the entire research community. |
Simulating ASR errors for training SLU systems (L18-1)
Copied to clipboard
| Challenge: | Existing methods to simulate automatic speech recognition errors from manual transcriptions are not available during training of the SLU model. |
| Approach: | They propose to use acoustic and linguistic word embeddings to define a similarity measure between words to predict ASR confusions. |
| Outcome: | The proposed method significantly improves the performance of spoken language understanding systems. |
Toward Qualitative Evaluation of Embeddings for Arabic Sentiment Analysis (2020.lrec-1)
Copied to clipboard
| Challenge: | Existing studies on Arabic sentiment analysis (SA) tasks focus on word embeddings to capture semantic and syntactic similarities, but Arabic language is characterized by its agglutination and morphological richness contributing to great sparsity. |
| Approach: | They propose several protocols to evaluate specific embeddings for Arabic sentiment analysis task. |
| Outcome: | The proposed embeddings are based on words and lemmas in Arabic sentiment analysis (SA) task. |
FrNewsLink : a corpus linking TV Broadcast News Segments and Press Articles (L18-1)
Copied to clipboard
Nathalie Camelin, Géraldine Damnati, Abdessalam Bouchekif, Anais Landeau, Delphine Charlet, Yannick Estève
| Challenge: | a corpus of TV Broadcast News resources is proposed to address several applicative tasks. |
| Approach: | They propose to use a corpus to address several applicative tasks that are made public . they propose to gather TVBN shows and press articles and use them to study semantic similarity . |
| Outcome: | The proposed corpus is based on 112 TVBN shows and press articles . it allows to study semantic similarity and multimedia News linking . |
Are Embedding Spaces Interpretable? Results of an Intrusion Detection Evaluation on a Large French Corpus (2022.lrec-1)
Copied to clipboard
| Challenge: | Word embedding methods use word co-occurrences to encode, syntactic and semantic information to describe vocabulary in a low-dimensional space. |
| Approach: | They evaluate word embedding interpretability using two methods . they use a word-in-space vector encoder and graph-based method SPINE . |
| Outcome: | The proposed methods show that they can be interpretable on a large French corpus. |