Papers by Salima Mdhaffar
A Multimodal Educational Corpus of Oral Courses: Annotation, Analysis and Case Study (2020.lrec-1)
Copied to clipboard
Salima Mdhaffar, Yannick Estève, Antoine Laurent, Nicolas Hernandez, Richard Dufour, Delphine Charlet, Geraldine Damnati, Solen Quiniou, Nathalie Camelin
| Challenge: | a corpus of spontaneous speech is being developed for educational use . the dataset will be freely available to the research community . |
| Approach: | They propose to use a French speech educational corpus to explore synchronous speech transcription and application in teaching situations. |
| Outcome: | The proposed corpus includes 10 hours of lectures, manually transcribed and segmented . the dataset will be freely available to the research community . |
The Spoken Language Understanding MEDIA Benchmark Dataset in the Era of Deep Learning: data updates, training and evaluation tools (2022.lrec-1)
Copied to clipboard
Gaëlle Laperrière, Valentin Pelloin, Antoine Caubrière, Salima Mdhaffar, Nathalie Camelin, Sahar Ghannay, Bassam Jabaian, Yannick Estève
| Challenge: | a growing number of studies address the spoken language understanding domain through a simple task like speech intent detection. |
| Approach: | They focus on the french MEDIA SLU dataset, which is distributed since 2005 . they propose a recipe for its use, including data preparation, training and evaluation scripts . |
| Outcome: | The MEDIA SLU dataset is used as a benchmark dataset for a large number of research projects. |
Impact Analysis of the Use of Speech and Language Models Pretrained by Self-Supersivion for Spoken Language Understanding (2022.lrec-1)
Copied to clipboard
Salima Mdhaffar, Valentin Pelloin, Antoine Caubrière, Gaëlle Laperriere, Sahar Ghannay, Bassam Jabaian, Nathalie Camelin, Yannick Estève
| Challenge: | Pretrained models have been introduced for both acoustic and language modeling. |
| Approach: | They present an error analysis of pretrained models using a french MEDIA benchmark dataset. |
| Outcome: | The proposed models have been able to improve on the french MEDIA benchmark dataset, which is one of the most challenging among all benchmarks accessible to the entire research community. |
TARIC-SLU: A Tunisian Benchmark Dataset for Spoken Language Understanding (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing SLU resources are limited in high-resource languages such as English, Mandarin and French. |
| Approach: | They propose to use a Tunisian dialect dataset to build a semantic model of the system that is continuously annotated with dialogue acts and slots. |
| Outcome: | The proposed dataset is based on train-based and ASR-based models of train-driven conversations in Tunisian dialect. |
Sonos Voice Control Bias Assessment Dataset: A Methodology for Demographic Bias Assessment in Voice Assistants (2024.lrec-main)
Copied to clipboard
Chloe Sekkat, Fanny Leroy, Salima Mdhaffar, Blake Perry Smith, Yannick Estève, Joseph Dureau, Alice Coucke
| Challenge: | Recent studies show voice assistants do not perform equally well for everyone . however, research on demographic robustness of speech technologies is still scarce . |
| Approach: | They propose a statistical method to detect demographic bias using a large dataset with controlled demographic tags. |
| Outcome: | The proposed method shows statistically significant differences in performance across age, dialectal region and ethnicity. |