Papers by Sophie Rosset
mALBERT: Is a Compact Multilingual BERT Model Still Worth It? (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing studies on the ethical and ecological impact of pre-trained language models raise questions about the temporal, financial, and environmental aspects of such models. |
| Approach: | They propose to focus on smaller models, such as compact models like ALBERT, which are more ecologically virtuous than these PLMs. |
| Outcome: | The proposed model is compared with classical multilingual models and is ethically virtuous. |
Where are we in Named Entity Recognition from Speech? (2020.lrec-1)
Copied to clipboard
| Challenge: | Named entity recognition is usually made through a pipeline process that consists of processing audio and applying a NER to the audio outputs. |
| Approach: | They propose an original 3-pass approach and explore the capability of an E2E system to do structured NER. |
| Outcome: | The proposed system performs better than the current pipeline approach. |
Small Language Models Are Good Too: An Empirical Study of Zero-Shot Classification (2024.lrec-main)
Copied to clipboard
| Challenge: | Using small language models, we challenge the dominance of large models in text classification by prompting. |
| Approach: | They compare the performance of small and large language models in a zero-shot context using different architectures and scoring functions. |
| Outcome: | The proposed model outperforms large models in a zero-shot context. |
Neural Networks approaches focused on French Spoken Language Understanding: application to the MEDIA Evaluation Task (2020.coling-main)
Copied to clipboard
| Challenge: | Recent studies have focused on English language and tasks, but few have explored the complexity of a SLU task. |
| Approach: | They propose to explore Neural Networks approaches for a French Spoken Language Understanding task. |
| Outcome: | The proposed approach outperforms classical Neural Network Architectures and achieves state-of-the-art results. |
New Semantic Task for the French Spoken Language Understanding MEDIA Benchmark (2024.lrec-main)
Copied to clipboard
| Challenge: | Intent classification and slot-filling tasks are essential tasks of Spoken Language Understanding (SLU). |
| Approach: | They propose to use a MEDIA SLU dataset to train a multilingual model to achieve both tasks jointly. |
| Outcome: | The proposed model can be trained on multiple datasets including the MEDIA dataset and extends to more tasks and use cases. |
Corpora with Part-of-Speech Annotations for Three Regional Languages of France: Alsatian, Occitan and Picard (L18-1)
Copied to clipboard
Delphine Bernhard, Anne-Laure Ligozat, Fanny Martin, Myriam Bras, Pierre Magistry, Marianne Vergez-Couret, Lucie Steiblé, Pascale Erhart, Nabil Hathout, Dominique Huck, Christophe Rey, Philippe Reynés, Sophie Rosset, Jean Sibille, Thomas Lavergne
| Challenge: | RESTAURE project aims to develop resources and tools for three regional languages of France: Alsatian, Occitan and Picard. |
| Approach: | They describe the creation of corpora with part-of-speech annotations for Alsatian, Occitan and Picard. |
| Outcome: | The authors describe the creation of annotated corpora for Alsatian, Occitan and Picard . the project is part of the RESTAURE project, which aims to develop resources and tools for these under-resourced French regional languages. |