Papers by Sophie Rosset

6 papers
mALBERT: Is a Compact Multilingual BERT Model Still Worth It? (2024.lrec-main)

Copied to clipboard

Challenge: Existing studies on the ethical and ecological impact of pre-trained language models raise questions about the temporal, financial, and environmental aspects of such models.
Approach: They propose to focus on smaller models, such as compact models like ALBERT, which are more ecologically virtuous than these PLMs.
Outcome: The proposed model is compared with classical multilingual models and is ethically virtuous.
Where are we in Named Entity Recognition from Speech? (2020.lrec-1)

Copied to clipboard

Challenge: Named entity recognition is usually made through a pipeline process that consists of processing audio and applying a NER to the audio outputs.
Approach: They propose an original 3-pass approach and explore the capability of an E2E system to do structured NER.
Outcome: The proposed system performs better than the current pipeline approach.
Small Language Models Are Good Too: An Empirical Study of Zero-Shot Classification (2024.lrec-main)

Copied to clipboard

Challenge: Using small language models, we challenge the dominance of large models in text classification by prompting.
Approach: They compare the performance of small and large language models in a zero-shot context using different architectures and scoring functions.
Outcome: The proposed model outperforms large models in a zero-shot context.
Neural Networks approaches focused on French Spoken Language Understanding: application to the MEDIA Evaluation Task (2020.coling-main)

Copied to clipboard

Challenge: Recent studies have focused on English language and tasks, but few have explored the complexity of a SLU task.
Approach: They propose to explore Neural Networks approaches for a French Spoken Language Understanding task.
Outcome: The proposed approach outperforms classical Neural Network Architectures and achieves state-of-the-art results.
New Semantic Task for the French Spoken Language Understanding MEDIA Benchmark (2024.lrec-main)

Copied to clipboard

Challenge: Intent classification and slot-filling tasks are essential tasks of Spoken Language Understanding (SLU).
Approach: They propose to use a MEDIA SLU dataset to train a multilingual model to achieve both tasks jointly.
Outcome: The proposed model can be trained on multiple datasets including the MEDIA dataset and extends to more tasks and use cases.
Corpora with Part-of-Speech Annotations for Three Regional Languages of France: Alsatian, Occitan and Picard (L18-1)

Copied to clipboard

Challenge: RESTAURE project aims to develop resources and tools for three regional languages of France: Alsatian, Occitan and Picard.
Approach: They describe the creation of corpora with part-of-speech annotations for Alsatian, Occitan and Picard.
Outcome: The authors describe the creation of annotated corpora for Alsatian, Occitan and Picard . the project is part of the RESTAURE project, which aims to develop resources and tools for these under-resourced French regional languages.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations