Papers by Solange Rossato

6 papers
Unraveling Spontaneous Speech Dimensions for Cross-Corpus ASR System Evaluation for French (2024.lrec-main)

Copied to clipboard

Challenge: 'spontaneous speech' is a catch-all term used for situations like speaking with a friend, being interviewed on radio/TV or giving a lecture.
Approach: They propose to use four dimensions to describe spontaneous speech variation in automatic speech recognition systems.
Outcome: The proposed system can be used to predict the WER of speech recognition systems on face-to-face interactions.
Identifying and typifying demographic unfairness in phoneme-level embeddings of self-supervised speech recognition models (2026.findings-acl)

Copied to clipboard

Challenge: a number of studies have been done to improve ASR for speaker groups, but there is still room for improvement . authors propose a framework typifying two types of error in phoneme embeddings .
Approach: They propose a framework typifying two types of error that can occur in phoneme modeling . they propose random error/high variance in phonemes embedding vs systematic error/embedding bias .
Outcome: The proposed framework typifies errors in phoneme modeling in ASR systems . it shows that training only on a single, typically disadvantaged SG improves performance .
Gender Representation in Open Source Speech Resources (2020.lrec-1)

Copied to clipboard

Challenge: Using open source corpora, we find that gender balance depends on other corpus characteristics such as elicited/non ellicite vs. non-eliciting speech, low/high resource language, speech task targeted.
Approach: They propose to use open source corpora to find gender information in spoken language systems . they propose metadata and recommendations for researchers to assure better transparency .
Outcome: The proposed method improves the quality and transparency of open source speech resources.
A Corpus of Spontaneous L2 English Speech for Real-situation Speaking Assessment (2024.lrec-main)

Copied to clipboard

Challenge: Existing automated scoring systems rely on highly controlled elicitation protocols, such as reading aloud isolated words or short sentences, limiting their ability to evaluate spontaneous speech.
Approach: They propose to collect a corpus of spontaneous L2 English speech from university students as part of a French national certificate in English.
Outcome: The results show that only 35.4% of the 6,350 targeted words had stress detected on the expected syllable, revealing a common stress shift to the final s.
Audiocite.net : A Large Spoken Read Dataset in French (2024.lrec-main)

Copied to clipboard

Challenge: Existing self-supervised learning methods for speech processing have proved difficult to apply to French due to the scarcity of large speech datasets.
Approach: They present a corpus of 6,682 hours of audiobooks from 130 readers . they describe the creation process and final statistics of the corpus .
Outcome: The proposed model based on the audiocite.net corpus, which contains 6,682 hours of audiobooks, was able to perform in 14k version.
Rhythmic Proximity Between Natives And Learners Of French - Evaluation of a metric based on the CEFC corpus (2020.lrec-1)

Copied to clipboard

Challenge: Among prosodic parameters, rhythm is one that varies noticeably from one language to another.
Approach: They propose to model rhythm in French using the corpus for l’Étude du Français Contemporain (CEFC) . they tested 146 native speakers, 37 non-native speakers and 29 non-Native Japanese learners of French .
Outcome: The proposed model is based on the corpus pour l’Étude du Français Contemporain (CEFC) which contains up to 300 hours of speech of a wide variety of speaker profiles and situations.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations