Papers by Solange Rossato
Unraveling Spontaneous Speech Dimensions for Cross-Corpus ASR System Evaluation for French (2024.lrec-main)
Copied to clipboard
| Challenge: | 'spontaneous speech' is a catch-all term used for situations like speaking with a friend, being interviewed on radio/TV or giving a lecture. |
| Approach: | They propose to use four dimensions to describe spontaneous speech variation in automatic speech recognition systems. |
| Outcome: | The proposed system can be used to predict the WER of speech recognition systems on face-to-face interactions. |
Identifying and typifying demographic unfairness in phoneme-level embeddings of self-supervised speech recognition models (2026.findings-acl)
Copied to clipboard
| Challenge: | a number of studies have been done to improve ASR for speaker groups, but there is still room for improvement . authors propose a framework typifying two types of error in phoneme embeddings . |
| Approach: | They propose a framework typifying two types of error that can occur in phoneme modeling . they propose random error/high variance in phonemes embedding vs systematic error/embedding bias . |
| Outcome: | The proposed framework typifies errors in phoneme modeling in ASR systems . it shows that training only on a single, typically disadvantaged SG improves performance . |
Gender Representation in Open Source Speech Resources (2020.lrec-1)
Copied to clipboard
| Challenge: | Using open source corpora, we find that gender balance depends on other corpus characteristics such as elicited/non ellicite vs. non-eliciting speech, low/high resource language, speech task targeted. |
| Approach: | They propose to use open source corpora to find gender information in spoken language systems . they propose metadata and recommendations for researchers to assure better transparency . |
| Outcome: | The proposed method improves the quality and transparency of open source speech resources. |
A Corpus of Spontaneous L2 English Speech for Real-situation Speaking Assessment (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing automated scoring systems rely on highly controlled elicitation protocols, such as reading aloud isolated words or short sentences, limiting their ability to evaluate spontaneous speech. |
| Approach: | They propose to collect a corpus of spontaneous L2 English speech from university students as part of a French national certificate in English. |
| Outcome: | The results show that only 35.4% of the 6,350 targeted words had stress detected on the expected syllable, revealing a common stress shift to the final s. |
Audiocite.net : A Large Spoken Read Dataset in French (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing self-supervised learning methods for speech processing have proved difficult to apply to French due to the scarcity of large speech datasets. |
| Approach: | They present a corpus of 6,682 hours of audiobooks from 130 readers . they describe the creation process and final statistics of the corpus . |
| Outcome: | The proposed model based on the audiocite.net corpus, which contains 6,682 hours of audiobooks, was able to perform in 14k version. |
Rhythmic Proximity Between Natives And Learners Of French - Evaluation of a metric based on the CEFC corpus (2020.lrec-1)
Copied to clipboard
| Challenge: | Among prosodic parameters, rhythm is one that varies noticeably from one language to another. |
| Approach: | They propose to model rhythm in French using the corpus for l’Étude du Français Contemporain (CEFC) . they tested 146 native speakers, 37 non-native speakers and 29 non-Native Japanese learners of French . |
| Outcome: | The proposed model is based on the corpus pour l’Étude du Français Contemporain (CEFC) which contains up to 300 hours of speech of a wide variety of speaker profiles and situations. |