Papers by Mahault Garnerin
Gender Representation in Open Source Speech Resources (2020.lrec-1)
Copied to clipboard
| Challenge: | Using open source corpora, we find that gender balance depends on other corpus characteristics such as elicited/non ellicite vs. non-eliciting speech, low/high resource language, speech task targeted. |
| Approach: | They propose to use open source corpora to find gender information in spoken language systems . they propose metadata and recommendations for researchers to assure better transparency . |
| Outcome: | The proposed method improves the quality and transparency of open source speech resources. |
MaSS: A Large and Clean Multilingual Corpus of Sentence-aligned Spoken Utterances Extracted from the Bible (2020.lrec-1)
Copied to clipboard
| Challenge: | The Bible is the same for all the languages, thus constituting a multilingual and comparable 2 spoken corpus, is not exploited to date. |
| Approach: | They propose to add multilingual links between small speech segments in different languages . they use a large dataset of 8,130 parallel spoken utterances across 8 languages - maSS . |
| Outcome: | The proposed model can build automatic speech recognition models for 700 languages. |