A Semi-Automatic Approach to Create Large Gender- and Age-Balanced Speaker Corpora: Usefulness of Speaker Diarization & Identification. (2022.lrec-1)
Copied to clipboard
Rémi Uro, David Doukhan, Albert Rilliard, Laetitia Larcher, Anissa-Claire Adgharouamane, Marie Tahon, Antoine Laurent
| Challenge: | Existing methods for creating diachronic corpus of voices are based on speaker characteristics and require human intervention. |
| Approach: | They propose to use a semi-automatic pipeline to create a diachronic corpus of voices balanced for speaker’s age, gender and recording period, according to 32 categories. |
| Outcome: | The proposed method cut down on manual annotations by ten and provides high quality speech for most of the selected excerpts. |
Similar Papers
ALLIES: A Speech Corpus for Segmentation, Speaker Diarization, Speech Recognition and Speaker Change Detection (2024.lrec-main)
Copied to clipboard
| Challenge: | a meta corpus of audio files is used to gather, annotate and transcribe speech . a large number of speech databases are needed to perform multi-speaker tasks such as speaker diarization and speaker change detection. |
| Approach: | They propose to use human feedback to homogenize and correct speaker labels among the audio files by integrating human feedback within a speaker verification system. |
| Outcome: | The proposed protocol evaluates speech segmentation, speaker diarization, speech transcription and speaker change detection using human feedback. |
Using Automatic Speech Recognition in Spoken Corpus Curation (2020.lrec-1)
Copied to clipboard
| Challenge: | Automatic Speech Recognition (ASR) is a new way to make audio-visual data accessible. |
| Approach: | They propose to use automatic speech recognition (ASR) to make audio-visual data accessible by systematic queries. |
| Outcome: | The proposed system has higher recognition scores for the north of Germany vs. lower scores for south of the country. |
AnnoTheia: A Semi-Automatic Annotation Toolkit for Audio-Visual Speech Technologies (2024.lrec-main)
Copied to clipboard
| Challenge: | a small fraction of the languages currently covered by speech technologies are mainly spoken in English. |
| Approach: | They present an annotation toolkit that detects when a person speaks on the scene and the corresponding transcription. |
| Outcome: | The proposed toolkit can speed up the annotation process by up to four times . it can be used in Spanish, and is available on github. |
Using a Knowledge Base to Automatically Annotate Speech Corpora and to Identify Sociolinguistic Variation (2022.lrec-1)
Copied to clipboard
| Challenge: | Speech characteristics vary from speaker to speaker due to many factors, including communication context, provenance, age, and social background. |
| Approach: | They propose a method that uses a knowledge base to provide speaker-specific information. |
| Outcome: | The proposed method can be used to enrich existing corpora with speaker-specific information and to correlate with diastratic features. |
SynPaFlex-Corpus: An Expressive French Audiobooks Corpus dedicated to expressive speech synthesis. (L18-1)
Copied to clipboard
| Challenge: | a French audiobooks corpus contains 87 hours of good audio quality speech . audiobooks provide mono-genre and multi-speaker speech whereas audiobooks usually provide a few hours of mono- and multispeakers . |
| Approach: | They present an expressive French audiobooks corpus containing eighty seven hours of speech . the corpus is annotated automatically and provides information as phone labels, phone boundaries, syllables, words or morpho-syntactic tagging. |
| Outcome: | The proposed corpus contains 87 hours of speech recorded by a single speaker . the data will allow developing models to better control expressiveness in speech synthesis . |
Overlaps and Gender Analysis in the Context of Broadcast Media (2022.lrec-1)
Copied to clipboard
| Challenge: | Using gender and overlap annotations, we characterise interactions between speakers according to their gender and role in broadcast media. |
| Approach: | They propose to characterise interactions between speakers according to their gender and role in broadcast media by using a small dataset of 93 recordings from LCP French channel. |
| Outcome: | The proposed method could improve the efficiency of qualitative studies conducted in human sciences. |
Computer-assisted Speaker Diarization: How to Evaluate Human Corrections (L18-1)
Copied to clipboard
| Challenge: | a framework to evaluate the human corrections of a speaker diarization is presented for the French National Audiovisual Institute (INA) the speaker diaarization task is a necessary pre-processing step for speaker identification and speech transcription. |
| Approach: | They propose a framework to evaluate the human corrections of a speaker diarization . they propose four elementary actions to correct the diarized speaker and an automaton to simulate the correction sequence. |
| Outcome: | The proposed framework copes with the needs of the French National Audiovisual Institute (INA) due to the increasing number of documents and the limited number of annotators, many documents remain undocumented or only partly documented. |
Artie Bias Corpus: An Open Dataset for Detecting Demographic Bias in Speech Applications (2020.lrec-1)
Copied to clipboard
| Challenge: | A speech technology exhibits demographic bias when performance is worse for one demographic group relative to another. |
| Approach: | They create an English dataset of expert-validated audio, transcript> pairs with demographic tags for age, gender, accent and open software which may be used to detect demographic bias in Automatic Speech Recognition systems. |
| Outcome: | The Artie Bias Corpus is a curated subset of the Mozilla Common Voice corpus, which is released under a Creative Commons CC0 license . |
GIL-GALaD: Gender Inclusive Language - German Auto-Assembled Large Database (2024.lrec-main)
Copied to clipboard
| Challenge: | grammatically gendered languages such as German pose unique challenges in generating gender-inclusive language for corrective model training or fine-tuning. |
| Approach: | a corpus of German gender-inclusive language is assembled to help improve model training . grammatically gendered languages such as german pose unique challenges . authors describe most common strategies for gender- inclusive language in german . |
| Outcome: | a corpus of German gender-inclusive language is assembled and will be included in the release. |
Under the Morphosyntactic Lens: A Multifaceted Evaluation of Gender Bias in Speech Translation (2022.acl-long)
Copied to clipboard
| Challenge: | grammatical gender languages are characterized by morphosyntactic chains of gender agreement marked on a variety of lexical items and parts-of-speech (POS). |
| Approach: | They propose to enrich the natural, gender-sensitive MuST-SHE corpus with two new linguistic annotation layers to explore gender bias. |
| Outcome: | The proposed models shed light on gender bias and its detection at several levels of granularity. |