| Challenge: | The corpus consists of 20 hours of video recordings of spontaneous talk-in-interaction typically involving 2-4 speakers. |
| Approach: | the corpus consists of 20 hours of video recordings of spontaneous talk-in-interaction typically involving 2-4 speakers. |
| Outcome: | the corpus consists of 20 hours of video recordings of spontaneous talk-in-interaction typically involving 2-4 speakers. |
Similar Papers
WeCanTalk: A New Multi-language, Multi-modal Resource for Speaker Recognition (2022.lrec-1)
Copied to clipboard
| Challenge: | The WeCanTalk corpus is a multi-modal, multi-language resource for speaker recognition. |
| Approach: | The WeCanTalk corpus is a multi-modal resource for speaker recognition. |
| Outcome: | The corpus contains data from 202 native speakers in Hong Kong who were fluent in at least one other language. |
CantoMap: a Hong Kong Cantonese MapTask Corpus (2020.lrec-1)
Copied to clipboard
| Challenge: | a corpus of connected spoken Hong Kong Cantonese is constructed to study the phonology and semantics of the language. |
| Approach: | They propose to build a corpus of connected spoken Hong Kong Cantonese with phonemic transcription and controlled elicitation tasks. |
| Outcome: | The proposed corpus contains 768 minutes of recordings and transcripts of forty speakers. |
PyCantonese: Cantonese Linguistics and NLP in Python (2022.lrec-1)
Copied to clipboard
| Challenge: | a limited number of Cantonese-specific datasets are available for PyCantones. |
| Approach: | They introduce PyCantonese, an open-source Python library for Cantonesi linguistics and natural language processing. |
| Outcome: | The proposed library is open-source and available for free for all purposes, including commercial ones. |
SpiCE: A New Open-Access Corpus of Conversational Bilingual Speech in Cantonese and English (2020.lrec-1)
Copied to clipboard
| Challenge: | SpiCE is a corpus of conversational Cantonese-English bilingual speech recorded in Vancouver, Canada . the corpus includes high-quality recordings of 34 early bilinguals in both English and Cantoneses . |
| Approach: | They describe the design, collection, orthographic transcription, and phonetic annotation of SpiCE . the corpus includes high-quality recordings of 34 early bilinguals in both English and Cantonese . |
| Outcome: | The SpiCE corpus includes high-quality recordings of 34 early bilinguals in both English and Cantonese . the corpus will promote bilingualism research for a typologically distinct pair of languages . |
A Crowdsourced Open-Source Kazakh Speech Corpus and Initial Speech Recognition Baseline (2021.eacl-main)
Copied to clipboard
Yerbolat Khassanov, Saida Mussakhojayeva, Almas Mirzakhmetov, Alen Adiyev, Mukhamet Nurpeiissov, Huseyin Atakan Varol
| Challenge: | The Kazakh speech corpus contains over 153,000 utterances spoken by participants from different regions and age groups, as well as both genders. |
| Approach: | They propose to build an open-source Kazakh speech corpus for the Kazakh language that contains over 153,000 transcribed audio . they describe the data collection and preprocessing procedures followed by a description of the database specifications. |
| Outcome: | The Kazakh speech corpus contains over 153,000 utterances spoken by participants from different regions and age groups, as well as both genders. |
CanVEC - the Canberra Vietnamese-English Code-switching Natural Speech Corpus (2020.lrec-1)
Copied to clipboard
| Challenge: | Using monolingual tools, code-switching is a problem in the natural language processing community. |
| Approach: | They introduce the Canberra Vietnamese-English Code-switching corpus (CanVEC) which is an original corpus of mixed speech annotated with language information, part of speech tags and Vietnamese translations. |
| Outcome: | The proposed corpus was annotated with language information, part of speech tags and Vietnamese translations using pipelining and monolingual toolkits. |
Automatic Speech Recognition Datasets in Cantonese: A Survey and New Dataset (2022.lrec-1)
Copied to clipboard
Tiezheng Yu, Rita Frieske, Peng Xu, Samuel Cahyawijaya, Cheuk Tung Yiu, Holy Lovenia, Wenliang Dai, Elham J. Barezi, Qifeng Chen, Xiaojuan Ma, Bertram Shi, Pascale Fung
| Challenge: | In this paper, we address the problem of data scarcity for the Hong Kong Cantonese language . due to the popularization of deep learning, ASR technology has led to a significant improvement in recognizing many languages. |
| Approach: | They propose to use a dataset to analyze the data available for the Hong Kong Cantonese language . they use zh-HK as a source and a state-of-the-art ASR model to build a powerful model . |
| Outcome: | The proposed model improves on the biggest existing dataset, Common Voice zh-HK. |
My Science Tutor (MyST)–a Large Corpus of Children’s Conversational Speech (2024.lrec-main)
Copied to clipboard
| Challenge: | a 13-year project was conducted between 2007 and 2019 to improve students' learning proficiency in elementary school science using conversational multimedia virtual tutor, Marni. |
| Approach: | They propose to use the corpus-name corpus to improve automatic speech recognition models and algorithms by training and developing a model on the training and development portion of the corpuse. |
| Outcome: | The corpus comprises 400 hours of speech, spanning some 230K utterances spread across about 10,500 virtual tutor sessions. |
Cifu: a Frequency Lexicon of Hong Kong Cantonese (2020.lrec-1)
Copied to clipboard
| Challenge: | lexical database for Hong Kong Cantonese offers phonological and orthographic information, frequency measures, and lexically neighborhood information for lexicals in HKC. |
| Approach: | They introduce a lexical database for Hong Kong Cantonese that offers phonological and orthographic information, frequency measures, and lexically neighborhood information for lexicals in HKC. |
| Outcome: | The proposed lexical database for Hong Kong Cantonese offers phonological and orthographic information, frequency measures, and lexically neighborhood information. |
ManaTTS Persian: a recipe for creating TTS datasets for lower resource languages (2025.naacl-long)
Copied to clipboard
| Challenge: | a new text-to-speech system is needed for visual impairments and the visually impaired . a text-based system is not available for all users, and is therefore limited to a limited audience. |
| Approach: | They propose to use ManaTTS, the most extensive publicly accessible Persian corpus . they use a fully transparent, MIT-licensed pipeline to collect transcribed speech datasets . |
| Outcome: | The proposed framework is the most extensive publicly accessible single-speaker Persian corpus . it includes tools for sentence tokenization, bounded audio segmentation, and forced alignment method . |