MYCanCor: A Video Corpus of spoken Malaysian Cantonese (L18-1)

Copied to clipboard

Challenge: The corpus consists of 20 hours of video recordings of spontaneous talk-in-interaction typically involving 2-4 speakers.
Approach: the corpus consists of 20 hours of video recordings of spontaneous talk-in-interaction typically involving 2-4 speakers.
Outcome: the corpus consists of 20 hours of video recordings of spontaneous talk-in-interaction typically involving 2-4 speakers.

Similar Papers

WeCanTalk: A New Multi-language, Multi-modal Resource for Speaker Recognition (2022.lrec-1)

Copied to clipboard

Challenge: The WeCanTalk corpus is a multi-modal, multi-language resource for speaker recognition.
Approach: The WeCanTalk corpus is a multi-modal resource for speaker recognition.
Outcome: The corpus contains data from 202 native speakers in Hong Kong who were fluent in at least one other language.
CantoMap: a Hong Kong Cantonese MapTask Corpus (2020.lrec-1)

Copied to clipboard

Challenge: a corpus of connected spoken Hong Kong Cantonese is constructed to study the phonology and semantics of the language.
Approach: They propose to build a corpus of connected spoken Hong Kong Cantonese with phonemic transcription and controlled elicitation tasks.
Outcome: The proposed corpus contains 768 minutes of recordings and transcripts of forty speakers.
PyCantonese: Cantonese Linguistics and NLP in Python (2022.lrec-1)

Copied to clipboard

Challenge: a limited number of Cantonese-specific datasets are available for PyCantones.
Approach: They introduce PyCantonese, an open-source Python library for Cantonesi linguistics and natural language processing.
Outcome: The proposed library is open-source and available for free for all purposes, including commercial ones.
SpiCE: A New Open-Access Corpus of Conversational Bilingual Speech in Cantonese and English (2020.lrec-1)

Copied to clipboard

Challenge: SpiCE is a corpus of conversational Cantonese-English bilingual speech recorded in Vancouver, Canada . the corpus includes high-quality recordings of 34 early bilinguals in both English and Cantoneses .
Approach: They describe the design, collection, orthographic transcription, and phonetic annotation of SpiCE . the corpus includes high-quality recordings of 34 early bilinguals in both English and Cantonese .
Outcome: The SpiCE corpus includes high-quality recordings of 34 early bilinguals in both English and Cantonese . the corpus will promote bilingualism research for a typologically distinct pair of languages .
A Crowdsourced Open-Source Kazakh Speech Corpus and Initial Speech Recognition Baseline (2021.eacl-main)

Copied to clipboard

Challenge: The Kazakh speech corpus contains over 153,000 utterances spoken by participants from different regions and age groups, as well as both genders.
Approach: They propose to build an open-source Kazakh speech corpus for the Kazakh language that contains over 153,000 transcribed audio . they describe the data collection and preprocessing procedures followed by a description of the database specifications.
Outcome: The Kazakh speech corpus contains over 153,000 utterances spoken by participants from different regions and age groups, as well as both genders.
CanVEC - the Canberra Vietnamese-English Code-switching Natural Speech Corpus (2020.lrec-1)

Copied to clipboard

Challenge: Using monolingual tools, code-switching is a problem in the natural language processing community.
Approach: They introduce the Canberra Vietnamese-English Code-switching corpus (CanVEC) which is an original corpus of mixed speech annotated with language information, part of speech tags and Vietnamese translations.
Outcome: The proposed corpus was annotated with language information, part of speech tags and Vietnamese translations using pipelining and monolingual toolkits.
Automatic Speech Recognition Datasets in Cantonese: A Survey and New Dataset (2022.lrec-1)

Copied to clipboard

Challenge: In this paper, we address the problem of data scarcity for the Hong Kong Cantonese language . due to the popularization of deep learning, ASR technology has led to a significant improvement in recognizing many languages.
Approach: They propose to use a dataset to analyze the data available for the Hong Kong Cantonese language . they use zh-HK as a source and a state-of-the-art ASR model to build a powerful model .
Outcome: The proposed model improves on the biggest existing dataset, Common Voice zh-HK.
My Science Tutor (MyST)–a Large Corpus of Children’s Conversational Speech (2024.lrec-main)

Copied to clipboard

Challenge: a 13-year project was conducted between 2007 and 2019 to improve students' learning proficiency in elementary school science using conversational multimedia virtual tutor, Marni.
Approach: They propose to use the corpus-name corpus to improve automatic speech recognition models and algorithms by training and developing a model on the training and development portion of the corpuse.
Outcome: The corpus comprises 400 hours of speech, spanning some 230K utterances spread across about 10,500 virtual tutor sessions.
Cifu: a Frequency Lexicon of Hong Kong Cantonese (2020.lrec-1)

Copied to clipboard

Challenge: lexical database for Hong Kong Cantonese offers phonological and orthographic information, frequency measures, and lexically neighborhood information for lexicals in HKC.
Approach: They introduce a lexical database for Hong Kong Cantonese that offers phonological and orthographic information, frequency measures, and lexically neighborhood information for lexicals in HKC.
Outcome: The proposed lexical database for Hong Kong Cantonese offers phonological and orthographic information, frequency measures, and lexically neighborhood information.
ManaTTS Persian: a recipe for creating TTS datasets for lower resource languages (2025.naacl-long)

Copied to clipboard

Challenge: a new text-to-speech system is needed for visual impairments and the visually impaired . a text-based system is not available for all users, and is therefore limited to a limited audience.
Approach: They propose to use ManaTTS, the most extensive publicly accessible Persian corpus . they use a fully transparent, MIT-licensed pipeline to collect transcribed speech datasets .
Outcome: The proposed framework is the most extensive publicly accessible single-speaker Persian corpus . it includes tools for sentence tokenization, bounded audio segmentation, and forced alignment method .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations