Phonetically Balanced Code-Mixed Speech Corpus for Hindi-English Automatic Speech Recognition (L18-1)
Copied to clipboard
Ayushi Pandey, Brij Mohan Lal Srivastava, Rohit Kumar, Bhanu Teja Nellore, Kasi Sai Teja, Suryakanth V. Gangashetty
| Challenge: | a phonetic balance in code-mixed Hindi-English corpus has been created . code-switching is a common phenomenon in multilingual and bilingual communities . |
| Approach: | They propose to create a phonetically balanced read speech corpus of code-mixed Hindi-English . they use a method to select sentences that contain triphones lower in frequency than a threshold . |
| Outcome: | The proposed corpus is phonetically balanced with a large code-mixed reference corpus. |
Similar Papers
Aggression-annotated Corpus of Hindi-English Code-mixed Data (L18-1)
Copied to clipboard
| Challenge: | a number of incidents of aggression and related events have increased over the web . the reach and extent of the Internet has given these events unprecedented power and influence to affect the lives of billions of people. |
| Approach: | They propose to develop an aggression tagset and an annotated corpus of Hindi-English code-mixed data from two of the most popular social networking / social media platforms in India -Twitter and Facebook. |
| Outcome: | The proposed dataset contains approximately 18k tweets and 21k facebook comments and is being released for further research in the field. |
MHE: Code-Mixed Corpora for Similar Language Identification (2022.lrec-1)
Copied to clipboard
| Challenge: | a new corpus of code-mixed data-sets is presented for similar language identification . the data-settings are based on a more-resourced minority language, Magahi . |
| Approach: | They propose a Magahi-Hindi-English code-mixed corpus for similar language identification . they discuss the complexity of the data-set and provide a few baselines . |
| Outcome: | The proposed corpus provides a language id at two levels: word and sentence. |
An Application for Building a Polish Telephone Speech Corpus (L18-1)
Copied to clipboard
| Challenge: | Specifically, we describe a tool designed to improve our Automatic Speech Recognition system performance. |
| Approach: | They propose to build a tool for speech corpus collection of a specific domain content. |
| Outcome: | The proposed tool can be used to gather 63 hours of speech recordings across several domains and achieve lower WER in two grammar-based speech recognition tasks. |
hinglishNorm - A Corpus of Hindi-English Code Mixed Sentences for Text Normalization (2020.coling-industry)
Copied to clipboard
| Challenge: | hinglishNorm is a human annotated corpus of Hindi-English code-mixed sentences for text normalization task. |
| Approach: | They propose to annotate sentences in Hindi-English code-mixed sentences using a human annotated normalized form. |
| Outcome: | The proposed corpus contains 13494 segments annotated for text normalization. |
Open-source Multi-speaker Speech Corpora for Building Gujarati, Kannada, Malayalam, Marathi, Tamil and Telugu Speech Synthesis Systems (2020.lrec-1)
Copied to clipboard
Fei He, Shan-Hui Cathy Chu, Oddur Kjartansson, Clara Rivera, Anna Katanova, Alexander Gutkin, Isin Demirsahin, Cibu Johny, Martin Jansche, Supheakmungkol Sarin, Knot Pipatsrisawat
| Challenge: | We present free high quality multi-speaker speech corpora for Gujarati, Kannada, Malayalam, Marathi, Tamil and Telugu . the datasets are primarily intended for use in text-to-speech applications, such as constructing multilingual voices or language adaptation. |
| Approach: | They present a free high quality multi-speaker speech corpora for Gujarati, Kannada, Malayalam, Marathi, Tamil and Telugu . they use it to build a multilingual text-to-speech model that can be scaled to other languages of interest. |
| Outcome: | The proposed model produces good quality voices with MOS > 3.6 for all the languages tested. |
Automatic Partitioning of a Code-Switched Speech Corpus Using Mixed-Integer Programming (2024.lrec-main)
Copied to clipboard
| Challenge: | Currently, partitioning speech corpora is done by hand, but this is not feasible for the dataset under investigation. |
| Approach: | They propose to partition a 41.6-hour corpus of code-switched speech into training, development and testing partitions using mixed-integer linear programming. |
| Outcome: | The proposed method allows to partition a 41.6-hour corpus of code-switched speech into training, development and testing partitions while maintaining a fixed number of speakers and a specific amount of codeswitching speech in the development and test partitions. |
CoMuMDR: Code-mixed Multi-modal Multi-domain corpus for Discourse paRsing in conversations (2025.findings-acl)
Copied to clipboard
| Challenge: | Discourse parsing datasets based on conversations are restricted to a single domain . a lack of discourse structures in audio-based conversations is a challenge . |
| Approach: | They introduce CoMuMDR: Code-mixed Multi-modal Multi-domain corpus for Discourse parsing in conversations. |
| Outcome: | The proposed corpus is code-mixed in Hindi and English and annotated with nine discourse relations. |
Urdu Pitch Accents and Intonation Patterns in Spontaneous Conversational Speech (2020.lrec-1)
Copied to clipboard
| Challenge: | Recent studies of Urdu intonation describe scripted and laboratory speech . |
| Approach: | They summarise Urdu pitch accents and their intonation patterns using a simplified version of the Rhythm and Pitch labelling system and a simple RAP system. |
| Outcome: | The analysis of a hand-labelled telephone conversation shows that low pitch accents play an important role in Urdu spontaneous speech. |
IndicSpeech: Text-to-Speech Corpus for Indian Languages (2020.lrec-1)
Copied to clipboard
| Challenge: | India has 22 languages, each of them being spoken by over a million people . the current state of the art text-to-speech systems for Indian languages are lacking in the multimedia domain . |
| Approach: | They propose to train a state-of-the-art TTS system for Hindi, Malayalam and Bengali and publish the results. |
| Outcome: | The proposed system trains neural text-to-speech systems for Hindi, Malayalam and Bengali and makes them publicly available. |
Humor Detection in English-Hindi Code-Mixed Social Media Content : Corpus and Baseline System (L18-1)
Copied to clipboard
| Challenge: | a growing number of social media users are using code-mixing to detect humor . linguistics researchers are looking for methods to detect humorous content in text . |
| Approach: | They analyze a corpus of English-Hindi code-mixed tweets annotated with humorous(H) tags. |
| Outcome: | The proposed method detects humor in code-mixed tweets in English-Hindi. |