Challenge: a phonetic balance in code-mixed Hindi-English corpus has been created . code-switching is a common phenomenon in multilingual and bilingual communities .
Approach: They propose to create a phonetically balanced read speech corpus of code-mixed Hindi-English . they use a method to select sentences that contain triphones lower in frequency than a threshold .
Outcome: The proposed corpus is phonetically balanced with a large code-mixed reference corpus.

Similar Papers

Aggression-annotated Corpus of Hindi-English Code-mixed Data (L18-1)

Copied to clipboard

Challenge: a number of incidents of aggression and related events have increased over the web . the reach and extent of the Internet has given these events unprecedented power and influence to affect the lives of billions of people.
Approach: They propose to develop an aggression tagset and an annotated corpus of Hindi-English code-mixed data from two of the most popular social networking / social media platforms in India -Twitter and Facebook.
Outcome: The proposed dataset contains approximately 18k tweets and 21k facebook comments and is being released for further research in the field.
MHE: Code-Mixed Corpora for Similar Language Identification (2022.lrec-1)

Copied to clipboard

Challenge: a new corpus of code-mixed data-sets is presented for similar language identification . the data-settings are based on a more-resourced minority language, Magahi .
Approach: They propose a Magahi-Hindi-English code-mixed corpus for similar language identification . they discuss the complexity of the data-set and provide a few baselines .
Outcome: The proposed corpus provides a language id at two levels: word and sentence.
An Application for Building a Polish Telephone Speech Corpus (L18-1)

Copied to clipboard

Challenge: Specifically, we describe a tool designed to improve our Automatic Speech Recognition system performance.
Approach: They propose to build a tool for speech corpus collection of a specific domain content.
Outcome: The proposed tool can be used to gather 63 hours of speech recordings across several domains and achieve lower WER in two grammar-based speech recognition tasks.
hinglishNorm - A Corpus of Hindi-English Code Mixed Sentences for Text Normalization (2020.coling-industry)

Copied to clipboard

Challenge: hinglishNorm is a human annotated corpus of Hindi-English code-mixed sentences for text normalization task.
Approach: They propose to annotate sentences in Hindi-English code-mixed sentences using a human annotated normalized form.
Outcome: The proposed corpus contains 13494 segments annotated for text normalization.
Open-source Multi-speaker Speech Corpora for Building Gujarati, Kannada, Malayalam, Marathi, Tamil and Telugu Speech Synthesis Systems (2020.lrec-1)

Copied to clipboard

Challenge: We present free high quality multi-speaker speech corpora for Gujarati, Kannada, Malayalam, Marathi, Tamil and Telugu . the datasets are primarily intended for use in text-to-speech applications, such as constructing multilingual voices or language adaptation.
Approach: They present a free high quality multi-speaker speech corpora for Gujarati, Kannada, Malayalam, Marathi, Tamil and Telugu . they use it to build a multilingual text-to-speech model that can be scaled to other languages of interest.
Outcome: The proposed model produces good quality voices with MOS > 3.6 for all the languages tested.
Automatic Partitioning of a Code-Switched Speech Corpus Using Mixed-Integer Programming (2024.lrec-main)

Copied to clipboard

Challenge: Currently, partitioning speech corpora is done by hand, but this is not feasible for the dataset under investigation.
Approach: They propose to partition a 41.6-hour corpus of code-switched speech into training, development and testing partitions using mixed-integer linear programming.
Outcome: The proposed method allows to partition a 41.6-hour corpus of code-switched speech into training, development and testing partitions while maintaining a fixed number of speakers and a specific amount of codeswitching speech in the development and test partitions.
CoMuMDR: Code-mixed Multi-modal Multi-domain corpus for Discourse paRsing in conversations (2025.findings-acl)

Copied to clipboard

Challenge: Discourse parsing datasets based on conversations are restricted to a single domain . a lack of discourse structures in audio-based conversations is a challenge .
Approach: They introduce CoMuMDR: Code-mixed Multi-modal Multi-domain corpus for Discourse parsing in conversations.
Outcome: The proposed corpus is code-mixed in Hindi and English and annotated with nine discourse relations.
Urdu Pitch Accents and Intonation Patterns in Spontaneous Conversational Speech (2020.lrec-1)

Copied to clipboard

Challenge: Recent studies of Urdu intonation describe scripted and laboratory speech .
Approach: They summarise Urdu pitch accents and their intonation patterns using a simplified version of the Rhythm and Pitch labelling system and a simple RAP system.
Outcome: The analysis of a hand-labelled telephone conversation shows that low pitch accents play an important role in Urdu spontaneous speech.
IndicSpeech: Text-to-Speech Corpus for Indian Languages (2020.lrec-1)

Copied to clipboard

Challenge: India has 22 languages, each of them being spoken by over a million people . the current state of the art text-to-speech systems for Indian languages are lacking in the multimedia domain .
Approach: They propose to train a state-of-the-art TTS system for Hindi, Malayalam and Bengali and publish the results.
Outcome: The proposed system trains neural text-to-speech systems for Hindi, Malayalam and Bengali and makes them publicly available.
Humor Detection in English-Hindi Code-Mixed Social Media Content : Corpus and Baseline System (L18-1)

Copied to clipboard

Challenge: a growing number of social media users are using code-mixing to detect humor . linguistics researchers are looking for methods to detect humorous content in text .
Approach: They analyze a corpus of English-Hindi code-mixed tweets annotated with humorous(H) tags.
Outcome: The proposed method detects humor in code-mixed tweets in English-Hindi.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations