Challenge: Using monolingual tools, code-switching is a problem in the natural language processing community.
Approach: They introduce the Canberra Vietnamese-English Code-switching corpus (CanVEC) which is an original corpus of mixed speech annotated with language information, part of speech tags and Vietnamese translations.
Outcome: The proposed corpus was annotated with language information, part of speech tags and Vietnamese translations using pipelining and monolingual toolkits.

Similar Papers

VietMix: A Naturally-Occurring Parallel Corpus and Augmentation Framework for Vietnamese-English Code-Mixed Machine Translation (2026.eacl-long)

Copied to clipboard

Challenge: Existing approaches to machine translation (MT) systems degrade when faced with code-mixed text.
Approach: They propose a system that can augment Vietnamese-English code-mixed text with iterative fine-tuning and targeted filtering.
Outcome: The proposed framework outperforms strong back-translation baselines and improves zero-shot models by up to +11.9 points.
ASCEND: A Spontaneous Chinese-English Dataset for Code-switching in Multi-turn Conversation (2022.lrec-1)

Copied to clipboard

Challenge: Code-switching is a speech phenomenon occurring when a speaker switches language during a conversation.
Approach: They propose to collect Mandarin Chinese-English code-switching corpus from read speech rather than spontaneous speech to address this phenomenon.
Outcome: ASCEND consists of 10.62 hours of clean speech, collected from 23 bilingual speakers of Chinese and English.
CoSSAT: Code-Switched Speech Annotation Tool (D19-59)

Copied to clipboard

Challenge: Code-switching is a phenomenon that occurs in multilingual societies where speakers who are fluent in two or more languages switch between these languages in the same conversation or utterance.
Approach: They propose an interface which helps annotators transcribe code-switched speech faster, more easily and more accurately than a traditional interface.
Outcome: The proposed interface can be used by 10 users to transcribe Hindi-English code-switched speech faster, easier and more accurately than a traditional interface.
A First South African Corpus of Multilingual Code-switched Soap Opera Speech (L18-1)

Copied to clipboard

Challenge: a corpus of code-switched speech from soap operas is compiled from soaps . the corpus contains 14.3 hours of annotated and segmented speech .
Approach: They propose a speech corpus containing multilingual code-switching from soap operas . the corpus contains English, isiZulu, isisXhosa, Setswana and Sesotho speech .
Outcome: The corpus contains 14.3 hours of annotated and segmented speech from soap operas . the speech rate is 1.22 to 1.83 times higher than prompted speech in the same languages .
TweetTaglish: A Dataset for Investigating Tagalog-English Code-Switching (2022.lrec-1)

Copied to clipboard

Challenge: a large dataset is available to study Tagalog-English code-switching in low-resource settings.
Approach: They propose to use a large dataset to investigate Tagalog-English code-switching . they use linguistic data from Tagalogue and Tagalit-English to investigate their results .
Outcome: The proposed dataset achieves a strong performance benchmark for Tagalog-English code-switching.
Collection and Analysis of Code-switch Egyptian Arabic-English Speech Corpus (L18-1)

Copied to clipboard

Challenge: despite of the great demand, there is still a huge shortage in available corpora for dialectal languages and code-switched speech.
Approach: They collect conversational Egyptian Arabic spontaneous speech, extract transcriptions and analyze it from a code-switching perspective.
Outcome: The authors collect conversational Egyptian Arabic spontaneous speech, extract transcriptions and analyze speech from the code-switching perspective.
ViNLI: A Vietnamese Corpus for Studies on Open-Domain Natural Language Inference (2022.coling-1)

Copied to clipboard

Challenge: a large-scale corpus is needed for studies on natural language inference (NLI) for Vietnamese, which can be considered a low-resource language.
Approach: They propose a corpus for evaluating Vietnamese natural language inference models . they use a human-annotated corpus extracted from more than 800 online news articles .
Outcome: The ViNLI corpus is created and evaluated with a strict process of quality control . the best system performance is still far from human performance (a 14.20% gap in accuracy).
Analyzing the Role of Part-of-Speech in Code-Switching: A Corpus-Based Study (2024.findings-eacl)

Copied to clipboard

Challenge: Code-switching (CS) is a common linguistic phenomenon wherein speakers fluidly transition between languages in conversation.
Approach: They propose to use a part-of-speech (POS)-based analysis of Spanish-English and Mandarin-English corpora to examine the propensity of bilinguals to engage in CS.
Outcome: The findings confirm the existence of a statistically significant connection between POS and the likelihood of CS across language pairs, but show that it diminishes as tokens distance themselves from CS instances.
SpiCE: A New Open-Access Corpus of Conversational Bilingual Speech in Cantonese and English (2020.lrec-1)

Copied to clipboard

Challenge: SpiCE is a corpus of conversational Cantonese-English bilingual speech recorded in Vancouver, Canada . the corpus includes high-quality recordings of 34 early bilinguals in both English and Cantoneses .
Approach: They describe the design, collection, orthographic transcription, and phonetic annotation of SpiCE . the corpus includes high-quality recordings of 34 early bilinguals in both English and Cantonese .
Outcome: The SpiCE corpus includes high-quality recordings of 34 early bilinguals in both English and Cantonese . the corpus will promote bilingualism research for a typologically distinct pair of languages .
End-to-End Speech Translation for Code Switched Speech (2022.findings-acl)

Copied to clipboard

Challenge: Code switching (CS) is the phenomenon of interchangeably using words and phrases from different languages.
Approach: They propose a new ST corpus that extends the joint transcription and translation setup.
Outcome: The proposed model performs well even when no training data is used.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations