CanVEC - the Canberra Vietnamese-English Code-switching Natural Speech Corpus (2020.lrec-1)
Copied to clipboard
| Challenge: | Using monolingual tools, code-switching is a problem in the natural language processing community. |
| Approach: | They introduce the Canberra Vietnamese-English Code-switching corpus (CanVEC) which is an original corpus of mixed speech annotated with language information, part of speech tags and Vietnamese translations. |
| Outcome: | The proposed corpus was annotated with language information, part of speech tags and Vietnamese translations using pipelining and monolingual toolkits. |
Similar Papers
VietMix: A Naturally-Occurring Parallel Corpus and Augmentation Framework for Vietnamese-English Code-Mixed Machine Translation (2026.eacl-long)
Copied to clipboard
| Challenge: | Existing approaches to machine translation (MT) systems degrade when faced with code-mixed text. |
| Approach: | They propose a system that can augment Vietnamese-English code-mixed text with iterative fine-tuning and targeted filtering. |
| Outcome: | The proposed framework outperforms strong back-translation baselines and improves zero-shot models by up to +11.9 points. |
ASCEND: A Spontaneous Chinese-English Dataset for Code-switching in Multi-turn Conversation (2022.lrec-1)
Copied to clipboard
Holy Lovenia, Samuel Cahyawijaya, Genta Winata, Peng Xu, Yan Xu, Zihan Liu, Rita Frieske, Tiezheng Yu, Wenliang Dai, Elham J. Barezi, Qifeng Chen, Xiaojuan Ma, Bertram Shi, Pascale Fung
| Challenge: | Code-switching is a speech phenomenon occurring when a speaker switches language during a conversation. |
| Approach: | They propose to collect Mandarin Chinese-English code-switching corpus from read speech rather than spontaneous speech to address this phenomenon. |
| Outcome: | ASCEND consists of 10.62 hours of clean speech, collected from 23 bilingual speakers of Chinese and English. |
CoSSAT: Code-Switched Speech Annotation Tool (D19-59)
Copied to clipboard
| Challenge: | Code-switching is a phenomenon that occurs in multilingual societies where speakers who are fluent in two or more languages switch between these languages in the same conversation or utterance. |
| Approach: | They propose an interface which helps annotators transcribe code-switched speech faster, more easily and more accurately than a traditional interface. |
| Outcome: | The proposed interface can be used by 10 users to transcribe Hindi-English code-switched speech faster, easier and more accurately than a traditional interface. |
A First South African Corpus of Multilingual Code-switched Soap Opera Speech (L18-1)
Copied to clipboard
| Challenge: | a corpus of code-switched speech from soap operas is compiled from soaps . the corpus contains 14.3 hours of annotated and segmented speech . |
| Approach: | They propose a speech corpus containing multilingual code-switching from soap operas . the corpus contains English, isiZulu, isisXhosa, Setswana and Sesotho speech . |
| Outcome: | The corpus contains 14.3 hours of annotated and segmented speech from soap operas . the speech rate is 1.22 to 1.83 times higher than prompted speech in the same languages . |
TweetTaglish: A Dataset for Investigating Tagalog-English Code-Switching (2022.lrec-1)
Copied to clipboard
| Challenge: | a large dataset is available to study Tagalog-English code-switching in low-resource settings. |
| Approach: | They propose to use a large dataset to investigate Tagalog-English code-switching . they use linguistic data from Tagalogue and Tagalit-English to investigate their results . |
| Outcome: | The proposed dataset achieves a strong performance benchmark for Tagalog-English code-switching. |
Collection and Analysis of Code-switch Egyptian Arabic-English Speech Corpus (L18-1)
Copied to clipboard
| Challenge: | despite of the great demand, there is still a huge shortage in available corpora for dialectal languages and code-switched speech. |
| Approach: | They collect conversational Egyptian Arabic spontaneous speech, extract transcriptions and analyze it from a code-switching perspective. |
| Outcome: | The authors collect conversational Egyptian Arabic spontaneous speech, extract transcriptions and analyze speech from the code-switching perspective. |
ViNLI: A Vietnamese Corpus for Studies on Open-Domain Natural Language Inference (2022.coling-1)
Copied to clipboard
| Challenge: | a large-scale corpus is needed for studies on natural language inference (NLI) for Vietnamese, which can be considered a low-resource language. |
| Approach: | They propose a corpus for evaluating Vietnamese natural language inference models . they use a human-annotated corpus extracted from more than 800 online news articles . |
| Outcome: | The ViNLI corpus is created and evaluated with a strict process of quality control . the best system performance is still far from human performance (a 14.20% gap in accuracy). |
Analyzing the Role of Part-of-Speech in Code-Switching: A Corpus-Based Study (2024.findings-eacl)
Copied to clipboard
| Challenge: | Code-switching (CS) is a common linguistic phenomenon wherein speakers fluidly transition between languages in conversation. |
| Approach: | They propose to use a part-of-speech (POS)-based analysis of Spanish-English and Mandarin-English corpora to examine the propensity of bilinguals to engage in CS. |
| Outcome: | The findings confirm the existence of a statistically significant connection between POS and the likelihood of CS across language pairs, but show that it diminishes as tokens distance themselves from CS instances. |
SpiCE: A New Open-Access Corpus of Conversational Bilingual Speech in Cantonese and English (2020.lrec-1)
Copied to clipboard
| Challenge: | SpiCE is a corpus of conversational Cantonese-English bilingual speech recorded in Vancouver, Canada . the corpus includes high-quality recordings of 34 early bilinguals in both English and Cantoneses . |
| Approach: | They describe the design, collection, orthographic transcription, and phonetic annotation of SpiCE . the corpus includes high-quality recordings of 34 early bilinguals in both English and Cantonese . |
| Outcome: | The SpiCE corpus includes high-quality recordings of 34 early bilinguals in both English and Cantonese . the corpus will promote bilingualism research for a typologically distinct pair of languages . |
End-to-End Speech Translation for Code Switched Speech (2022.findings-acl)
Copied to clipboard
Orion Weller, Matthias Sperber, Telmo Pires, Hendra Setiawan, Christian Gollan, Dominic Telaar, Matthias Paulik
| Challenge: | Code switching (CS) is the phenomenon of interchangeably using words and phrases from different languages. |
| Approach: | They propose a new ST corpus that extends the joint transcription and translation setup. |
| Outcome: | The proposed model performs well even when no training data is used. |