CoSSAT: Code-Switched Speech Annotation Tool (D19-59)

Copied to clipboard

Challenge: Code-switching is a phenomenon that occurs in multilingual societies where speakers who are fluent in two or more languages switch between these languages in the same conversation or utterance.
Approach: They propose an interface which helps annotators transcribe code-switched speech faster, more easily and more accurately than a traditional interface.
Outcome: The proposed interface can be used by 10 users to transcribe Hindi-English code-switched speech faster, easier and more accurately than a traditional interface.

Similar Papers

Automatic Identification of Code-Switching Functions in Speech Transcripts (2023.findings-acl)

Copied to clipboard

Challenge: Code-switching, or switching between languages, occurs for many reasons and has important linguistic, sociological, and cultural implications.
Approach: They build a system to identify a wide range of functions for which speakers code-switch in everyday speech with an accuracy of 75% . they use a dataset of Hindi-English code-witched data to analyze their results .
Outcome: The proposed system can identify a wide range of functions for which speakers code-switch in everyday speech, with an accuracy of 75% across all functions.
PolyWER: A Holistic Evaluation Framework for Code-Switched Speech Recognition (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing methods for measuring accuracy, such as Word Error Rate (WER), are too strict to address this challenge.
Approach: They propose a framework for evaluating speech recognition systems to handle language-mixing by appending annotations to a publicly available Arabic-English code-switched dataset.
Outcome: The proposed framework evaluates speech recognition systems against human judgement and a publicly available Arabic-English code-switched dataset.
End-to-End Speech Translation for Code Switched Speech (2022.findings-acl)

Copied to clipboard

Challenge: Code switching (CS) is the phenomenon of interchangeably using words and phrases from different languages.
Approach: They propose a new ST corpus that extends the joint transcription and translation setup.
Outcome: The proposed model performs well even when no training data is used.
CoSTA: Code-Switched Speech Translation using Aligned Speech-Text Interleaving (2025.coling-main)

Copied to clipboard

Challenge: More than half of the world's population is presumed to be bilingual . spoken translation of code-switched speech has been under-explored .
Approach: They propose an end-to-end model architecture CoSTA that scaffolds on pretrained ASR and MT modules.
Outcome: The proposed model outperforms existing models by 3.5 BLEU points in spoken translation of code-switched speech.
Collection and Analysis of Code-switch Egyptian Arabic-English Speech Corpus (L18-1)

Copied to clipboard

Challenge: despite of the great demand, there is still a huge shortage in available corpora for dialectal languages and code-switched speech.
Approach: They collect conversational Egyptian Arabic spontaneous speech, extract transcriptions and analyze it from a code-switching perspective.
Outcome: The authors collect conversational Egyptian Arabic spontaneous speech, extract transcriptions and analyze speech from the code-switching perspective.
CoVoSwitch: Machine Translation of Synthetic Code-Switched Text Based on Intonation Units (2024.acl-srw)

Copied to clipboard

Challenge: Multilingual code-switching research is often hindered by the lack and linguistically biased status of available datasets.
Approach: They synthesize code-switching data by replacing intonation units detected through PSST, a speech segmentation model fine-tuned from OpenAI’s Whisper, using a language-to-text translation dataset, CoVoST 2.
Outcome: The proposed model outperforms two monolingual models and is better at code-switching translation into English than non-English.
CoVoST: A Diverse Multilingual Speech-To-Text Translation Corpus (2020.lrec-1)

Copied to clipboard

Challenge: Existing datasets involve language pairs with English as source language, are low resource or lack labeled data.
Approach: They propose a multilingual speech-to-text translation corpus from 11 languages into English . they provide empirical evidence of the quality of the data and provide initial benchmarks .
Outcome: The proposed model is the first end-to-end multilingual model for spoken language translation.
Part-of-Speech Tagging for Code-Switched, Transliterated Texts without Explicit Language Identification (D18-1)

Copied to clipboard

Challenge: Code-switching is a challenge for NLP due to the lack of representative data for training models.
Approach: They propose a model that is trained exclusively on monolingual resources but can be applied to unseen code-switched text at inference time.
Outcome: The proposed model outperforms standard models on Hindi-English part-of-speech tagging and on unannotated code-switched text with alternate scripts.
Universal Dependency Parsing for Hindi-English Code-Switching (N18-1)

Copied to clipboard

Challenge: Code-switching data often need additional processes such as language identification, normalization and/or back-transliteration to be processed.
Approach: They propose a neural stacking model that leverages part-of-speech tags and syntactic tree annotations in tweets to parse code-switching data.
Outcome: The proposed model is 1.5% better than the augmented model and 3.8% better than one which uses first-best normalization and/or back-transliteration.
UniCoM: A Universal Code-Switching Speech Generator (2025.findings-emnlp)

Copied to clipboard

Challenge: Code-switching (CS) is a common phenomenon in real-world conversations and poses significant challenges for multilingual speech technology.
Approach: They propose a pipeline for generating high-quality, natural CS samples without altering sentence semantics.
Outcome: The proposed pipeline generates high-quality, natural CS samples without altering sentence semantics without alteration of sentence semantic.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations