The French-Algerian Code-Switching Triggered audio corpus (FACST) (L18-1)

Copied to clipboard

Challenge: The French Algerian Code-Switching Triggered corpus is a corpus of spontaneous CS utterances . it is used to support linguistic and phonetic studies in phonetics and prosody .
Approach: They propose to use a triggering protocol to elicit CS in natural conversations . they propose to do data segmentation and annotation in each language .
Outcome: The proposed corpus is based on a code-switching protocol and is well-suited for linguistic and acoustic-phonetic studies.

Similar Papers

Cairo Student Code-Switch (CSCS) Corpus: An Annotated Egyptian Arabic-English Corpus (2020.lrec-1)

Copied to clipboard

Challenge: Code-switching is a phenomenon commonly observed in the Arabicspeaking world . there is still a huge gap in the available resources and NLP applications .
Approach: They propose a corpus of Egyptian- Arabic code-switch speech data that is fully tokenized, lemmatized and annotated for part-of-speech tags.
Outcome: The proposed corpus of Egyptian- Arabic code-switch speech data is fully tokenized, lemmatized and annotated for part-of-speech tags.
Collection and Analysis of Code-switch Egyptian Arabic-English Speech Corpus (L18-1)

Copied to clipboard

Challenge: despite of the great demand, there is still a huge shortage in available corpora for dialectal languages and code-switched speech.
Approach: They collect conversational Egyptian Arabic spontaneous speech, extract transcriptions and analyze it from a code-switching perspective.
Outcome: The authors collect conversational Egyptian Arabic spontaneous speech, extract transcriptions and analyze speech from the code-switching perspective.
UniCoM: A Universal Code-Switching Speech Generator (2025.findings-emnlp)

Copied to clipboard

Challenge: Code-switching (CS) is a common phenomenon in real-world conversations and poses significant challenges for multilingual speech technology.
Approach: They propose a pipeline for generating high-quality, natural CS samples without altering sentence semantics.
Outcome: The proposed pipeline generates high-quality, natural CS samples without altering sentence semantics without alteration of sentence semantic.
ArzEn: A Speech Corpus for Code-switched Egyptian Arabic-English (2020.lrec-1)

Copied to clipboard

Challenge: a corpus of Arabic-English code-switching (CS) spontaneous speech is collected in an Egyptian university soundproof room . the language in Egypt is rather complex and poses many challenges to natural language processing (NLP)
Approach: They present an Egyptian Arabic-English code-switching (CS) spontaneous speech corpus.
Outcome: The proposed corpus is designed to be used in automatic speech recognition systems . it provides a useful resource for analyzing the CS phenomenon from linguistic, sociological, and psychological perspectives.
A Survey of Code-switched Arabic NLP: Progress, Challenges, and Future Directions (2025.coling-main)

Copied to clipboard

Challenge: Code-switching (CSW) is a common linguistic phenomenon in multilingual societies . current literature on CSW in the arab world is limited to the Arabic language .
Approach: They present a review of the literature in the field of code-switched Arabic NLP . they propose recommendations for future research .
Outcome: This review provides a broad perspective on the current literature in the field of code-switched Arabic NLP . it also provides recommendations for future research .
Identifying Sentiments in Algerian Code-switched User-generated Comments (2020.lrec-1)

Copied to clipboard

Challenge: a recent study has focused on sentiment analysis for the Arabic variety, but it has been extended to other domains.
Approach: They build a corpus of 36,000 code-switched user-generated comments annotated for sentiments in Algerian Arabic.
Outcome: The proposed model performs better on unedited code-switched and unbalanced data across sentiment classes.
FAB: The French Absolute Beginner Corpus for Pronunciation Training (2020.lrec-1)

Copied to clipboard

Challenge: French Absolute Beginner corpus is intended for the development and study of Computer-Assisted Pronunciation Training (CAPT) tools for absolute beginner learners.
Approach: They introduce the French Absolute Beginner (FAB) speech corpus which is intended for the development and study of Computer-Assisted Pronunciation Training tools for absolute beginner learners.
Outcome: The proposed corpus is intended for the development and study of Computer-Assisted Pronunciation Training tools for absolute beginner learners.
Automatic Identification of Code-Switching Functions in Speech Transcripts (2023.findings-acl)

Copied to clipboard

Challenge: Code-switching, or switching between languages, occurs for many reasons and has important linguistic, sociological, and cultural implications.
Approach: They build a system to identify a wide range of functions for which speakers code-switch in everyday speech with an accuracy of 75% . they use a dataset of Hindi-English code-witched data to analyze their results .
Outcome: The proposed system can identify a wide range of functions for which speakers code-switch in everyday speech, with an accuracy of 75% across all functions.
Analyzing the Role of Part-of-Speech in Code-Switching: A Corpus-Based Study (2024.findings-eacl)

Copied to clipboard

Challenge: Code-switching (CS) is a common linguistic phenomenon wherein speakers fluidly transition between languages in conversation.
Approach: They propose to use a part-of-speech (POS)-based analysis of Spanish-English and Mandarin-English corpora to examine the propensity of bilinguals to engage in CS.
Outcome: The findings confirm the existence of a statistically significant connection between POS and the likelihood of CS across language pairs, but show that it diminishes as tokens distance themselves from CS instances.
Building a User-Generated Content North-African Arabizi Treebank: Tackling Hell (2020.acl-main)

Copied to clipboard

Challenge: a treebank for a north-African Arabic dialect known for code-switching is made freely available . authors: geopolitical events are a factor highlighting a language deficiency in terms of natural language processing resources .
Approach: They propose to make a treebank for a romanized user-generated content variety of Algerian . they supplement it with 50k unlabeled sentences from common crawl and web-crawled data .
Outcome: The proposed treebank is made of 1500 sentences, fully annotated in morpho-syntax and universal dependency syntax, with full translation at both the word and sentence levels.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations