Challenge: a corpus of Arabic-English code-switching (CS) spontaneous speech is collected in an Egyptian university soundproof room . the language in Egypt is rather complex and poses many challenges to natural language processing (NLP)
Approach: They present an Egyptian Arabic-English code-switching (CS) spontaneous speech corpus.
Outcome: The proposed corpus is designed to be used in automatic speech recognition systems . it provides a useful resource for analyzing the CS phenomenon from linguistic, sociological, and psychological perspectives.

Similar Papers

Collection and Analysis of Code-switch Egyptian Arabic-English Speech Corpus (L18-1)

Copied to clipboard

Challenge: despite of the great demand, there is still a huge shortage in available corpora for dialectal languages and code-switched speech.
Approach: They collect conversational Egyptian Arabic spontaneous speech, extract transcriptions and analyze it from a code-switching perspective.
Outcome: The authors collect conversational Egyptian Arabic spontaneous speech, extract transcriptions and analyze speech from the code-switching perspective.
Cairo Student Code-Switch (CSCS) Corpus: An Annotated Egyptian Arabic-English Corpus (2020.lrec-1)

Copied to clipboard

Challenge: Code-switching is a phenomenon commonly observed in the Arabicspeaking world . there is still a huge gap in the available resources and NLP applications .
Approach: They propose a corpus of Egyptian- Arabic code-switch speech data that is fully tokenized, lemmatized and annotated for part-of-speech tags.
Outcome: The proposed corpus of Egyptian- Arabic code-switch speech data is fully tokenized, lemmatized and annotated for part-of-speech tags.
Arabic Speech Rhythm Corpus: Read and Spontaneous Speaking Styles (2020.lrec-1)

Copied to clipboard

Challenge: a corpus of Arabic speech recordings has been built to allow comparisons between Arabic and other languages.
Approach: They propose to build a corpus of Arabic speech recordings that can be compared with other languages.
Outcome: The proposed corpus can be used for forensic phonetic research and casework applications.
ZAEBUC-Spoken: A Multilingual Multidialectal Arabic-English Speech Corpus (2024.lrec-main)

Copied to clipboard

Challenge: a corpus of multilingual Arabic-English speech is presented in a new paper . a major bottleneck is the lack of data needed for training NLP models .
Approach: They propose a multilingual multidialectal Arabic-English speech corpus with a set of guidelines for automatic speech recognition.
Outcome: The proposed corpus includes two languages with Arabic and English spoken in multiple variants and Arabic and Arabic with various accents.
The WAW Corpus: The First Corpus of Interpreted Speeches and their Translations for English and Arabic (L18-1)

Copied to clipboard

Challenge: Using the corpus, we study the characteristics of interpreters' work and train machine translation systems.
Approach: They propose to build an interpreting corpus for Arabic and an Arabic corpus to study interpreters' work.
Outcome: The proposed corpus can be used for teaching interpreters and to train machine translation systems.
A Survey of Code-switched Arabic NLP: Progress, Challenges, and Future Directions (2025.coling-main)

Copied to clipboard

Challenge: Code-switching (CSW) is a common linguistic phenomenon in multilingual societies . current literature on CSW in the arab world is limited to the Arabic language .
Approach: They present a review of the literature in the field of code-switched Arabic NLP . they propose recommendations for future research .
Outcome: This review provides a broad perspective on the current literature in the field of code-switched Arabic NLP . it also provides recommendations for future research .
The French-Algerian Code-Switching Triggered audio corpus (FACST) (L18-1)

Copied to clipboard

Challenge: The French Algerian Code-Switching Triggered corpus is a corpus of spontaneous CS utterances . it is used to support linguistic and phonetic studies in phonetics and prosody .
Approach: They propose to use a triggering protocol to elicit CS in natural conversations . they propose to do data segmentation and annotation in each language .
Outcome: The proposed corpus is based on a code-switching protocol and is well-suited for linguistic and acoustic-phonetic studies.
TArC: Tunisian Arabish Corpus, First complete release (2022.lrec-1)

Copied to clipboard

Challenge: a project focused on Tunisian Arabic encoded in Arabizi is a hybrid approach to linguistics and linguistic research . Arabic dialects are notoriously under-resourced linguistic systems .
Approach: They propose to use Arabic script as a linguistic corpus and a neural network architecture to annotate the latter with various levels of linguistic information.
Outcome: The proposed approach is hybrid and combines linguistic and linguistic tools . the proposed approach produces in cascade different levels of annotation .
DECM: Evaluating Bilingual ASR Performance on a Code-switching/mixing Benchmark (2024.lrec-main)

Copied to clipboard

Challenge: Code-switched (CSW) speech is a linguistic phenomenon that occurs when spoken utterances switch languages between sentences.
Approach: They propose to use a dataset to evaluate German-English CSW speech . they show that the dataset includes splits with varying degrees of CSW .
Outcome: The proposed dataset includes spontaneous speech from diverse domains, enabling realistic CSW evaluation in German-English.
Processing and Understanding Mixed Language Data (D19-2)

Copied to clipboard

Challenge: Multilingual communities exhibit code-mixing, mixing of two or more languages in a single conversation . social media and other informal interactive platforms are allowing code-switching in user-generated text .
Approach: a tutorial aims to provide a foundation for researchers to study code-mixing in multilingual communities.
Outcome: a tutorial aims to provide new researchers with a foundation in linguistics and computational aspects of code-mixing.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations