Nunc profana tractemus. Detecting Code-Switching in a Large Corpus of 16th Century Letters (2022.lrec-1)
Copied to clipboard
Martin Volk, Lukas Fischer, Patricia Scheurer, Bernard Silvan Schroffenegger, Raphael Schwitter, Phillip Ströbel, Benjamin Suter
| Challenge: | a corpus of 16th century letters from and to the Zurich reformer Heinrich Bullinger has been preserved . a recent study investigated code-switching in these 8600 letters . |
| Approach: | They investigate the automatic detection of code-switching in a 16th century letter exchange . they use a popular language identifier to bootstrap a word-based language classifier . |
| Outcome: | The proposed language classifier bootstraps with a popular identifier on a small training corpus of 150 sentences per language. |
Similar Papers
Detecting de minimis Code-Switching in Historical German Books (2020.coling-main)
Copied to clipboard
| Challenge: | Code-switching has drawn scholarly attention in computational linguistics and natural language processing from many different perspectives. |
| Approach: | They propose to compare informal code-switching to its appearance in more formal registers by annotating and inspecting the German textarchives. |
| Outcome: | The proposed classifiers can help reduce errors when speech recognition is applied to a large corpus with rare embedded languages. |
The Decades Progress on Code-Switching Research in NLP: A Systematic Survey on Trends and Challenges (2023.findings-acl)
Copied to clipboard
| Challenge: | Code-Switching is a common phenomenon in written text and conversation . it is not so common to observe code-switching in spoken language and not in written language . |
| Approach: | They present a systematic survey on code-switching research in natural language processing to understand the progress of the past decades and conceptualize the challenges and tasks on the topic. |
| Outcome: | The proposed model combines linguistic theories and machine learning techniques to understand the code-switching phenomenon. |
Code-Switched Language Identification is Harder Than You Think (2024.eacl-long)
Copied to clipboard
| Challenge: | Code switching (CS) is a common phenomenon in written and spoken communication, but is handled poorly by many NLP applications. |
| Approach: | They propose to use CS language identification for corpus building to make it more realistic by scaling it to more languages and considering models with simpler architectures for faster inference. |
| Outcome: | The proposed system is based on a sentence-level multi-label tagging problem and provides recommendations for future work. |
The DReaM Corpus: A Multilingual Annotated Corpus of Grammars for the World’s Languages (2020.lrec-1)
Copied to clipboard
| Challenge: | Until recently, language descriptions were available in paper form only, with indexes as the only search aid. |
| Approach: | They propose to digitize a multilingual corpus of language descriptions and annotate it with various meta, word, and text attributes to make searching and analysis easier and more useful. |
| Outcome: | The proposed corpus is searchable through a couple of well-established corpus infrastructures. |
Automatic Identification of Code-Switching Functions in Speech Transcripts (2023.findings-acl)
Copied to clipboard
| Challenge: | Code-switching, or switching between languages, occurs for many reasons and has important linguistic, sociological, and cultural implications. |
| Approach: | They build a system to identify a wide range of functions for which speakers code-switch in everyday speech with an accuracy of 75% . they use a dataset of Hindi-English code-witched data to analyze their results . |
| Outcome: | The proposed system can identify a wide range of functions for which speakers code-switch in everyday speech, with an accuracy of 75% across all functions. |
Code-Mixed Probes Show How Pre-Trained Models Generalise on Code-Switched Text (2024.lrec-main)
Copied to clipboard
| Challenge: | Code-switching is a prevalent linguistic phenomenon in which multilingual individuals seamlessly alternate between languages. |
| Approach: | They propose to use pre-trained language models to generalise to code-switched text . they use a dataset of well-formed naturalistic code-witched texts and parallel translations into the source languages to examine their results. |
| Outcome: | The proposed model generalises to code-switched text, shedding light on their ability to generalise representations to CS corpora. |
Collection and Analysis of Code-switch Egyptian Arabic-English Speech Corpus (L18-1)
Copied to clipboard
| Challenge: | despite of the great demand, there is still a huge shortage in available corpora for dialectal languages and code-switched speech. |
| Approach: | They collect conversational Egyptian Arabic spontaneous speech, extract transcriptions and analyze it from a code-switching perspective. |
| Outcome: | The authors collect conversational Egyptian Arabic spontaneous speech, extract transcriptions and analyze speech from the code-switching perspective. |
TweetTaglish: A Dataset for Investigating Tagalog-English Code-Switching (2022.lrec-1)
Copied to clipboard
| Challenge: | a large dataset is available to study Tagalog-English code-switching in low-resource settings. |
| Approach: | They propose to use a large dataset to investigate Tagalog-English code-switching . they use linguistic data from Tagalogue and Tagalit-English to investigate their results . |
| Outcome: | The proposed dataset achieves a strong performance benchmark for Tagalog-English code-switching. |
Cairo Student Code-Switch (CSCS) Corpus: An Annotated Egyptian Arabic-English Corpus (2020.lrec-1)
Copied to clipboard
| Challenge: | Code-switching is a phenomenon commonly observed in the Arabicspeaking world . there is still a huge gap in the available resources and NLP applications . |
| Approach: | They propose a corpus of Egyptian- Arabic code-switch speech data that is fully tokenized, lemmatized and annotated for part-of-speech tags. |
| Outcome: | The proposed corpus of Egyptian- Arabic code-switch speech data is fully tokenized, lemmatized and annotated for part-of-speech tags. |
Introducing a Parsed Corpus of Historical High German (2024.lrec-main)
Copied to clipboard
| Challenge: | outlines the development of the Indiana Parsed Corpus of (Historical) High German . outlines selection of texts, decisions on part-of-speech tags and other labels . |
| Approach: | They propose to build a parsed German corpus that spans Germanic from 1050 to 1950 . they propose to use Penn-style treebanks to capture syntactic relationships between words . |
| Outcome: | The proposed corpus spans Germanic languages from 1050 to 1950 and illustrative annotation issues unique to the language. |