Papers with CS
Copied to clipboard
| Challenge: | Multilingual speakers outnumber monolingual speakers in the world . CS is a frequent habit in both spoken and written informal communications . |
| Approach: | They evaluate the efficacy of cross-lingual transfer learning with mBERT for NLU on a Basque-Spanish CS chatbot corpus. |
| Outcome: | The proposed model outperforms models trained on Basque and Spanish without CS on a basque-Spanish chatbot corpus. |
Copied to clipboard
| Challenge: | Existing studies fail to provide comprehensive service satisfaction analysis . Existing models fail to include satisfaction polarity classification and sentimental utterance identification . |
| Approach: | They propose a model that predicts customer sentiments and aggregates them into service satisfaction polarity. |
| Outcome: | The proposed model predicts customer sentiments and aggregates them into service satisfaction polarity and reasoning clues. |
Copied to clipboard
| Challenge: | Recent advances in automatic speech recognition (ASR) have pushed error rates below 5% on standard monolingual benchmarks. |
| Approach: | They propose a framework for the evaluation of multilingual ASR models using loanword labels and a hierarchical CS-level labeling scheme that allows for fine-tuning with synthetic CS data. |
| Outcome: | The proposed framework provides a means for the precise evaluation of multilingual ASR models and fosters research in the field. |
Copied to clipboard
| Challenge: | Code switching (CS) is a common phenomenon in written and spoken communication, but is handled poorly by many NLP applications. |
| Approach: | They propose to use CS language identification for corpus building to make it more realistic by scaling it to more languages and considering models with simpler architectures for faster inference. |
| Outcome: | The proposed system is based on a sentence-level multi-label tagging problem and provides recommendations for future work. |
Copied to clipboard
| Challenge: | Sentence-level LIDs are classifiers trained on monolingual texts to provide single labels, typically using a softmax layer to turn scores into probabilities. |
| Approach: | They propose a simple yet effective code-switching language identification method that uses the LID itself to mask features associated with L1 and L2 in the next round. |
| Outcome: | The proposed method is based on two open-source LIDs based in the FastText architecture and does not require any external resources. |
Copied to clipboard
| Challenge: | Using culture-agnostic subsets, performance drops in many LMMs when evaluated in Japanese. |
| Approach: | They introduce a Japanese benchmark to evaluate large multimodal models on expert-level tasks based on the Japanese cultural context. |
| Outcome: | The proposed benchmark enables comparisons with other benchmarks in other languages based on cultural contexts. |
Copied to clipboard
| Challenge: | Large language models (LLMs) are increasingly embedded in Computer Science classrooms to automate code generation, feedback, and assessment. |
| Approach: | They propose a guardrail framework for educational AI systems that can handle unsafe and irrelevant prompts. |
| Outcome: | The proposed framework reduces potentially harmful or policy-violating code completions by 30-65% without degrading performance on legitimate educational tasks. |
Copied to clipboard
| Challenge: | Recent advances in language modeling have caused disruptive shifts throughout AI research, spurring discussion about how the field is changing and how it should change. |
| Approach: | They analyze a dataset of 16,979 LLM-related arXiv papers and examine industry and academic publishing trends. |
| Outcome: | The authors examine the impact of large language models on AI research in 2023 and 2022. |
Copied to clipboard
| Challenge: | a crowdsourcing effort to annotate and link parallel texts has been unsuccessful . a data set of parallel texts in eleven languages is presented . |
| Approach: | They present a wikified data set of English sentences linked to Wikipedia pages . they use crowdsourcing to annotate the texts and perform crowdsourcing for complex annotations . |
| Outcome: | The proposed data set is valuable as it constitutes a rich resource . it includes annotated data of English sentences linked to translations in eleven languages . |
Copied to clipboard
| Challenge: | Code-switching (CS) is the alternation of languages within an utterance or conversation. |
| Approach: | They propose to use translation-and-align and augment with a generation model followed by match-and filter to improve CS generalizability of cross-lingual models when data for only one language is available. |
| Outcome: | The proposed models improve when only English data is available alongside zero or a few CS training instances. |
Copied to clipboard
| Challenge: | Code summarization (CS) is a promising area in recent language understanding . previous work using structurebased traversal or non-sequential models to learn structural program semantics has shown no performance gain . |
| Approach: | They propose to use a structure-based traversal model to learn structural program semantics to generate human language automatically for programming language in the format of source code. |
| Outcome: | Experiments show that the proposed method achieves state-of-the-art on benchmarks. |
Copied to clipboard
| Challenge: | Large-scale introductory CS courses struggle to provide personalized support and encourage active participation. |
| Approach: | They propose to use predictive query management to generate student questions and answers ahead of lectures and to engage in interactive conversations with a tutoring model. |
| Outcome: | The proposed learning assistant generates student questions and answers ahead of lectures and interacts with students via the same interface. |
Copied to clipboard
| Challenge: | Customer-oriented behaviour (COB) is often hindered by a lack of clarity in its definition and lack of robust analytical, categorization, and computational approaches. |
| Approach: | They propose a conceptual and empirical framework for customer-oriented behaviour in call centre interactions . they aim to identify facets of COB that positively impact on Customer Satisfaction . |
| Outcome: | The proposed framework improves our understanding of the dynamics shaping sales strategies in call centres and holds promise for practical applications in optimising customer-agent interactions. |
Copied to clipboard
| Challenge: | Code switching (CS) is the phenomenon of interchangeably using words and phrases from different languages. |
| Approach: | They propose a new ST corpus that extends the joint transcription and translation setup. |
| Outcome: | The proposed model performs well even when no training data is used. |
Copied to clipboard
| Challenge: | a novel MT pipeline that considers the intra-data relation is proposed . previous MT systems have demonstrated relatively low performance, making them hardly utilized as another data source. |
| Approach: | They propose a new MT pipeline that considers the intra-data relation . they propose CS and IT to enhance the intra data relation based on a data point . |
| Outcome: | The proposed pipeline improves translation quality and training data compared with the existing approach . it yields better training data and better translation quality than previous approaches . |
Copied to clipboard
| Challenge: | Code-switching (CS) is a common linguistic phenomenon wherein speakers fluidly transition between languages in conversation. |
| Approach: | They propose to use a part-of-speech (POS)-based analysis of Spanish-English and Mandarin-English corpora to examine the propensity of bilinguals to engage in CS. |
| Outcome: | The findings confirm the existence of a statistically significant connection between POS and the likelihood of CS across language pairs, but show that it diminishes as tokens distance themselves from CS instances. |
Copied to clipboard
| Challenge: | Using CS/CM as a linguistic phenomenon could be a sign of tension in Holocaust survivors’ interviews. |
| Approach: | They annotated CS/CM codes and annotate silence situations in an open corpus . they found that most annotations were captured in the tension places . |
| Outcome: | The proposed method shows that annotations are captured in the tension places . the study calls for more research endeavors on tension detection . |
Copied to clipboard
| Challenge: | Cued Speech (CS) is a visual communication system developed for people with hearing loss to complement speech reading at the phonetic level. |
| Approach: | They propose a method to phonemize written corpora so that each word is aligned with the corresponding CS key(s) this method is part of a wider project aimed at creating an augmented reality system displaying a virtual coding hand where the user will be able to choose a text upon its complexity for cueing. |
| Outcome: | The proposed method is part of a wider project aimed at creating an augmented reality system displaying a virtual coding hand where the user can choose a text upon its complexity for cueing. |
Copied to clipboard
| Challenge: | a major barrier to research on CS has been the lack of large multilingual, multi-genre CS-annotated corpora. |
| Approach: | They propose a web-based annotation system that manages large-scale CS data annotation. |
| Outcome: | The proposed system can manage large-scale multilingual code switching (CS) data annotation. |
Copied to clipboard
| Challenge: | Social media data can be used to improve natural language processing performance, but it is often overlooked by lexical normalization systems. |
| Approach: | They propose three lexical normalization models specifically designed to handle code-switched data and evaluate their performance on POS tags. |
| Outcome: | The proposed models outperform monolingual models and lead to 5.4% performance increase for POS tagging compared to unnormalized input. |
Copied to clipboard
| Challenge: | Code-switching (CS) is a phenomenon of alternating between two or more languages in conversations . if at least one language is morphologically rich, a large number of words can be composed of morphemes from more than one language. |
| Approach: | They propose to extend the language identification task to the subword level by splitting mixed words while tagging each part with a language ID. |
| Outcome: | The proposed model outperforms the baseline on a Spanish–Wixarika and adapted German–Turkish datasets. |
Copied to clipboard
| Challenge: | Recent trends in NLP research have raised an interest in linguistic code-switching . however, many of these approaches are limited to a few language pairs and a specific domain . |
| Approach: | They propose a centralized benchmark for Linguistic Code-switching Evaluation that combines eleven corpora covering four different code-switch languages and four tasks. |
| Outcome: | The proposed benchmark provides a centralized benchmark and compares with other benchmarks in real-time. |
Copied to clipboard
| Challenge: | The French Algerian Code-Switching Triggered corpus is a corpus of spontaneous CS utterances . it is used to support linguistic and phonetic studies in phonetics and prosody . |
| Approach: | They propose to use a triggering protocol to elicit CS in natural conversations . they propose to do data segmentation and annotation in each language . |
| Outcome: | The proposed corpus is based on a code-switching protocol and is well-suited for linguistic and acoustic-phonetic studies. |
Copied to clipboard
| Challenge: | a computational model for code-switching text is lacking in the corpus of real text. |
| Approach: | They propose a neural machine translation model to generate Hindi-English code-switched sentences using monolingual Hindi sentences. |
| Outcome: | The proposed model reduces perplexity on a language modeling task and improves on linguistic inference tasks. |
Copied to clipboard
| Challenge: | Existing methods for generating evidence-supported counterspeech lack clear guidance with a core claim for organizing evidence. |
| Approach: | They propose a Factuality and Faithfulness Reinforcement Learning framework for generating claim-guided and evidence-supported counterspeech (F2RL) they generate counter-claims based on hate speech and design a self-evaluation mechanism to select the most appropriate one. |
| Outcome: | The proposed framework achieves excellent performance on three benchmark datasets with strong factuality and faithfulness. |
Copied to clipboard
| Challenge: | Code-switching (CS) is a problem in machine translation, but its performance is not investigated for CS settings. |
| Approach: | They propose to use morphological segmentation techniques for machine translation tasks . they compare morphology-based and frequency-based segmentation for MT tasks based on data size . |
| Outcome: | The proposed approach performs best in MT tasks but under-performs in other languages. |
Copied to clipboard
| Challenge: | Linguistic studies on code-switching focus on the "how" and "why" of CS . a new model aims to derive CS functions from local and global properties of the code-witched discourse . |
| Approach: | They propose a model that integrates CS phenomena and modalities into a representation that includes local and global properties of the code-switched discourse. |
| Outcome: | The proposed model simplifies the analysis of English/Hindi CS datasets and provides a flexible framework for further studies. |
Copied to clipboard
| Challenge: | Existing studies on large language models have shown that they are poorly aligned in practice. |
| Approach: | They propose a framework to evaluate safety in large language models . they propose two new metrics to quantify fake alignment and obtain corrected performance estimation. |
| Outcome: | The proposed framework and two metrics show that some models with purported safety are poorly aligned in practice. |
Copied to clipboard
| Challenge: | Current clinical LLM benchmarks fail to evaluate advanced clinical skills in AI and large language models (LLMs). |
| Approach: | They propose a framework to evaluate large language models (LLMs) using two instruction-following tasks designed to reflect real clinical scenarios. |
| Outcome: | The proposed framework evaluates LLMs through two instruction-following tasks designed to reflect real clinical scenarios. |
Copied to clipboard
| Challenge: | Code-switching (CS) is the process of speakers switching between two or more languages in spoken or written language. |
| Approach: | They propose to use the Matrix Language Frame theory to describe CS speech . they compare MLID of English/Mandarin and English/Spanish CS to acoustic language identity . |
| Outcome: | The proposed models outperform monolingual models in acoustic language identity recognition tasks. |
Copied to clipboard
| Challenge: | Existing terminology constraint test sets are blind to this issue due to oversimplified settings . PH methods retain high constraint accuracy but lower translation quality . |
| Approach: | They propose a method that replaces terminology terms with ordered labels . placeholder methods are better at retaining high constraint accuracy but lower translation quality . |
| Outcome: | The proposed method achieves high accuracy and translation quality regardless of the number or length of constraints. |
Copied to clipboard
| Challenge: | despite advances in English-Thai MT, common MT approaches often underperform in the medical field due to their inability to precisely translate medical terminologies. |
| Approach: | They propose to maintain medical terminology in English within translated text through code-switched translation. |
| Outcome: | The proposed method shows that medical professionals prefer CS translations that maintain critical English terms accurately, even if it slightly compromises fluency. |
Copied to clipboard
| Challenge: | Code-switching (CS) is a phenomenon of switching between multiple languages . current models cannot handle CS due to lack of annotated data and limited resources. |
| Approach: | They propose a self-training method to repurpose existing models using a switch-point bias by leveraging unannotated data to reduce the gap between the switch point performance and retain overall performance on two distinct language pairs. |
| Outcome: | The proposed model reduces the gap between the switch point performance while retaining the overall performance on two distinct language pairs. |
Copied to clipboard
| Challenge: | Existing methods for CS use dictionaries or parallel sentences with word-alignment to generate CS data by randomly switching words in a sentence. |
| Approach: | They propose a method that focuses on Entity-level Code-Switching to capture fine-grained cross-lingual semantics without corrupting syntax. |
| Outcome: | The proposed method captures fine-grained cross-lingual semantics without corrupting syntax. |
Copied to clipboard
| Challenge: | a corpus of Arabic-English code-switching (CS) spontaneous speech is collected in an Egyptian university soundproof room . the language in Egypt is rather complex and poses many challenges to natural language processing (NLP) |
| Approach: | They present an Egyptian Arabic-English code-switching (CS) spontaneous speech corpus. |
| Outcome: | The proposed corpus is designed to be used in automatic speech recognition systems . it provides a useful resource for analyzing the CS phenomenon from linguistic, sociological, and psychological perspectives. |
Copied to clipboard
| Challenge: | Existing cognitive stimulation systems lack data on how to integrate emotional support and therapy principles into chit-chat dialogue systems. |
| Approach: | They propose a multi-source knowledge fusion method for CS dialogue to generate open-ended responses guided by the therapy principle and emotional support strategy. |
| Outcome: | The proposed method generates open-ended responses guided by the therapy principle and emotional support strategy of the target response. |
Copied to clipboard
| Challenge: | Existing theories of code-switching (CS) have been refuted in subsequent investigations. |
| Approach: | They propose to use syntactic information to predict where bilinguals switch languages . they find that syntax alone is sufficient for an automatic system to distinguish between sentences in minimal pairs of CS, to the same degree as bilingual humans. |
| Outcome: | The proposed model can explain why bilinguals switch languages more often than in others, but there is no large-scale, multi-language, cross-phenomena experiment that tests this claim. |
Copied to clipboard
| Challenge: | Code-switching (CS) data is ubiquitous in today’s globalized world, but the dearth of annotated datasets in code-switch tasks poses a significant challenge for transfer learning in limited-resource setups. |
| Approach: | They propose a prompt composition technique that outperforms prompt-tuning and fine-tuned prompt-based prompt composition techniques for CS tasks that combine language and task knowledge. |
| Outcome: | The proposed approach outperforms prompt-tuning and fine-tuned approaches on 10 datasets across 4 languages and achieves competitive results in low-resource cross-lingual and cross-task setting. |
Copied to clipboard
| Challenge: | a retrieval shortcut in conversational search (CS) relies on partial history to retrieve relevant passages . naively trained dense retrievers heavily exploit the shortcut and perform poorly when asked to answer history-independent questions. |
| Approach: | They propose to exploit a retrieval shortcut in conversational search (CS) that allows models to only use partial history to retrieve relevant passages while disregarding the latest question. |
| Outcome: | The proposed model outperforms the previous state-of-the-art model by 11.0 on QReCC. |
Copied to clipboard
| Challenge: | Code-switching (CS) is a common phenomenon in real-world conversations and poses significant challenges for multilingual speech technology. |
| Approach: | They propose a pipeline for generating high-quality, natural CS samples without altering sentence semantics. |
| Outcome: | The proposed pipeline generates high-quality, natural CS samples without altering sentence semantics without alteration of sentence semantic. |
Copied to clipboard
| Challenge: | Linguistic code-switching (CS) is an understudied area in natural language processing . lack of resources and annotated data makes it difficult to strive for progress in CS-related tasks. |
| Approach: | They propose a method to adapt monolingual models to code-switched text in various tasks . they transfer English knowledge from a pre-trained ELMo model to different code-paired languages . |
| Outcome: | The proposed method outperforms multilingual BERT and homologous CS-unaware models and provides state-of-the-art in CS tasks. |
Copied to clipboard
| Challenge: | Existing approaches to decode text to the most probable sequence have been proposed to address these challenges by improving coherence, diversity, and resemblance to human-generated text. |
| Approach: | They propose a novel decoding strategy that extends contrastive search by incorporating an adaptive degeneration penalty informed by the model’s estimated uncertainty at each generation step. |
| Outcome: | The proposed approach improves creativity and coherence while maintaining coherency across model architectures, languages, and datasets. |
Copied to clipboard
| Challenge: | Existing methods do not have wide language coverage, fail to account for the diverse range of CS phenomena, or do not scale. |
| Approach: | They propose to use minimal pairs of CS to estimate the extent to which large language models (LLMs) use code-switching in the same way as bilinguals. |
| Outcome: | The proposed model assigns higher probability to the naturally occurring CS sentence than to the variant for every language pair. |
Copied to clipboard
| Challenge: | Using FreqRank, we localize malicious components in outputs for triggered inputs and their corresponding backdoor triggers. |
| Approach: | They propose a mutation-based defense to localize malicious components in LLM outputs and their corresponding backdoor triggers. |
| Outcome: | The proposed defense has an average attack success rate (ASR) of 86.6% and can localize the backdoor triggers in 98% of cases. |
Copied to clipboard
| Challenge: | Existing measures of code-switching (CS) complexity are word-based, meaning any word is equally likely to switch between any two words. |
| Approach: | They adapt two NLP metrics, multilinguality and CS probability, and put forward Intonation Units (IUs) as basic tokens for transcribed bilingual speech. |
| Outcome: | The proposed measures account for prosodic and prosodic constraints on CS in bilingual speech. |
Copied to clipboard
| Challenge: | Existing datasets for scientific NLI are derived from various computer science domains, whereas non-CS domains are completely ignored. |
| Approach: | They propose a scientific natural language inference benchmark called MisMatched that incorporates sentence pairs having an implicit scientific NLI relation into model training. |
| Outcome: | The proposed benchmark covers three non-CS domains and contains 2,700 human annotated sentence pairs. |
Copied to clipboard
| Challenge: | Existing research on generative AI security is driven by mutually reinforcing attack and defense methodologies grounded in empirical experience. |
| Approach: | They propose a new algorithm that uses a random sampling algorithm to control risk. |
| Outcome: | The proposed algorithm improves robustness and utility while maintaining latency comparable to existing algorithms. |
Copied to clipboard
| Challenge: | Recent large language models (LLMs) demonstrate multilingual abilities, yet they are English-centric due to dominance of English in training corpora. |
| Approach: | They propose to use a synthetic English-korean CS question-answering dataset to investigate this potential. |
| Outcome: | The proposed model can activate, identify and leverage knowledge for reasoning in low-resource languages. |