Challenge: Recent studies have focused on cross-lingual transfer between languages with similar typology and languages of different scripts.
Approach: They propose to annotate Algerian user-generated comments with parallel annotations . they also investigate the effect of script vs. language similarity in cross-lingual transfer .
Outcome: The proposed model fine-tunes multi-lingual models on Algerian language and scripts . it shows that script vs. language similarity is important for part-of-speech tagging and sentiment analysis .

Similar Papers

Unknown Script: Impact of Script on Cross-Lingual Transfer (2024.naacl-srw)

Copied to clipboard

Challenge: Existing models for high-resource languages are not available for all languages, and the vast majority of the world's languages are excluded from these models.
Approach: They propose to use pre-trained models to analyze the effect of the target language and its script on cross-lingual transfer.
Outcome: The proposed model is based on six models pre-trained on NER and POS tasks in the original script and romanized version.
Analyzing the Effect of Linguistic Similarity on Cross-Lingual Transfer: Tasks and Experimental Setups Matter (2025.findings-acl)

Copied to clipboard

Challenge: Prior work on cross-lingual transfer often focuses on a small set of languages from a few language families and/or a single task.
Approach: They analyze cross-lingual transfer for 263 languages from a wide variety of language families . they include three popular NLP tasks: POS tagging, dependency parsing, topic classification .
Outcome: The proposed approach is based on linguistic similarity measures for 263 languages . the results show that the effect of linguistic similarities on transfer performance depends on a range of factors .
Cross-Cultural Similarity Features for Cross-Lingual Transfer Learning of Pragmatically Motivated Tasks (2021.eacl-main)

Copied to clipboard

Challenge: a large amount of work on cross-lingual transfer learning focused on typological and genealogical similarities between languages.
Approach: They propose three features that capture cross-cultural similarities that manifest in linguistic patterns and quantify distinct aspects of language pragmatics.
Outcome: The proposed features capture cross-cultural similarities manifest in linguistic patterns and quantify aspects of language pragmatics.
An Automatic Learning of an Algerian Dialect Lexicon by using Multilingual Word Embeddings (L18-1)

Copied to clipboard

Challenge: a study on the Algerian Arabic dialect aims to build a lexicon of words written in Arabic or Latin script . multilinguality of the corpus is due to the fact that people use several languages to post comments . stretched letters, misspelled words, emoticons, condensed writing are among the problems .
Approach: They propose to build automatically from a social network an Algerian dialect lexicon.
Outcome: The proposed method leads to a score of 73% on a test lexicon . the study is based on analyzing a lexical corpus of an Algerian dialect .
Identifying Sentiments in Algerian Code-switched User-generated Comments (2020.lrec-1)

Copied to clipboard

Challenge: a recent study has focused on sentiment analysis for the Arabic variety, but it has been extended to other domains.
Approach: They build a corpus of 36,000 code-switched user-generated comments annotated for sentiments in Algerian Arabic.
Outcome: The proposed model performs better on unedited code-switched and unbalanced data across sentiment classes.
TransliCo: A Contrastive Learning Framework to Address the Script Barrier in Multilingual Pretrained Language Models (2024.acl-long)

Copied to clipboard

Challenge: The world’s more than 7000 languages are written in at least 293 scripts, which poses a difficulty for multilingual pretrained language models in learning crosslingual knowledge through lexical overlap.
Approach: They propose a framework that optimizes the Transliteration Contrastive Modeling objective to fine-tune an mPLM by contrasting sentences in its training data and transliterations in a unified script.
Outcome: The proposed model outperforms Glot500-m on zero-shot crosslingual transfer tasks while retaining uniformity across scripts.
Through the Looking Glass of Multilingual AI: Contrasting Language- and Name Script-Dependent Ethnic Hierarchies in GPT and DeepSeek (2026.acl-srw)

Copied to clipboard

Challenge: a recent study found that large language models are biased overwhelmingly Anglocentric . a stereotype perceptual map is a framework for analyzing how ethnic groups are positioned along evaluative dimensions.
Approach: They use a stereotype perceptual map to examine how ethnic groups are positioned along evaluative dimensions.
Outcome: The stereotype perceptual map analyzes model behavior across languages, scripts, evaluative domains and models.
Exploring Anisotropy and Outliers in Multilingual Language Models for Cross-Lingual Semantic Sentence Similarity (2023.findings-acl)

Copied to clipboard

Challenge: Recent studies have shown that contextual language models display outlier dimensions . this is true for monolingual and multilingual models, but little work has been done on multilingual contexts .
Approach: They investigate outlier dimensions and their relationship to anisotropy in multilingual contexts . they focus on cross-lingual semantic similarity tasks .
Outcome: The proposed model improves on cross-lingual semantic similarity tasks.
Cross-lingual Editing in Multilingual Language Models (2024.findings-eacl)

Copied to clipboard

Challenge: Existing models editing techniques (METs) can efficiently update outdated LLMs without retraining.
Approach: They propose a cross-lingual model editing paradigm where a fact is edited in one language and the subsequent update propagation is observed across other languages.
Outcome: The proposed techniques perform well in multilingual models with knowledge stored in multiple languages.
Happiness is Sharing a Vocabulary: A Study of Transliteration Methods (2026.eacl-long)

Copied to clipboard

Challenge: a key problem in multilingual NLP is script barrier, which makes it difficult to share knowledge between languages . a new study shows that transliteration can be useful for languages using non-Latin scripts .
Approach: They propose to use romanization, phonemic transcription, and substitution ciphers to evaluate models . romanization outperforms other input types in 7 out of 8 evaluation settings .
Outcome: The proposed approach outperforms other input types on three tasks and is the most effective . romanization outperformed other input type in 7 out of 8 evaluation settings .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations