Challenge: a recent study has focused on sentiment analysis for the Arabic variety, but it has been extended to other domains.
Approach: They build a corpus of 36,000 code-switched user-generated comments annotated for sentiments in Algerian Arabic.
Outcome: The proposed model performs better on unedited code-switched and unbalanced data across sentiment classes.

Similar Papers

NaijaSenti: A Nigerian Twitter Sentiment Corpus for Multilingual Sentiment Analysis (2022.lrec-1)

Copied to clipboard

Challenge: Sentiment analysis is one of the most widely studied applications in NLP, but most work focuses on languages with large amounts of data.
Approach: They propose a large-scale human-annotated Twitter sentiment dataset for the four most widely spoken languages in Nigeria.
Outcome: The proposed dataset includes 30,000 tweets and a significant fraction of code-mixed tweets.
Normalising Non-standardised Orthography in Algerian Code-switched User-generated Data (D19-55)

Copied to clipboard

Challenge: a new corpus of unstructured data from social media is presenting challenges to NLP research . standardisation is neither natural nor universal, it is rather a human invention.
Approach: They compile a parallel corpus of Arabic textual data matched with human annotations . they use a deep neural model designed to deal with context-dependent spelling correction .
Outcome: The proposed model performs best with two CNN sub-network encoders and an LSTM decoder . pre-processing data token-by-token with edit-distance aligner significantly improves performance .
An Algerian Corpus and an Annotation Platform for Opinion and Emotion Analysis (2020.lrec-1)

Copied to clipboard

Challenge: Currently, there are more than 4 billion Internet users worldwide . the number of social media users in Algeria has tripled over a year .
Approach: They propose a platform for crowdsourcing annotation of tweets at different levels of granularity.
Outcome: The proposed platform can be used to create the largest Algerian dialect subjectivity lexicon of about 9,000 entries.
Building a User-Generated Content North-African Arabizi Treebank: Tackling Hell (2020.acl-main)

Copied to clipboard

Challenge: a treebank for a north-African Arabic dialect known for code-switching is made freely available . authors: geopolitical events are a factor highlighting a language deficiency in terms of natural language processing resources .
Approach: They propose to make a treebank for a romanized user-generated content variety of Algerian . they supplement it with 50k unlabeled sentences from common crawl and web-crawled data .
Outcome: The proposed treebank is made of 1500 sentences, fully annotated in morpho-syntax and universal dependency syntax, with full translation at both the word and sentence levels.
NERDz: A Preliminary Dataset of Named Entities for Algerian (2022.aacl-short)

Copied to clipboard

Challenge: NER is a fundamental task in information extraction and natural language processing.
Approach: They propose to build a manually annotated Algerian vernacular dataset using a recent extension to the Algerian NArabizi Treebank.
Outcome: The proposed dataset is the first of its kind for the Algerian vernacular dialect.
Wetin dey with these comments? Modeling Sociolinguistic Factors Affecting Code-switching Behavior in Nigerian Online Discussions (P19-1)

Copied to clipboard

Challenge: Multilingual individuals code switch between languages as part of a complex communication process.
Approach: They propose to model the social and contextual factors eliciting code switching in a rich contextual environment by analyzing 330K articles and 389K comments labeled for code switching behavior.
Outcome: The proposed model shows that topic-driven variation, tribal affiliation, emotional valence, and audience design all play complementary roles in behavior.
SentiArabic: A Sentiment Analyzer for Standard Arabic (L18-1)

Copied to clipboard

Challenge: Sentiment analysis is a process of applying computational approaches to identify attitudes, emotions and opinions in text, speech and visual data.
Approach: They propose a sentiment analyzer that identifies the overall contextual polarity for Arabic text.
Outcome: The proposed system achieves an F-score of 76.5% when evaluated on a blind test set.
Evaluating Emotion Arcs Across Languages: Bridging the Global Divide in Sentiment Analysis (2023.findings-emnlp)

Copied to clipboard

Challenge: Emotion arcs capture how an individual (or a population) feels over time.
Approach: They compare machine-learning and Lexicon-Only methods to generate emotion arcs . they run experiments on 18 diverse datasets in 9 languages .
Outcome: The proposed method is poor at instance level emotion classification, but highly accurate when aggregating information from hundreds of instances.
An Automatic Learning of an Algerian Dialect Lexicon by using Multilingual Word Embeddings (L18-1)

Copied to clipboard

Challenge: a study on the Algerian Arabic dialect aims to build a lexicon of words written in Arabic or Latin script . multilinguality of the corpus is due to the fact that people use several languages to post comments . stretched letters, misspelled words, emoticons, condensed writing are among the problems .
Approach: They propose to build automatically from a social network an Algerian dialect lexicon.
Outcome: The proposed method leads to a score of 73% on a test lexicon . the study is based on analyzing a lexical corpus of an Algerian dialect .
A Survey of Code-switched Arabic NLP: Progress, Challenges, and Future Directions (2025.coling-main)

Copied to clipboard

Challenge: Code-switching (CSW) is a common linguistic phenomenon in multilingual societies . current literature on CSW in the arab world is limited to the Arabic language .
Approach: They present a review of the literature in the field of code-switched Arabic NLP . they propose recommendations for future research .
Outcome: This review provides a broad perspective on the current literature in the field of code-switched Arabic NLP . it also provides recommendations for future research .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations