| Challenge: | Arabizi is a written form of spoken Arabic, relying on Latin characters and digits. |
| Approach: | They propose to use Arabizi as a written form of spoken Arabic in online social networks . they use a corpus of 7.7M tweets written in Arabizi and a subset of SALAD to train a model in Arabic . |
| Outcome: | The proposed model outperforms state-of-the-art models on sentiment analysis task using arabizi . the proposed model is based on a corpus of 7.7M tweets written in arabizi and a subset of LAD manually annotated for sentiment analysis. |
Similar Papers
NaijaSenti: A Nigerian Twitter Sentiment Corpus for Multilingual Sentiment Analysis (2022.lrec-1)
Copied to clipboard
Shamsuddeen Hassan Muhammad, David Ifeoluwa Adelani, Sebastian Ruder, Ibrahim Sa’id Ahmad, Idris Abdulmumin, Bello Shehu Bello, Monojit Choudhury, Chris Chinenye Emezue, Saheed Salahudeen Abdullahi, Anuoluwapo Aremu, Alípio Jorge, Pavel Brazdil
| Challenge: | Sentiment analysis is one of the most widely studied applications in NLP, but most work focuses on languages with large amounts of data. |
| Approach: | They propose a large-scale human-annotated Twitter sentiment dataset for the four most widely spoken languages in Nigeria. |
| Outcome: | The proposed dataset includes 30,000 tweets and a significant fraction of code-mixed tweets. |
Multi-source Multi-domain Sentiment Analysis with BERT-based Models (2022.lrec-1)
Copied to clipboard
| Challenge: | Sentiment analysis is a widely studied task in natural language processing. |
| Approach: | They propose to improve BERT-based models for sentiment analysis on italian corpora and evaluate their performance on the basis of eight corpors. |
| Outcome: | The proposed model is evaluated over eight sentiment analysis corpora from different domains and sources on the prediction of positive, negative and neutral classes. |
An Algerian Corpus and an Annotation Platform for Opinion and Emotion Analysis (2020.lrec-1)
Copied to clipboard
| Challenge: | Currently, there are more than 4 billion Internet users worldwide . the number of social media users in Algeria has tripled over a year . |
| Approach: | They propose a platform for crowdsourcing annotation of tweets at different levels of granularity. |
| Outcome: | The proposed platform can be used to create the largest Algerian dialect subjectivity lexicon of about 9,000 entries. |
A Dataset and BERT-based Models for Targeted Sentiment Analysis on Turkish Texts (2022.acl-srw)
Copied to clipboard
| Challenge: | Sentiment analysis is a field that is growing due to the availability of the Internet and the growing number of online platforms. |
| Approach: | They propose an annotated Turkish dataset suitable for targeted sentiment analysis. |
| Outcome: | The proposed models outperform the traditional models for the targeted sentiment analysis task. |
XLM-T: Multilingual Language Models in Twitter for Sentiment Analysis and Beyond (2022.lrec-1)
Copied to clipboard
| Challenge: | Language models are ubiquitous in NLP, but current analyses focus on (multilingual variants of) standard benchmarks and task-specific corpora as multilingual signals. |
| Approach: | They propose a model to train and evaluate multilingual language models in Twitter using a set of Twitter datasets in eight different languages and a XLM-T model. |
| Outcome: | The proposed model trains and evaluates multilingual models on Twitter. |
Swan and ArabicMTEB: Dialect-Aware, Arabic-Centric, Cross-Lingual, and Cross-Cultural Embedding Models and Benchmarks (2025.findings-naacl)
Copied to clipboard
Gagan Bhatia, El Moatez Billah Nagoudi, Abdellah El Mekki, Fakhraddin Alwajih, Muhammad Abdul-Mageed
| Challenge: | In this paper, we introduce a family of embedding models addressing both small-scale and large-scale use cases. |
| Approach: | They propose to use ArabicMTEB to evaluate Arabic text embedding models . they propose to build a benchmark suite that assesses cross-lingual, multi-dialectal, multidomain, and multi-cultural Arabic text embedded models. |
| Outcome: | The proposed models outperform Multilingual-E5-large and Swan-Large in most Arabic tasks while remaining dialectally and culturally aware. |
Exploring Multilingual Pre-trained Language Model for Aspect-based Sentiment Analysis (2026.findings-acl)
Copied to clipboard
| Challenge: | Aspect-based sentiment analysis studies have focused on English datasets, but labeled data is scarce. |
| Approach: | They propose a multilingual pre-trained language model that leverages bilingual pre-training to leverage aspects-based sentiment analysis. |
| Outcome: | The proposed model outperforms state-of-the-art models across multiple languages. |
Identifying Sentiments in Algerian Code-switched User-generated Comments (2020.lrec-1)
Copied to clipboard
| Challenge: | a recent study has focused on sentiment analysis for the Arabic variety, but it has been extended to other domains. |
| Approach: | They build a corpus of 36,000 code-switched user-generated comments annotated for sentiments in Algerian Arabic. |
| Outcome: | The proposed model performs better on unedited code-switched and unbalanced data across sentiment classes. |
MultiBooked: A Corpus of Basque and Catalan Hotel Reviews Annotated for Aspect-level Sentiment Classification (L18-1)
Copied to clipboard
| Challenge: | sentiment analysis research has focused on unsupervised or semi-supervised approaches, but these still require a large number of resources and do not reach the performance of supervised approaches. |
| Approach: | They propose two datasets for supervised aspect-level sentiment analysis in Basque and Catalan. |
| Outcome: | The proposed datasets are based on two under-resourced languages, basque and catalan. |
AfriSenti: A Twitter Sentiment Analysis Benchmark for African Languages (2023.emnlp-main)
Copied to clipboard
Shamsuddeen Muhammad, Idris Abdulmumin, Abinew Ayele, Nedjma Ousidhoum, David Adelani, Seid Yimam, Ibrahim Ahmad, Meriem Beloucif, Saif Mohammad, Sebastian Ruder, Oumaima Hourrane, Alipio Jorge, Pavel Brazdil, Felermino Ali, Davis David, Salomey Osei, Bello Shehu-Bello, Falalu Lawan, Tajuddeen Gwadabe, Samuel Rutunda, Tadesse Belay, Wendimu Messelle, Hailu Balcha, Sisay Chala, Hagos Gebremichael, Bernard Opoku, Stephen Arthur
| Challenge: | Africa has the highest linguistic diversity among all continents. |
| Approach: | They introduce a sentiment analysis benchmark that contains >110,000 tweets in 14 African languages . they describe the data collection methodology, annotation process, and challenges . |
| Outcome: | The proposed dataset contains >110,000 tweets in 14 African languages . the tweets were annotated by native speakers and used in the shared task . |