An Automatic Learning of an Algerian Dialect Lexicon by using Multilingual Word Embeddings (L18-1)
Copied to clipboard
| Challenge: | a study on the Algerian Arabic dialect aims to build a lexicon of words written in Arabic or Latin script . multilinguality of the corpus is due to the fact that people use several languages to post comments . stretched letters, misspelled words, emoticons, condensed writing are among the problems . |
| Approach: | They propose to build automatically from a social network an Algerian dialect lexicon. |
| Outcome: | The proposed method leads to a score of 73% on a test lexicon . the study is based on analyzing a lexical corpus of an Algerian dialect . |
Similar Papers
Normalising Non-standardised Orthography in Algerian Code-switched User-generated Data (D19-55)
Copied to clipboard
| Challenge: | a new corpus of unstructured data from social media is presenting challenges to NLP research . standardisation is neither natural nor universal, it is rather a human invention. |
| Approach: | They compile a parallel corpus of Arabic textual data matched with human annotations . they use a deep neural model designed to deal with context-dependent spelling correction . |
| Outcome: | The proposed model performs best with two CNN sub-network encoders and an LSTM decoder . pre-processing data token-by-token with edit-distance aligner significantly improves performance . |
Automatic Identification of Maghreb Dialects Using a Dictionary-Based Approach (L18-1)
Copied to clipboard
| Challenge: | Automatic identification of Arabic dialects in texts is difficult, especially for Maghreb languages and when they are written in Arabic or Latin characters (Arabizi). |
| Approach: | They propose a dictionary-based approach to detect Arabic dialects in texts . they focus on transliteration of Arabicizi into Latin script and code-switching . |
| Outcome: | The proposed approach shows that it is possible to detect dialects in Arabic and Latin scripts. |
The interplay between language similarity and script on a novel multi-layer Algerian dialect corpus (2021.findings-acl)
Copied to clipboard
| Challenge: | Recent studies have focused on cross-lingual transfer between languages with similar typology and languages of different scripts. |
| Approach: | They propose to annotate Algerian user-generated comments with parallel annotations . they also investigate the effect of script vs. language similarity in cross-lingual transfer . |
| Outcome: | The proposed model fine-tunes multi-lingual models on Algerian language and scripts . it shows that script vs. language similarity is important for part-of-speech tagging and sentiment analysis . |
An Algerian Corpus and an Annotation Platform for Opinion and Emotion Analysis (2020.lrec-1)
Copied to clipboard
| Challenge: | Currently, there are more than 4 billion Internet users worldwide . the number of social media users in Algeria has tripled over a year . |
| Approach: | They propose a platform for crowdsourcing annotation of tweets at different levels of granularity. |
| Outcome: | The proposed platform can be used to create the largest Algerian dialect subjectivity lexicon of about 9,000 entries. |
Identifying Sentiments in Algerian Code-switched User-generated Comments (2020.lrec-1)
Copied to clipboard
| Challenge: | a recent study has focused on sentiment analysis for the Arabic variety, but it has been extended to other domains. |
| Approach: | They build a corpus of 36,000 code-switched user-generated comments annotated for sentiments in Algerian Arabic. |
| Outcome: | The proposed model performs better on unedited code-switched and unbalanced data across sentiment classes. |
Standardisation of Dialect Comments in Social Networks in View of Sentiment Analysis : Case of Tunisian Dialect (2022.lrec-1)
Copied to clipboard
| Challenge: | Using the internet, the spoken Arabic dialect language becomes informal languages written in social media . this linguistic situation inhibits mutual understanding and makes computational approaches difficult . we present a pipeline to standardize the written texts in social networks by translating them to MSA . |
| Approach: | They propose a pipeline to standardize Arabic written texts by translating them to MSA . they use a bert-based model to select Tunisian Dialect from MSA and other dialects . |
| Outcome: | The proposed pipeline achieves the best score for the standardization of written texts in social networks . the proposed pipeline includes the translated TD and the original text written in MSA . |
Habibi - a multi Dialect multi National Arabic Song Lyrics Corpus (2020.lrec-1)
Copied to clipboard
| Challenge: | Unlike western music, Arabic songs are poorly classified and the majority of the songs available online are classified under Modern Arabic Pop genre or what is now known as Franco-Arabic . |
| Approach: | They introduce Habibi the first Arabic Song Lyrics corpus for singers from 18 different Arabic countries. |
| Outcome: | The proposed corpus contains more than 30,000 Arabic song lyrics in 6 Arabic dialects for singers from 18 different arab countries. |
NERDz: A Preliminary Dataset of Named Entities for Algerian (2022.aacl-short)
Copied to clipboard
| Challenge: | NER is a fundamental task in information extraction and natural language processing. |
| Approach: | They propose to build a manually annotated Algerian vernacular dataset using a recent extension to the Algerian NArabizi Treebank. |
| Outcome: | The proposed dataset is the first of its kind for the Algerian vernacular dialect. |
The MADAR Arabic Dialect Corpus and Lexicon (L18-1)
Copied to clipboard
Houda Bouamor, Nizar Habash, Mohammad Salameh, Wajdi Zaghouani, Owen Rambow, Dana Abdulrahim, Ossama Obeid, Salam Khalifa, Fadhl Eryani, Alexander Erdmann, Kemal Oflazer
| Challenge: | Using a corpus of 25 Arabic city dialects and a lexicon of 1,045 concepts, we study 25 cities in a travel domain . focus on cities opens new avenues for research from dialectology to dialect identification and machine translation. |
| Approach: | They present two Arabic language resources that are part of the Multi Arabic Dialect Applications and Resources project. |
| Outcome: | The proposed resources are the first of their kind in terms of their coverage and fine granularity. |
ALDi: Quantifying the Arabic Level of Dialectness of Text (2023.emnlp-main)
Copied to clipboard
| Challenge: | Existing work on Dialect Identification (DI) on the sentence level has focused on binary tasks, whereas ALDi treats the task as binary. |
| Approach: | They propose a dataset which contains 127,835 sentences manually labeled with their level of dialectness. |
| Outcome: | The proposed model can identify dialectness on a range of other corpora, providing a more nuanced picture than traditional DI systems. |