| Challenge: | Sentiment lexica are vital for sentiment analysis in absence of document-level annotations . linguistic resources are limited for at least a few hundred languages, putting them at risk of extinction . |
| Approach: | They introduce UniSent universal sentiment lexica for 1000+ languages . they use a Bible corpus to project sentiment information from English to other languages based on Twitter data . |
| Outcome: | The proposed method mitigates domain mismatch between Bible and Twitter by using embeddings . it compares to other sentiment seeding methods in a subset of languages with ground truth available . |
Similar Papers
A Multilingual BPE Embedding Space for Universal Sentiment Lexicon Induction (P19-1)
Copied to clipboard
| Challenge: | Existing methods for sentiment lexicon induction are limited to low-resource languages. |
| Approach: | They propose a method for sentiment lexicon induction that is applicable to the entire range of typological diversity of the world's languages. |
| Outcome: | The proposed method is applicable to the entire range of typological diversity of the world's languages. |
Learning and Evaluating Emotion Lexicons for 91 Languages (2020.acl-main)
Copied to clipboard
| Challenge: | Emotion lexicons describe the affective meaning of words but are limited in coverage for most languages. |
| Approach: | They propose a method for creating arbitrarily large emotion lexicons for any target language. |
| Outcome: | The proposed method exceeds human reliability for some languages and variables. |
NaijaSenti: A Nigerian Twitter Sentiment Corpus for Multilingual Sentiment Analysis (2022.lrec-1)
Copied to clipboard
Shamsuddeen Hassan Muhammad, David Ifeoluwa Adelani, Sebastian Ruder, Ibrahim Sa’id Ahmad, Idris Abdulmumin, Bello Shehu Bello, Monojit Choudhury, Chris Chinenye Emezue, Saheed Salahudeen Abdullahi, Anuoluwapo Aremu, Alípio Jorge, Pavel Brazdil
| Challenge: | Sentiment analysis is one of the most widely studied applications in NLP, but most work focuses on languages with large amounts of data. |
| Approach: | They propose a large-scale human-annotated Twitter sentiment dataset for the four most widely spoken languages in Nigeria. |
| Outcome: | The proposed dataset includes 30,000 tweets and a significant fraction of code-mixed tweets. |
Building Sentiment Lexicons for Mainland Scandinavian Languages Using Machine Translation and Sentence Embeddings (2022.lrec-1)
Copied to clipboard
| Challenge: | a simple but effective method to build sentiment lexicons for the three Mainland Scandinavian languages is proposed . a number of experiments with Scandinavian language datasets yield state-of-the-art results using a rule-based sentiment analysis algorithm. |
| Approach: | They propose a simple but effective method to build sentiment lexicons for the three Mainland Scandinavian languages. |
| Outcome: | The proposed method is based on the English Sentiwordnet and a thesaurus in one of the target languages. |
How Universal are Universal Dependencies? Exploiting Syntax for Multilingual Clause-level Sentiment Detection (2020.lrec-1)
Copied to clipboard
| Challenge: | a new method for clause-level sentiment detection is proposed for multilingual use cases. |
| Approach: | They propose a pipeline method that makes the most of syntactic structures based on Universal Dependencies. |
| Outcome: | The proposed method achieves high precision in sentiment detection for 17 languages . it avoids machine-learning approaches that may cause obstacles to its use cases . |
Odi et Amo. Creating, Evaluating and Extending Sentiment Lexicons for Latin. (2020.lrec-1)
Copied to clipboard
| Challenge: | a new paper aims to provide sentiment analysis tools for ancient languages . the current sentiment analysis resources only cover modern languages based on textual typologies . |
| Approach: | They propose to use manually-curated Latin lexicons to evaluate sentiment analysis tools . they propose a gold standard and a silver standard for evaluating lexical items . |
| Outcome: | The proposed lexicons are evaluated using a gold standard and a silver standard for sentiment analysis. |
AfriSenti: A Twitter Sentiment Analysis Benchmark for African Languages (2023.emnlp-main)
Copied to clipboard
Shamsuddeen Muhammad, Idris Abdulmumin, Abinew Ayele, Nedjma Ousidhoum, David Adelani, Seid Yimam, Ibrahim Ahmad, Meriem Beloucif, Saif Mohammad, Sebastian Ruder, Oumaima Hourrane, Alipio Jorge, Pavel Brazdil, Felermino Ali, Davis David, Salomey Osei, Bello Shehu-Bello, Falalu Lawan, Tajuddeen Gwadabe, Samuel Rutunda, Tadesse Belay, Wendimu Messelle, Hailu Balcha, Sisay Chala, Hagos Gebremichael, Bernard Opoku, Stephen Arthur
| Challenge: | Africa has the highest linguistic diversity among all continents. |
| Approach: | They introduce a sentiment analysis benchmark that contains >110,000 tweets in 14 African languages . they describe the data collection methodology, annotation process, and challenges . |
| Outcome: | The proposed dataset contains >110,000 tweets in 14 African languages . the tweets were annotated by native speakers and used in the shared task . |
GATITOS: Using a New Multilingual Lexicon for Low-resource Machine Translation (2023.emnlp-main)
Copied to clipboard
| Challenge: | a new study explores the effectiveness of bilingual lexica in machine translation models . cross-lingual vocabulary alignment is still highly imperfect in these models, despite the success of supervised and self-supervised training. |
| Approach: | They use a resource to improve translation performance on 200-language models . they show that lexica is more reliable than human-translated data . |
| Outcome: | The proposed approach improves on 200-language translation models with lexical data augmentation . the proposed approach is open-source and has 168 tail languages . |
A Comparative Analysis of Unsupervised Language Adaptation Methods (D19-61)
Copied to clipboard
| Challenge: | Recent proposed approaches to perform unsupervised language adaptation lack annotated resources in less-resourced languages. |
| Approach: | They propose to use Adversarial Training, Sentence Encoder Alignment and Shared-Private Architecture to perform unsupervised language adaptation without using aligned sentences. |
| Outcome: | The proposed approaches are more suitable when the source and target language datasets contain other variations in content besides the language shift. |
Twitter corpus of Resource-Scarce Languages for Sentiment Analysis and Multilingual Emoji Prediction (C18-1)
Copied to clipboard
| Challenge: | a majority of research studies on twitter focus on English tweets, despite the fact that English dominates the mix of languages. |
| Approach: | They leverage social media platforms such as twitter for developing corpus across multiple languages . they use tweets to collect data for sentiment analysis and emoji prediction . |
| Outcome: | The proposed method is applicable for resource-scarce languages provided speakers of that particular language are active users on social media platforms. |