Challenge: Sentiment lexica are vital for sentiment analysis in absence of document-level annotations . linguistic resources are limited for at least a few hundred languages, putting them at risk of extinction .
Approach: They introduce UniSent universal sentiment lexica for 1000+ languages . they use a Bible corpus to project sentiment information from English to other languages based on Twitter data .
Outcome: The proposed method mitigates domain mismatch between Bible and Twitter by using embeddings . it compares to other sentiment seeding methods in a subset of languages with ground truth available .

Similar Papers

A Multilingual BPE Embedding Space for Universal Sentiment Lexicon Induction (P19-1)

Copied to clipboard

Challenge: Existing methods for sentiment lexicon induction are limited to low-resource languages.
Approach: They propose a method for sentiment lexicon induction that is applicable to the entire range of typological diversity of the world's languages.
Outcome: The proposed method is applicable to the entire range of typological diversity of the world's languages.
Learning and Evaluating Emotion Lexicons for 91 Languages (2020.acl-main)

Copied to clipboard

Challenge: Emotion lexicons describe the affective meaning of words but are limited in coverage for most languages.
Approach: They propose a method for creating arbitrarily large emotion lexicons for any target language.
Outcome: The proposed method exceeds human reliability for some languages and variables.
NaijaSenti: A Nigerian Twitter Sentiment Corpus for Multilingual Sentiment Analysis (2022.lrec-1)

Copied to clipboard

Challenge: Sentiment analysis is one of the most widely studied applications in NLP, but most work focuses on languages with large amounts of data.
Approach: They propose a large-scale human-annotated Twitter sentiment dataset for the four most widely spoken languages in Nigeria.
Outcome: The proposed dataset includes 30,000 tweets and a significant fraction of code-mixed tweets.
Building Sentiment Lexicons for Mainland Scandinavian Languages Using Machine Translation and Sentence Embeddings (2022.lrec-1)

Copied to clipboard

Challenge: a simple but effective method to build sentiment lexicons for the three Mainland Scandinavian languages is proposed . a number of experiments with Scandinavian language datasets yield state-of-the-art results using a rule-based sentiment analysis algorithm.
Approach: They propose a simple but effective method to build sentiment lexicons for the three Mainland Scandinavian languages.
Outcome: The proposed method is based on the English Sentiwordnet and a thesaurus in one of the target languages.
How Universal are Universal Dependencies? Exploiting Syntax for Multilingual Clause-level Sentiment Detection (2020.lrec-1)

Copied to clipboard

Challenge: a new method for clause-level sentiment detection is proposed for multilingual use cases.
Approach: They propose a pipeline method that makes the most of syntactic structures based on Universal Dependencies.
Outcome: The proposed method achieves high precision in sentiment detection for 17 languages . it avoids machine-learning approaches that may cause obstacles to its use cases .
Odi et Amo. Creating, Evaluating and Extending Sentiment Lexicons for Latin. (2020.lrec-1)

Copied to clipboard

Challenge: a new paper aims to provide sentiment analysis tools for ancient languages . the current sentiment analysis resources only cover modern languages based on textual typologies .
Approach: They propose to use manually-curated Latin lexicons to evaluate sentiment analysis tools . they propose a gold standard and a silver standard for evaluating lexical items .
Outcome: The proposed lexicons are evaluated using a gold standard and a silver standard for sentiment analysis.
AfriSenti: A Twitter Sentiment Analysis Benchmark for African Languages (2023.emnlp-main)

Copied to clipboard

Challenge: Africa has the highest linguistic diversity among all continents.
Approach: They introduce a sentiment analysis benchmark that contains >110,000 tweets in 14 African languages . they describe the data collection methodology, annotation process, and challenges .
Outcome: The proposed dataset contains >110,000 tweets in 14 African languages . the tweets were annotated by native speakers and used in the shared task .
GATITOS: Using a New Multilingual Lexicon for Low-resource Machine Translation (2023.emnlp-main)

Copied to clipboard

Challenge: a new study explores the effectiveness of bilingual lexica in machine translation models . cross-lingual vocabulary alignment is still highly imperfect in these models, despite the success of supervised and self-supervised training.
Approach: They use a resource to improve translation performance on 200-language models . they show that lexica is more reliable than human-translated data .
Outcome: The proposed approach improves on 200-language translation models with lexical data augmentation . the proposed approach is open-source and has 168 tail languages .
A Comparative Analysis of Unsupervised Language Adaptation Methods (D19-61)

Copied to clipboard

Challenge: Recent proposed approaches to perform unsupervised language adaptation lack annotated resources in less-resourced languages.
Approach: They propose to use Adversarial Training, Sentence Encoder Alignment and Shared-Private Architecture to perform unsupervised language adaptation without using aligned sentences.
Outcome: The proposed approaches are more suitable when the source and target language datasets contain other variations in content besides the language shift.
Twitter corpus of Resource-Scarce Languages for Sentiment Analysis and Multilingual Emoji Prediction (C18-1)

Copied to clipboard

Challenge: a majority of research studies on twitter focus on English tweets, despite the fact that English dominates the mix of languages.
Approach: They leverage social media platforms such as twitter for developing corpus across multiple languages . they use tweets to collect data for sentiment analysis and emoji prediction .
Outcome: The proposed method is applicable for resource-scarce languages provided speakers of that particular language are active users on social media platforms.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations