Challenge: Existing studies on sentiment analysis in low-resource languages have focused on major languages and emotionally laden text genres like social media and reviews.
Approach: They propose to use GPT-4 for sentiment analysis on Faroese news texts using a multi-class approach with 225 sentences analysed in 170 articles.
Outcome: The proposed model performs remarkably well on 225 sentences and 170 articles compared to human annotators .

Similar Papers

Evaluating the Potential of Language-family-specific Generative Models for Low-resource Data Augmentation: A Faroese Case Study (2024.lrec-main)

Copied to clipboard

Challenge: generative language models have shown promising results for translation in zero, one, and fewshot learning settings, among other types of tasks.
Approach: They propose to prompt a generative language model for the Nordic languages for Faroese to English translation in a zero, one, and few-shot setting and challenge its Farose language understanding capabilities on a small dataset.
Outcome: The proposed model can translate Faroese to English in a zero, one, and few-shot setting and then use it to create an annotated Farose semantic textual similarity (STS) dataset.
A Fine-grained Sentiment Dataset for Norwegian (2020.lrec-1)

Copied to clipboard

Challenge: Using a dataset for fine-grained sentiment analysis in Norwegian, we examine the annotation effort and provide an overview of the developed annotation guidelines.
Approach: They propose a dataset for fine-grained sentiment analysis in Norwegian . they provide an overview of the developed annotation guidelines and analyze inter-annotator agreement .
Outcome: The proposed dataset is the first of its kind for Norwegian and is available online.
GoodNewsEveryone: A Corpus of News Headlines Annotated with Emotions, Semantic Roles, and Reader Perception (2020.lrec-1)

Copied to clipboard

Challenge: Fewer studies address emotions as a phenomenon to be tackled with structured learning, which can be explained by the lack of relevant datasets.
Approach: They propose to annotate 5000 English news headlines with their associated emotions, the corresponding emotion experiencers and textual cues, related emotion causes and targets, and the reader’s perception of the emotion of the headline.
Outcome: The proposed method enables further research on emotion classification, emotion intensity prediction, emotion cause detection and supports qualitative studies.
Is GPT-4 a Good Data Analyst? (2023.findings-emnlp)

Copied to clipboard

Challenge: Large language models (LLMs) have shown their powerful capabilities in plenty of domains and tasks, including context understanding, code generation, language generation, data storytelling, etc.
Approach: They propose to use GPT-4 as a data analyst to perform end-to-end data analysis with databases from a wide range of domains.
Outcome: The proposed framework compares GPT-4 with human data analysts to perform end-to-end data analysis with databases from a wide range of domains.
DynaSent: A Dynamic Benchmark for Sentiment Analysis (2021.acl-long)

Copied to clipboard

Challenge: Sentiment analysis is an early success story for NLP, in both a technical and an industrial sense.
Approach: They propose to combine naturally occurring sentences with sentences created using the open-source Dynabench Platform, which facilities human-and-model-in-the-loop dataset creation.
Outcome: The proposed model is more coherent than comparable models and motivates training models from scratch over successive fine-tuning.
Information Extraction from Legal Wills: How Well Does GPT-4 Do? (2023.findings-emnlp)

Copied to clipboard

Challenge: Using information extraction from legal wills is an important application of artificial intelligence (AI)
Approach: They propose a manually annotated dataset for Information Extraction (IE) from legal wills . they also use it to evaluate the performance of large language models (LLMs)
Outcome: The proposed dataset can be used to evaluate large language models on IE from legal wills . it shows that the model performs reasonably well, but inconsistent outputs and overgeneralization are observed .
MarathiEmoExplain: A Dataset for Sentiment, Emotion, and Explanation in Low-Resource Marathi (2025.findings-emnlp)

Copied to clipboard

Challenge: Marathi is the third most widely spoken language in India with over 83 million native speakers . available Marath datasets are limited to coarse sentiment labels and lack fine-grained emotional categorization or interpretability through explanations.
Approach: They propose to annotate Marathi sentences labeled with sentiment, emotion and a corresponding natural language justification.
Outcome: The proposed dataset provides a benchmark for future research in multilingual and explainable NLP.
Sentiment Analysis in the Era of Large Language Models: A Reality Check (2024.findings-naacl)

Copied to clipboard

Challenge: Sentiment analysis (SA) has been a long-standing research area in natural language processing.
Approach: They propose a benchmark to evaluate LLMs' SA abilities and propose 'sentiEval' benchmark to be used for a more comprehensive evaluation.
Outcome: The proposed benchmark outperforms small language models on 26 datasets on 13 tasks and compared them with LLMs trained on domain-specific datasets.
Multi-source Multi-domain Sentiment Analysis with BERT-based Models (2022.lrec-1)

Copied to clipboard

Challenge: Sentiment analysis is a widely studied task in natural language processing.
Approach: They propose to improve BERT-based models for sentiment analysis on italian corpora and evaluate their performance on the basis of eight corpors.
Outcome: The proposed model is evaluated over eight sentiment analysis corpora from different domains and sources on the prediction of positive, negative and neutral classes.
MAD-TSC: A Multilingual Aligned News Dataset for Target-dependent Sentiment Classification (2023.acl-long)

Copied to clipboard

Challenge: Sentiment classification is a task that requires domain-specific datasets.
Approach: They propose a new dataset which includes aligned examples in eight languages . they show that machine translations can replace manual ones and that results match English .
Outcome: The proposed dataset compares the performance of the proposed model with existing datasets in eight languages and human and machine translations.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations