RoBERTa Low Resource Fine Tuning for Sentiment Analysis in Albanian (2024.lrec-main)
Copied to clipboard
| Challenge: | Recent advances in the education domain have provided new opportunities for solving interesting, but difficult problems. |
| Approach: | They propose to use EduSenti to fine-tune language models for assigning sentiment to reviews of educators' performance annotated for sentiment, emotion and educational topic. |
| Outcome: | The proposed model is compared with an Albanian masked language trained model from the last XLM-RoBERTa checkpoint and shows that it is a good fit for the proposed model. |
Similar Papers
The ParlaSent Multilingual Training Dataset for Sentiment Identification in Parliamentary Proceedings (2024.lrec-main)
Copied to clipboard
| Challenge: | The paper presents a new training dataset of sentences in 7 languages, manually annotated for sentiment, which is used in a series of experiments focused on training a robust sentiment identifier for parliamentary proceedings. |
| Approach: | They propose to use a dataset of sentences manually annotated for sentiment to train a robust sentiment identifier for parliamentary proceedings. |
| Outcome: | The proposed model performs very well on languages not seen during fine-tuning and additional fine- tuning data from other languages significantly improves the target parliament’s results. |
LoRETTA: Low-Rank Economic Tensor-Train Adaptation for Ultra-Low-Parameter Fine-Tuning of Large Language Models (2024.naacl-long)
Copied to clipboard
| Challenge: | Existing methods for parameter-efficient fine-tuning are limited by the growing number of trainable parameters with the rapid deployment of Large Language Models (LLMs). |
| Approach: | They propose a parameter-efficient framework that reduces trainable parameters through tensor-train decomposition. |
| Outcome: | The proposed methods achieve comparable or better performance than most widely used methods with up to 100 fewer parameters on the LLaMA-2-7B models. |
XED: A Multilingual Dataset for Sentiment Analysis and Emotion Detection (2020.coling-main)
Copied to clipboard
| Challenge: | XED is a multilingual fine-grained emotion dataset for English and other low-resource languages. |
| Approach: | They propose a multilingual fine-grained emotion dataset using Plutchik's Wheel of Emotions and a projection scheme to annotate Finnish and English sentences. |
| Outcome: | The proposed dataset is based on human-annotated Finnish and English sentences and projected annotations for 30 additional languages. |
EduBench: A Comprehensive Benchmarking Dataset for Evaluating Large Language Models in Diverse Educational Scenarios (2026.acl-long)
Copied to clipboard
Bin Xu, Yu Bai, Huashan Sun, Yiguan Lin, Siming Liu, Xinyue Liang, Yaolin Li, Zhuangzhi Dong, Jingren Zhang, Yufan Deng, Xinyu Zou, Yang Gao, Heyan Huang
| Challenge: | Existing benchmarks that focus on knowledge-intensive tasks do not reflect diverse educational scenarios. |
| Approach: | They propose a benchmark that incorporates 9 major scenarios and 4,000 educational contexts. |
| Outcome: | The proposed model performs comparable to state-of-the-art large models on the test set. |
Multi-source Multi-domain Sentiment Analysis with BERT-based Models (2022.lrec-1)
Copied to clipboard
| Challenge: | Sentiment analysis is a widely studied task in natural language processing. |
| Approach: | They propose to improve BERT-based models for sentiment analysis on italian corpora and evaluate their performance on the basis of eight corpors. |
| Outcome: | The proposed model is evaluated over eight sentiment analysis corpora from different domains and sources on the prediction of positive, negative and neutral classes. |
A Warm Start and a Clean Crawled Corpus - A Recipe for Good Language Models (2022.lrec-1)
Copied to clipboard
Vésteinn Snæbjarnarson, Haukur Barri Símonarson, Pétur Orri Ragnarsson, Svanhvít Lilja Ingólfsdóttir, Haukur Jónsson, Vilhjalmur Thorsteinsson, Hafsteinn Einarsson
| Challenge: | Pre-trained neural language models have shown impressive results when adapted for a variety of classification and text generation tasks. |
| Approach: | They propose to use Icelandic's Icelandic Common Crawl Corpus to train language models that achieve state-of-the-art performance in downstream tasks. |
| Outcome: | The proposed models achieve state-of-the-art in a variety of downstream tasks including part-of speech tagging, named entity recognition and constituency parsing. |
MultiBooked: A Corpus of Basque and Catalan Hotel Reviews Annotated for Aspect-level Sentiment Classification (L18-1)
Copied to clipboard
| Challenge: | sentiment analysis research has focused on unsupervised or semi-supervised approaches, but these still require a large number of resources and do not reach the performance of supervised approaches. |
| Approach: | They propose two datasets for supervised aspect-level sentiment analysis in Basque and Catalan. |
| Outcome: | The proposed datasets are based on two under-resourced languages, basque and catalan. |
EENLP: Cross-lingual Eastern European NLP Index (2022.lrec-1)
Copied to clipboard
Alexey Tikhonov, Alex Malkhasov, Andrey Manoshin, George-Andrei Dima, Réka Cserháti, Md.Sadek Hossain Asif, Matt Sárdi
| Challenge: | Existing NLP resources for Eastern European languages are sparse. |
| Approach: | They propose to use existing Eastern European language resources to build cross-lingual datasets for five different semantic tasks to support commonsense reasoning. |
| Outcome: | The proposed model trains on 104 languages and shows impressive results on text analysis tasks. |
Does syntax matter? A strong baseline for Aspect-based Sentiment Analysis with RoBERTa (2021.naacl-main)
Copied to clipboard
| Challenge: | Aspect-based sentiment analysis (ABSA) is a fine-grained task in sentiment analysis. |
| Approach: | They compare a model with a dependency parser and a tree from a fine-tuned RoBERTa model to find the polarities for aspects in a sentence. |
| Outcome: | The proposed model outperforms the parser-provided tree on six datasets across four languages. |
MALAMUTE: A Multilingual, Highly-granular, Template-free, Education-based Probing Dataset (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing cloze-style benchmarks for language models lack specific, granular areas of knowledge and often rely on templates that can bias models. |
| Approach: | They propose a multilingual, template-free, and highly granular probing dataset comprising expert-written, peer-reviewed probes from 71 university-level textbooks across three languages. |
| Outcome: | The proposed dataset covers eight domains, each with up to 14 subdomains, further broken down into concepts and concept-based prompts. |