Challenge: Recent advances in the education domain have provided new opportunities for solving interesting, but difficult problems.
Approach: They propose to use EduSenti to fine-tune language models for assigning sentiment to reviews of educators' performance annotated for sentiment, emotion and educational topic.
Outcome: The proposed model is compared with an Albanian masked language trained model from the last XLM-RoBERTa checkpoint and shows that it is a good fit for the proposed model.

Similar Papers

The ParlaSent Multilingual Training Dataset for Sentiment Identification in Parliamentary Proceedings (2024.lrec-main)

Copied to clipboard

Challenge: The paper presents a new training dataset of sentences in 7 languages, manually annotated for sentiment, which is used in a series of experiments focused on training a robust sentiment identifier for parliamentary proceedings.
Approach: They propose to use a dataset of sentences manually annotated for sentiment to train a robust sentiment identifier for parliamentary proceedings.
Outcome: The proposed model performs very well on languages not seen during fine-tuning and additional fine- tuning data from other languages significantly improves the target parliament’s results.
LoRETTA: Low-Rank Economic Tensor-Train Adaptation for Ultra-Low-Parameter Fine-Tuning of Large Language Models (2024.naacl-long)

Copied to clipboard

Challenge: Existing methods for parameter-efficient fine-tuning are limited by the growing number of trainable parameters with the rapid deployment of Large Language Models (LLMs).
Approach: They propose a parameter-efficient framework that reduces trainable parameters through tensor-train decomposition.
Outcome: The proposed methods achieve comparable or better performance than most widely used methods with up to 100 fewer parameters on the LLaMA-2-7B models.
XED: A Multilingual Dataset for Sentiment Analysis and Emotion Detection (2020.coling-main)

Copied to clipboard

Challenge: XED is a multilingual fine-grained emotion dataset for English and other low-resource languages.
Approach: They propose a multilingual fine-grained emotion dataset using Plutchik's Wheel of Emotions and a projection scheme to annotate Finnish and English sentences.
Outcome: The proposed dataset is based on human-annotated Finnish and English sentences and projected annotations for 30 additional languages.
EduBench: A Comprehensive Benchmarking Dataset for Evaluating Large Language Models in Diverse Educational Scenarios (2026.acl-long)

Copied to clipboard

Challenge: Existing benchmarks that focus on knowledge-intensive tasks do not reflect diverse educational scenarios.
Approach: They propose a benchmark that incorporates 9 major scenarios and 4,000 educational contexts.
Outcome: The proposed model performs comparable to state-of-the-art large models on the test set.
Multi-source Multi-domain Sentiment Analysis with BERT-based Models (2022.lrec-1)

Copied to clipboard

Challenge: Sentiment analysis is a widely studied task in natural language processing.
Approach: They propose to improve BERT-based models for sentiment analysis on italian corpora and evaluate their performance on the basis of eight corpors.
Outcome: The proposed model is evaluated over eight sentiment analysis corpora from different domains and sources on the prediction of positive, negative and neutral classes.
A Warm Start and a Clean Crawled Corpus - A Recipe for Good Language Models (2022.lrec-1)

Copied to clipboard

Challenge: Pre-trained neural language models have shown impressive results when adapted for a variety of classification and text generation tasks.
Approach: They propose to use Icelandic's Icelandic Common Crawl Corpus to train language models that achieve state-of-the-art performance in downstream tasks.
Outcome: The proposed models achieve state-of-the-art in a variety of downstream tasks including part-of speech tagging, named entity recognition and constituency parsing.
MultiBooked: A Corpus of Basque and Catalan Hotel Reviews Annotated for Aspect-level Sentiment Classification (L18-1)

Copied to clipboard

Challenge: sentiment analysis research has focused on unsupervised or semi-supervised approaches, but these still require a large number of resources and do not reach the performance of supervised approaches.
Approach: They propose two datasets for supervised aspect-level sentiment analysis in Basque and Catalan.
Outcome: The proposed datasets are based on two under-resourced languages, basque and catalan.
EENLP: Cross-lingual Eastern European NLP Index (2022.lrec-1)

Copied to clipboard

Challenge: Existing NLP resources for Eastern European languages are sparse.
Approach: They propose to use existing Eastern European language resources to build cross-lingual datasets for five different semantic tasks to support commonsense reasoning.
Outcome: The proposed model trains on 104 languages and shows impressive results on text analysis tasks.
Does syntax matter? A strong baseline for Aspect-based Sentiment Analysis with RoBERTa (2021.naacl-main)

Copied to clipboard

Challenge: Aspect-based sentiment analysis (ABSA) is a fine-grained task in sentiment analysis.
Approach: They compare a model with a dependency parser and a tree from a fine-tuned RoBERTa model to find the polarities for aspects in a sentence.
Outcome: The proposed model outperforms the parser-provided tree on six datasets across four languages.
MALAMUTE: A Multilingual, Highly-granular, Template-free, Education-based Probing Dataset (2025.findings-acl)

Copied to clipboard

Challenge: Existing cloze-style benchmarks for language models lack specific, granular areas of knowledge and often rely on templates that can bias models.
Approach: They propose a multilingual, template-free, and highly granular probing dataset comprising expert-written, peer-reviewed probes from 71 university-level textbooks across three languages.
Outcome: The proposed dataset covers eight domains, each with up to 14 subdomains, further broken down into concepts and concept-based prompts.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations