Evaluating the Capabilities of Large Language Models for Multi-label Emotion Understanding (2025.coling-main)
Copied to clipboard
Tadesse Destaw Belay, Israel Abebe Azime, Abinew Ali Ayele, Grigori Sidorov, Dietrich Klakow, Philip Slusallek, Olga Kolesnikova, Seid Muhie Yimam
| Challenge: | Emotion classification is one of the most challenging tasks in large language models. |
| Approach: | They propose to use a multi-label emotion classification dataset for four Ethiopian languages to evaluate their ability to learn and reason. |
| Outcome: | The proposed model improves the understanding of emotions in language models and how people convey emotions through various languages. |
Similar Papers
EthioLLM: Multilingual Large Language Models for Ethiopian Languages with Task Evaluation (2024.lrec-main)
Copied to clipboard
Atnafu Lambebo Tonja, Israel Abebe Azime, Tadesse Destaw Belay, Mesay Gemeda Yigezu, Moges Ahmed Ah Mehamed, Abinew Ali Ayele, Ebrahim Chekol Jibril, Michael Melese Woldeyohannis, Olga Kolesnikova, Philipp Slusallek, Dietrich Klakow, Seid Muhie Yimam
| Challenge: | Low-resource languages are lagging behind current state-of-the-art (SOTA) developments in the field of NLP due to insufficient resources to train LLMs. |
| Approach: | They propose to use multilingual large language models for five Ethiopian languages and a benchmark dataset to evaluate their performance. |
| Outcome: | The proposed models outperform existing models in five Ethiopian languages and a benchmark dataset for various downstream NLP tasks. |
Are Large Language Model-based Evaluators the Solution to Scaling Up Multilingual Evaluation? (2024.findings-eacl)
Copied to clipboard
Rishav Hada, Varun Gumma, Adrian Wynter, Harshita Diddee, Mohamed Ahmed, Monojit Choudhury, Kalika Bali, Sunayana Sitaram
| Challenge: | Large Language Models (LLMs) excel in various tasks, but their evaluation, especially in languages beyond the top 20, remains inadequate due to existing benchmarks and metrics limitations. |
| Approach: | They propose to use Large Language Models as evaluators to rank or score other models’ outputs by calibrating them against 20K human judgments across three text-generation tasks, five metrics, and eight languages. |
| Outcome: | The proposed evaluation methods can be used to improve multilingual evaluation by calibrating them against 20K human judgments across three text-generation tasks, five metrics, and eight languages. |
SpanEmo: Casting Multi-label Emotion Classification as Span-prediction (2021.eacl-main)
Copied to clipboard
| Challenge: | Current approaches to ER ignore potential ambiguities, in which multiple emotions overlap. |
| Approach: | They propose a model "SpanEmo" which casts multi-label emotion classification as span-prediction and introduces a loss function focused on modelling multiple co-existing emotions in a sentence. |
| Outcome: | The proposed model can predict multiple co-existing emotions in a sentence and improve model performance and learning meaningful associations between labels and words in the sentence. |
Characterizing and Evaluating Working Emotion Vocabularies in Multilingual Large Language Models (2026.acl-long)
Copied to clipboard
| Challenge: | Prior work evaluating emotion and affective understanding in large language models rely on predetermined label sets or focus on a singular evaluation task. |
| Approach: | They examine the ability of multilingual language models to predict any term used by an author to label their own feelings or emotions. |
| Outcome: | The proposed models perform poorly on three different tasks in English and Spanish. |
EmoBench: Evaluating the Emotional Intelligence of Large Language Models (2024.acl-long)
Copied to clipboard
Sahand Sabour, Siyang Liu, Zheyuan Zhang, June Liu, Jinfeng Zhou, Alvionna Sunaryo, Tatia Lee, Rada Mihalcea, Minlie Huang
| Challenge: | Existing benchmarks for Emotional Intelligence (EI) focus on emotion recognition, neglecting essential EI capabilities. |
| Approach: | They propose a benchmark that proposes a comprehensive definition for machine EI . they propose 400 hand-crafted questions in English and Chinese to evaluate EI. |
| Outcome: | The proposed benchmarks focus on emotion recognition, neglecting EI capabilities . they are constructed from existing datasets, which include frequent patterns and errors . the proposed benchmark includes questions in English and Chinese that require thorough reasoning and understanding . |
Large Language Models Do Multi-Label Classification Differently (2025.emnlp-main)
Copied to clipboard
| Challenge: | Multi-label classification is prevalent in real-world settings, but the behavior of Large Language Models (LLMs) in this setting is understudied. |
| Approach: | They propose to use initial probability distributions to analyze output distributions of LLMs at each label generation step to find out how LLM models perform multi-label classification. |
| Outcome: | The proposed methods improve alignment and predictive performance over existing methods. |
Cross-lingual Emotion Detection (2022.lrec-1)
Copied to clipboard
| Challenge: | Emotion detection is a useful tool for understanding human behavior, but constructing annotated datasets to train models can be expensive. |
| Approach: | They propose to use English as the source language with Arabic and Spanish as target languages to train models for emotion detection in a target language. |
| Outcome: | The proposed approaches surpass state-of-the-art models in Arabic and Spanish by 4% and 5% respectively. |
BRIGHTER: BRIdging the Gap in Human-Annotated Textual Emotion Recognition Datasets for 28 Languages (2025.acl-long)
Copied to clipboard
Shamsuddeen Hassan Muhammad, Nedjma Ousidhoum, Idris Abdulmumin, Jan Philip Wahle, Terry Ruas, Meriem Beloucif, Christine de Kock, Nirmal Surange, Daniela Teodorescu, Ibrahim Said Ahmad, David Ifeoluwa Adelani, Alham Fikri Aji, Felermino D. M. A. Ali, Ilseyar Alimova, Vladimir Araujo, Nikolay Babakov, Naomi Baes, Ana-Maria Bucur, Andiswa Bukula, Guanqun Cao, Rodrigo Tufiño, Rendi Chevi, Chiamaka Ijeoma Chukwuneke, Alexandra Ciobotaru, Daryna Dementieva, Murja Sani Gadanya, Robert Geislinger, Bela Gipp, Oumaima Hourrane, Oana Ignat, Falalu Ibrahim Lawan, Rooweither Mabuya, Rahmad Mahendra, Vukosi Marivate, Alexander Panchenko, Andrew Piper, Charles Henrique Porto Ferreira, Vitaly Protasov, Samuel Rutunda, Manish Shrivastava, Aura Cristina Udrea, Lilian Diana Awuor Wanzare, Sophie Wu, Florian Valentin Wunderlich, Hanif Muhammad Zhafran, Tianhui Zhang, Yi Zhou, Saif M. Mohammad
| Challenge: | Emotion recognition is an umbrella term for several NLP tasks, but most work on high-resource languages has focused on low-resourced languages. |
| Approach: | They propose to use emotion recognition to describe perceived emotions in 28 different languages and across several domains to identify and annotate the datasets. |
| Outcome: | The proposed datasets cover low-resource languages from Africa, Asia, Eastern Europe, and Latin America, with instances labeled by fluent speakers. |
GlotEval: A Test Suite for Massively Multilingual Evaluation of Large Language Models (2025.emnlp-demos)
Copied to clipboard
Hengyu Luo, Zihao Li, Joseph Attieh, Sawal Devkota, Ona de Gibert, Xu Huang, Shaoxiong Ji, Peiqin Lin, Bhavani Sai Praneeth Varma Mantina, Ananda Sreenidhi, Raúl Vázquez, Mengjie Wang, Samea Yusofi, Fei Yuan, Jörg Tiedemann
| Challenge: | Existing evaluation frameworks focus on English and a handful of high-resource languages, thereby overlooking the realistic performance of large language models in multilingual and lower-resourced scenarios. |
| Approach: | They propose a unified and lightweight framework that integrates 27 benchmarks under a standard ISO 639-3 language identifier system to enable seamless incorporation of new benchmarks. |
| Outcome: | The proposed framework integrates 27 benchmarks under a standard ISO 639-3 language identifier system, allowing for seamless incorporation of new benchmarks. |
7 Points to Tsinghua but 10 Points to ? Assessing Large Language Models in Agentic Multilingual National Bias (2025.findings-acl)
Copied to clipboard
| Challenge: | Large Language Models have garnered significant attention for their capabilities in multilingual natural language processing, but studies on risks associated with cross biases are limited to immediate context preferences. |
| Approach: | They investigate multilingual bias in state-of-the-art Large Language Models by analyzing their responses to decision-making tasks across multiple languages. |
| Outcome: | The proposed model can provide personalized advice across university applications, travel, and relocation scenarios. |