At Your Own PACE: A Causal Framework for Evaluating EQ in LLMs (2026.findings-acl)
Copied to clipboard
| Challenge: | Emotional Quotient (EQ) has emerged as a competency for seamless human-AI integration. |
| Approach: | They propose a framework for a closed-loop EQ evaluation using a PACE taxonomy to define four dimensions of LLM EQ. |
| Outcome: | The proposed framework achieves high alignment of 89.31% with human preferences while maintaining robust consistency of 83.6%. |
Similar Papers
Deep Associations, High Creativity: A Simple yet Effective Metric for Evaluating Large Language Models (2025.emnlp-main)
Copied to clipboard
| Challenge: | Recent studies evaluate the creative capabilities of large language models (LLMs) through diverse tasks, aiming to understand their strengths and limitations. |
| Approach: | They propose to ask LLMs to generate Parallel Chains of Associations to Evaluate their creativity. |
| Outcome: | The proposed framework minimizes the risk of data contamination and offers a highly efficient evaluation. |
EmotionQueen: A Benchmark for Evaluating Empathy of Large Language Models (2024.findings-acl)
Copied to clipboard
| Challenge: | Existing evaluations of emotional intelligence in large language models (LLMs) focus on basic sentiment analysis tasks, such as emotion recognition, which is not enough to evaluate LLMs’ overall emotional intelligence. |
| Approach: | They propose a framework for evaluating the emotional intelligence of large language models (LLMs) that includes four distinct tasks: Key Event Recognition, Mixed Event Recognition and Implicit Emotional Recognition. |
| Outcome: | The proposed framework includes four distinct tasks: Key Event Recognition, Mixed Event Recognition and Implicit Emotional Recognition. |
AEQ-Bench: Measuring Empathy of Omni-Modal Large Models (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing benchmarks focus on cognitive abilities, such as knowledge retrieval, complex reasoning, and instruction following, largely overlooking empathy evaluation. |
| Approach: | They propose to benchmark two core empathetic capabilities of omnimodal large models (OLMs) generating empatries by comprehending affective cues from multi-modal inputs and judging empathy of audio responses without relying on text transcription. |
| Outcome: | The proposed benchmark outperforms existing models with audio output capabilities but is unreliable for evaluating fine-grained paralinguistic expressiveness. |
Bridging Internal Consistency and External Alignment: A Causal and Dynamic Interpretability Framework for LLM Generation (2026.acl-long)
Copied to clipboard
| Challenge: | Existing interpretability methods focus on internal and external aspects of the model . existing explanations often focus on surface correlations or static dependencies . |
| Approach: | They propose a causal and dynamic interpretability framework for Large Language Models . they characterize backdoor-adjusted causal effects of generated prefix and prompt . |
| Outcome: | The proposed framework provides a unified causal view of internal consistency and external alignment in LLM generation dynamics. |
METER: Evaluating Multi-Level Contextual Causal Reasoning in Large Language Models (2026.acl-long)
Copied to clipboard
| Challenge: | Existing benchmarks evaluate contextual causal reasoning in fragmented settings, failing to ensure context consistency or cover the full causal hierarchy. |
| Approach: | They use a unified context to benchmark large language models' contextual causal reasoning skills. |
| Outcome: | The proposed benchmarks show that LLMs are susceptible to distraction by irrelevant but factually correct information at lower level of causality. |
CausalGraph2LLM: Evaluating LLMs for Causal Queries (2025.findings-naacl)
Copied to clipboard
| Challenge: | Recent advances in Large Language Models (LLMs) have opened up new avenues for their use beyond standard Natural Language Processing tasks. |
| Approach: | They propose a benchmark to evaluate the capabilities of Large Language Models (LLMs) they use over 700k queries to compare their encoding capabilities. |
| Outcome: | The proposed benchmark compared LLMs on graph-level and node-level queries and open-sourced and closed models. |
RUBRIC-MQM : Span-Level LLM-as-judge in Machine Translation For High-End Models (2025.acl-industry)
Copied to clipboard
| Challenge: | Existing LLMs are unable to match outputs due to their open-ended nature . |
| Approach: | They propose a meta-evaluation strategy PromptCUE to evaluate cutting-edge LAJ-MT models such as GEMBA-MQM and a rubric-style prompt tailored to the characteristics of LLMs. |
| Outcome: | The proposed model is able to predict scores or identify errors for individual sentences and is reliable in the real world. |
Causal-ESC: Reliable Policy Learning for Emotional Support Conversation via Causal Inference (2026.acl-long)
Copied to clipboard
| Challenge: | Existing approaches to Emotional Support Conversation (ESC) are mechanistically opaque and lacks a causal mechanism between dialogue features and effective empathic strategies. |
| Approach: | They propose a framework that uses Doubly Robust learning to model causal effects of utterance features on strategy selection. |
| Outcome: | The proposed framework outperforms state-of-the-art baselines in empathy and helpfulness and provides a theoretically grounded, interpretable solution to the mechanistic interpretability dilemma in affective computing. |
ESC-Judge: A Framework for Comparing Emotional Support Conversational Agents (2025.emnlp-main)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) increasingly power mental-health chatbots . yet the field lacks a scalable, theory-grounded way to decide which model is more effective to deploy. |
| Approach: | They propose a framework that grounds head-to-head comparisons of Emotional-Support LLMs in Hill’s Exploration–Insight–Action counselling model. |
| Outcome: | The proposed framework matches PhD-level annotators in 85% of Exploration, 83% of Insight, and 86% of Action decisions, demonstrating human-level reliability at a fraction of the cost. |
When Can We Trust LLMs in Mental Health? Large-Scale Benchmarks for Reliable LLM Evaluation (2026.eacl-long)
Copied to clipboard
Abeer Badawi, Elahe Rahimi, Md Tahmid Rahman Laskar, Sheri Grach, Lindsay Bertrand, Lames Danok, Prathiba Dhanesh, Jimmy Huang, Frank Rudzicz, Elham Dolatabadi
| Challenge: | Existing benchmarks for large language models are limited in scale, authenticity, and reliability due to the emotionally complex nature of therapeutic dialogue. |
| Approach: | They propose two benchmarks that provide a framework for evaluating large language models for mental health support. |
| Outcome: | The proposed framework provides a framework for generation and evaluation of large-scale authentic dialogue datasets and judge-reliability assessments. |