RubRIX: Rubric-Driven Risk Mitigation in Caregiver-AI Interactions (2026.findings-acl)
Copied to clipboard
Drishti Goel, Jeongah Lee, Qiuyue Zhong, Violeta J. Rodriguez, Daniel S. Brown, Ravi Karkar, Dong Whi Yoo, Koustuv Saha
| Challenge: | Existing frameworks focused on general risks may not adequately capture nuanced risks of LLMs in caregiving contexts. |
| Approach: | They propose a theory-driven, clinician-validated framework for evaluating risks in LLMs . RubRIX operationalizes five empirically-derived risk dimensions: Inattention, Bias Stigma, Information Inaccuracy, Uncritical Affirmation, and Epistemic Arrogance. |
| Outcome: | The proposed framework reduces risk components by 45-98% after one iteration across models. |
Similar Papers
MHSafeEval: Role-Aware Interaction-Level Evaluation of Mental Health Safety in Large Language Models (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing evaluation frameworks assess isolated responses using coarse-grained taxonomies or static datasets. |
| Approach: | They propose a role-aware mental health safety taxonomy that characterizes clinically significant harm in terms of interactional roles an AI counselor adopts. |
| Outcome: | The proposed framework significantly improves failure-mode coverage and diagnostic granularity. |
SAGE: A Generic Framework for LLM Safety Evaluation (2025.emnlp-industry)
Copied to clipboard
| Challenge: | Current safety evaluation methodologies focus on single-turn interactions with generic policies, failing to capture conversational dynamics of real-world usage and application-specific harms. |
| Approach: | They propose a framework for customized and dynamic harm evaluations that employs prompted adversarial agents with diverse personalities based on the Big Five model. |
| Outcome: | The proposed framework enables system-aware multi-turn conversations that adapt to target applications and harm policies. |
Responsible Evaluation of AI for Mental Health (2026.acl-long)
Copied to clipboard
Hiba Arnaout, Anmol Goel, H. Andrew Schwartz, Steffen T. Eberhardt, Dana Atzil-Slonim, Gavin Doherty, Brian Schwartz, Wolfgang Lutz, Tim Althoff, Munmun De Choudhury, Hamidreza Jamalabadi, Raj Sanjay Shah, Flor Miriam Plaza-del-Arco, Dirk Hovy, Maria Liakata, Iryna Gurevych
| Challenge: | Existing approaches to evaluating AI tools in this domain remain fragmented and inconsistent. |
| Approach: | They propose a taxonomy of AI mental health support types that integrates clinical soundness, social context, and equity to provide a structured basis for evaluation. |
| Outcome: | The proposed framework integrates clinical soundness, social context, and equity, providing a structured basis for evaluation. |
Like a Therapist, But Not: Reddit Narratives of AI in Mental Health Contexts (2026.findings-acl)
Copied to clipboard
| Challenge: | Large language models are increasingly used for emotional support and mental health–related interactions outside clinical settings. |
| Approach: | They analyze 5,126 Reddit posts describing use of AI for emotional support or therapy . positive sentiment is most strongly associated with task and goal alignment, they say . |
| Outcome: | The proposed framework analyzes language, adoption-related attitudes, and relational alignment at scale. positive sentiment is most strongly associated with task and goal alignment. |
Persona-Grounded Safety Evaluation of AI Companions in Multi-Turn Conversations (2026.acl-long)
Copied to clipboard
| Challenge: | Existing safety evaluations rely on self-reported user data or interviews . a recent study evaluated how Replika responds to high-risk user groups . |
| Approach: | They propose a framework for controlled simulation and safety evaluation of multi-turn interactions with AI companion applications. |
| Outcome: | The proposed framework evaluates how Replika responds to high-risk user groups . it incorporates emotion modeling and LLM-assisted utterance-and harm-level classification . |
RubricBench: Aligning Model-Generated Rubrics with Human Standards (2026.acl-long)
Copied to clipboard
Junyi Zhou, Qiyuan Zhang, Yufei Wang, Fuyuan Lyu, Yidong Ming, Can Xu, Qingfeng Sun, Kai Zheng, Peng Kang, Xue Liu, Chen Ma
| Challenge: | Existing benchmarks lack discriminative complexity and ground-truth rubric annotations required for rigorous evaluation. |
| Approach: | They propose a curated benchmark with 1,147 pairwise comparisons to assess the reliability of rubric-based evaluation. |
| Outcome: | The proposed benchmarks show that they support diverse domains, exhibit discriminative ability, provide high-quality annotations, and include human-authored rubrics. |
StealthGraph: Exposing Domain-Specific Risks in LLMs through Knowledge-Graph-Guided Harmful Prompt Generation (2026.acl-long)
Copied to clipboard
| Challenge: | Domain-specific datasets of harmful prompts are scarce and often rely on manual construction. Existing efforts to improve domain knowledge and reduce harmful prompt generation are lacking. |
| Approach: | They propose a framework that transforms domain knowledge into actionable constraints and increases the implicitness of generated harmful prompts. |
| Outcome: | The proposed framework yields high-quality datasets combining strong domain relevance with implicitness, enabling more realistic red-teaming and advancing LLM safety research. |
At Your Own PACE: A Causal Framework for Evaluating EQ in LLMs (2026.findings-acl)
Copied to clipboard
| Challenge: | Emotional Quotient (EQ) has emerged as a competency for seamless human-AI integration. |
| Approach: | They propose a framework for a closed-loop EQ evaluation using a PACE taxonomy to define four dimensions of LLM EQ. |
| Outcome: | The proposed framework achieves high alignment of 89.31% with human preferences while maintaining robust consistency of 83.6%. |
Are LLMs effective psychological assessors? Leveraging adaptive RAG for interpretable mental health screening through psychometric practice (2025.acl-long)
Copied to clipboard
| Challenge: | standardized questionnaires are essential tools for mental health screening, but computational approaches bypass these tools in favor of black-box classification. |
| Approach: | They propose a questionnaire-guided screening framework that bridges psychological practice and computational methods through adaptive Retrieval-Augmented Generation. |
| Outcome: | The proposed framework matches or outperforms state-of-the-art performance on Reddit-based benchmarks and extends to self-harm screening. |
Can AI Relate: Testing Large Language Model Response for Mental Health Support (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Large language models (LLMs) are already being piloted for clinical use in hospitals . recent failures of the Tessa chatbot have led to doubts about their reliability in high-stakes settings. |
| Approach: | They propose safety guidelines for the potential deployment of large language models for mental health response. |
| Outcome: | The proposed framework measures equity in empathy and adherence of LLM responses to motivational interviewing theory. |