Challenge: Existing frameworks focused on general risks may not adequately capture nuanced risks of LLMs in caregiving contexts.
Approach: They propose a theory-driven, clinician-validated framework for evaluating risks in LLMs . RubRIX operationalizes five empirically-derived risk dimensions: Inattention, Bias Stigma, Information Inaccuracy, Uncritical Affirmation, and Epistemic Arrogance.
Outcome: The proposed framework reduces risk components by 45-98% after one iteration across models.

Similar Papers

MHSafeEval: Role-Aware Interaction-Level Evaluation of Mental Health Safety in Large Language Models (2026.findings-acl)

Copied to clipboard

Challenge: Existing evaluation frameworks assess isolated responses using coarse-grained taxonomies or static datasets.
Approach: They propose a role-aware mental health safety taxonomy that characterizes clinically significant harm in terms of interactional roles an AI counselor adopts.
Outcome: The proposed framework significantly improves failure-mode coverage and diagnostic granularity.
SAGE: A Generic Framework for LLM Safety Evaluation (2025.emnlp-industry)

Copied to clipboard

Challenge: Current safety evaluation methodologies focus on single-turn interactions with generic policies, failing to capture conversational dynamics of real-world usage and application-specific harms.
Approach: They propose a framework for customized and dynamic harm evaluations that employs prompted adversarial agents with diverse personalities based on the Big Five model.
Outcome: The proposed framework enables system-aware multi-turn conversations that adapt to target applications and harm policies.
Responsible Evaluation of AI for Mental Health (2026.acl-long)

Copied to clipboard

Challenge: Existing approaches to evaluating AI tools in this domain remain fragmented and inconsistent.
Approach: They propose a taxonomy of AI mental health support types that integrates clinical soundness, social context, and equity to provide a structured basis for evaluation.
Outcome: The proposed framework integrates clinical soundness, social context, and equity, providing a structured basis for evaluation.
Like a Therapist, But Not: Reddit Narratives of AI in Mental Health Contexts (2026.findings-acl)

Copied to clipboard

Challenge: Large language models are increasingly used for emotional support and mental health–related interactions outside clinical settings.
Approach: They analyze 5,126 Reddit posts describing use of AI for emotional support or therapy . positive sentiment is most strongly associated with task and goal alignment, they say .
Outcome: The proposed framework analyzes language, adoption-related attitudes, and relational alignment at scale. positive sentiment is most strongly associated with task and goal alignment.
Persona-Grounded Safety Evaluation of AI Companions in Multi-Turn Conversations (2026.acl-long)

Copied to clipboard

Challenge: Existing safety evaluations rely on self-reported user data or interviews . a recent study evaluated how Replika responds to high-risk user groups .
Approach: They propose a framework for controlled simulation and safety evaluation of multi-turn interactions with AI companion applications.
Outcome: The proposed framework evaluates how Replika responds to high-risk user groups . it incorporates emotion modeling and LLM-assisted utterance-and harm-level classification .
RubricBench: Aligning Model-Generated Rubrics with Human Standards (2026.acl-long)

Copied to clipboard

Challenge: Existing benchmarks lack discriminative complexity and ground-truth rubric annotations required for rigorous evaluation.
Approach: They propose a curated benchmark with 1,147 pairwise comparisons to assess the reliability of rubric-based evaluation.
Outcome: The proposed benchmarks show that they support diverse domains, exhibit discriminative ability, provide high-quality annotations, and include human-authored rubrics.
StealthGraph: Exposing Domain-Specific Risks in LLMs through Knowledge-Graph-Guided Harmful Prompt Generation (2026.acl-long)

Copied to clipboard

Challenge: Domain-specific datasets of harmful prompts are scarce and often rely on manual construction. Existing efforts to improve domain knowledge and reduce harmful prompt generation are lacking.
Approach: They propose a framework that transforms domain knowledge into actionable constraints and increases the implicitness of generated harmful prompts.
Outcome: The proposed framework yields high-quality datasets combining strong domain relevance with implicitness, enabling more realistic red-teaming and advancing LLM safety research.
At Your Own PACE: A Causal Framework for Evaluating EQ in LLMs (2026.findings-acl)

Copied to clipboard

Challenge: Emotional Quotient (EQ) has emerged as a competency for seamless human-AI integration.
Approach: They propose a framework for a closed-loop EQ evaluation using a PACE taxonomy to define four dimensions of LLM EQ.
Outcome: The proposed framework achieves high alignment of 89.31% with human preferences while maintaining robust consistency of 83.6%.
Are LLMs effective psychological assessors? Leveraging adaptive RAG for interpretable mental health screening through psychometric practice (2025.acl-long)

Copied to clipboard

Challenge: standardized questionnaires are essential tools for mental health screening, but computational approaches bypass these tools in favor of black-box classification.
Approach: They propose a questionnaire-guided screening framework that bridges psychological practice and computational methods through adaptive Retrieval-Augmented Generation.
Outcome: The proposed framework matches or outperforms state-of-the-art performance on Reddit-based benchmarks and extends to self-harm screening.
Can AI Relate: Testing Large Language Model Response for Mental Health Support (2024.findings-emnlp)

Copied to clipboard

Challenge: Large language models (LLMs) are already being piloted for clinical use in hospitals . recent failures of the Tessa chatbot have led to doubts about their reliability in high-stakes settings.
Approach: They propose safety guidelines for the potential deployment of large language models for mental health response.
Outcome: The proposed framework measures equity in empathy and adherence of LLM responses to motivational interviewing theory.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations