PerMemSafe: Benchmarking Implicit Personalized Safety of Long Horizon Self-Evolving Agents (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing self-evolving agents have a low safety rate in long-horizon interactions . however, this reliance on context-independent safety evaluations is insufficient . |
| Approach: | They propose a framework that explicitly models personalized risk inference and memory evolution. |
| Outcome: | The proposed framework improves implicit personalized safety by 23.8% over prior frameworks while maintaining helpfulness in long-horizon interactions. |
Similar Papers
When Personalization Legitimizes Risks: Uncovering Safety Vulnerabilities in Personalized Dialogue Agents (2026.acl-long)
Copied to clipboard
Jiahe Guo, Xiangran Guo, Yulin Hu, Zimo Long, Xingyu Sui, Xuda Zhi, Yongbo Huang, Hao He, Weixiang Zhao, Yanyan Zhao, Bing Qin
| Challenge: | Existing research on personalized LLM agents focuses on the effectiveness of personalized responses. |
| Approach: | They propose a benchmark to quantify intent legitimation in personalized interactions . they propose 'detection-reflection' method that detects intent legititimation from internal representation space . |
| Outcome: | The proposed method reduces safety degradation by using internal representation space. |
On Safety Risks in Experience-Driven Self-Evolving Agents (2026.findings-acl)
Copied to clipboard
Weixiang Zhao, Yichen Zhang, Yingshuo Wang, Yang Deng, Yanyan Zhao, Xuda Zhi, Yongbo Huang, Hao He, Wanxiang Che, Bing Qin, Ting Liu
| Challenge: | Experience-driven self-evolution has emerged as a promising paradigm for improving the autonomy of large language model agents, yet its reliance on self-curated experience introduces underexplored safety risks. |
| Approach: | They investigate how experience accumulation and utilization in self-evolving agents affect safety performance across web-based and embodied environments. |
| Outcome: | The findings expose inherent limitations of current self-evolving agents and call for more principled strategies to ensure safe and reliable adaptation. |
SafetyMem: Adaptive Jailbreak Defense via Dual-Component Safety Memory (2026.acl-long)
Copied to clipboard
| Challenge: | Existing defenses for Large Language Models suffer from a 'memory gap' parameter-modifying methods are computationally expensive and inference-time filters cannot retain or reuse defense knowledge across interactions. |
| Approach: | They propose a framework that secures Large Language Models through a dual-component safety memory system. |
| Outcome: | The proposed framework significantly reduces attack success rates while preserving interpretability and efficiency. |
A Self-Evolving LLM Agent Framework for Role-Based Norm Compliance in Healthcare (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing systems treat roles as static prompts and rely on one-shot safety filters . a self-evolving LLM agent is proposed that learns from role-based social experience . |
| Approach: | They propose a self-evolving LLM agent that learns from role-based social experience and explicitly models communicator-level individual traits informed by prior communication questionnaires and clinical literature. |
| Outcome: | The proposed agent learns from role-based social experience and models communicator-level individual traits informed by prior communication questionnaires and clinical literature. |
AgenticEval: Toward Agentic and Self-Evolving Safety Evaluation of Large Language Models (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing static safety evaluation methods are ill-equipped to address dynamic nature of AI risks and evolving regulations, creating a critical safety gap. |
| Approach: | They propose a new paradigm of agentic safety evaluation reframing evaluation as a continuous and self-evolving process rather than a one-time audit. |
| Outcome: | The proposed framework shows a consistent decline in model safety as the evaluation hardens. |
VestaBench: An Embodied Benchmark for Safe Long-Horizon Planning Under Multi-Constraint and Adversarial Settings (2025.emnlp-industry)
Copied to clipboard
| Challenge: | Existing safety benchmarks do not represent a diverse range of multi-constraint tasks that require long-horizon planning with a focus on safety. |
| Approach: | They propose a benchmark to assess the safety of embodied AI agents under multiple constraints. |
| Outcome: | The proposed benchmarks show that LLMs perform poorly against their tasks . they also suffer significantly compromised safety outcomes . |
ConVerse: Benchmarking Contextual Safety in Agent-to-Agent Conversations (2026.findings-eacl)
Copied to clipboard
| Challenge: | Large language models (LLMs) are rapidly transitioning from passive text generators to autonomous agents that act and communicate on behalf of users. |
| Approach: | a new benchmark evaluates privacy and security risks in agent–agent interactions . a converse model enables attackers to embed malicious requests within plausible discourse . the model is based on a three-tier taxonomy assessing abstraction quality . |
| Outcome: | ConVerse tests privacy and security risks in agent–agent interactions with 12 user personas and over 864 contextually grounded attacks. |
Is Safety Standard Same for Everyone? User-Specific Safety Evaluation of Large Language Models (2025.findings-emnlp)
Copied to clipboard
Yeonjun In, Wonjoong Kim, Kanghoon Yoon, Sungchul Kim, Mehrab Tanjim, Sangwu Park, Kibum Kim, Chanyoung Park
| Challenge: | Extensive benchmarks evaluate LLM safety relying heavily on general standards . no benchmark datasets exist to evaluate the user-specific safety of LLMs . |
| Approach: | a new benchmark is designed to assess user-specific aspect of LLM safety . authors propose a simple remedy based on chain-of-thought to improve user-specified safety. |
| Outcome: | a new benchmark assesses the user-specific aspect of LLM safety . the proposed solution improves user-specified safety by chain-of-thought . |
From Recall to Forgetting: Benchmarking Long-Term Memory for Personalized Agents (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing benchmarks frame long-term memory evaluation as fact retrieval from past conversations, providing limited insight into agents’ ability to consolidate memory over time or handle frequent knowledge updates. |
| Approach: | They propose a long-term memory benchmark that evaluates three memory-grounded tasks: remembering, reasoning, and recommending. |
| Outcome: | The proposed benchmarks evaluate three tasks: remembering, reasoning, and recommending. |
ImplicitMemBench: Measuring Unconscious Behavioral Adaptation in Large Language Models (2026.acl-long)
Copied to clipboard
| Challenge: | Existing memory benchmarks for LLMs evaluate explicit recall of facts, yet overlook implicit memory where experience becomes automated behavior without conscious retrieval. |
| Approach: | They propose a benchmark that evaluates implicit memory using three constructs from non-declarative memory. |
| Outcome: | The new benchmark reframes evaluation from "what agents recall" to "what they automatically enact" no model exceeds 66% overall, with top performers far below human baselines . |