Challenge: Existing self-evolving agents have a low safety rate in long-horizon interactions . however, this reliance on context-independent safety evaluations is insufficient .
Approach: They propose a framework that explicitly models personalized risk inference and memory evolution.
Outcome: The proposed framework improves implicit personalized safety by 23.8% over prior frameworks while maintaining helpfulness in long-horizon interactions.

Similar Papers

When Personalization Legitimizes Risks: Uncovering Safety Vulnerabilities in Personalized Dialogue Agents (2026.acl-long)

Copied to clipboard

Challenge: Existing research on personalized LLM agents focuses on the effectiveness of personalized responses.
Approach: They propose a benchmark to quantify intent legitimation in personalized interactions . they propose 'detection-reflection' method that detects intent legititimation from internal representation space .
Outcome: The proposed method reduces safety degradation by using internal representation space.
On Safety Risks in Experience-Driven Self-Evolving Agents (2026.findings-acl)

Copied to clipboard

Challenge: Experience-driven self-evolution has emerged as a promising paradigm for improving the autonomy of large language model agents, yet its reliance on self-curated experience introduces underexplored safety risks.
Approach: They investigate how experience accumulation and utilization in self-evolving agents affect safety performance across web-based and embodied environments.
Outcome: The findings expose inherent limitations of current self-evolving agents and call for more principled strategies to ensure safe and reliable adaptation.
SafetyMem: Adaptive Jailbreak Defense via Dual-Component Safety Memory (2026.acl-long)

Copied to clipboard

Challenge: Existing defenses for Large Language Models suffer from a 'memory gap' parameter-modifying methods are computationally expensive and inference-time filters cannot retain or reuse defense knowledge across interactions.
Approach: They propose a framework that secures Large Language Models through a dual-component safety memory system.
Outcome: The proposed framework significantly reduces attack success rates while preserving interpretability and efficiency.
A Self-Evolving LLM Agent Framework for Role-Based Norm Compliance in Healthcare (2026.findings-acl)

Copied to clipboard

Challenge: Existing systems treat roles as static prompts and rely on one-shot safety filters . a self-evolving LLM agent is proposed that learns from role-based social experience .
Approach: They propose a self-evolving LLM agent that learns from role-based social experience and explicitly models communicator-level individual traits informed by prior communication questionnaires and clinical literature.
Outcome: The proposed agent learns from role-based social experience and models communicator-level individual traits informed by prior communication questionnaires and clinical literature.
AgenticEval: Toward Agentic and Self-Evolving Safety Evaluation of Large Language Models (2026.findings-acl)

Copied to clipboard

Challenge: Existing static safety evaluation methods are ill-equipped to address dynamic nature of AI risks and evolving regulations, creating a critical safety gap.
Approach: They propose a new paradigm of agentic safety evaluation reframing evaluation as a continuous and self-evolving process rather than a one-time audit.
Outcome: The proposed framework shows a consistent decline in model safety as the evaluation hardens.
VestaBench: An Embodied Benchmark for Safe Long-Horizon Planning Under Multi-Constraint and Adversarial Settings (2025.emnlp-industry)

Copied to clipboard

Challenge: Existing safety benchmarks do not represent a diverse range of multi-constraint tasks that require long-horizon planning with a focus on safety.
Approach: They propose a benchmark to assess the safety of embodied AI agents under multiple constraints.
Outcome: The proposed benchmarks show that LLMs perform poorly against their tasks . they also suffer significantly compromised safety outcomes .
ConVerse: Benchmarking Contextual Safety in Agent-to-Agent Conversations (2026.findings-eacl)

Copied to clipboard

Challenge: Large language models (LLMs) are rapidly transitioning from passive text generators to autonomous agents that act and communicate on behalf of users.
Approach: a new benchmark evaluates privacy and security risks in agent–agent interactions . a converse model enables attackers to embed malicious requests within plausible discourse . the model is based on a three-tier taxonomy assessing abstraction quality .
Outcome: ConVerse tests privacy and security risks in agent–agent interactions with 12 user personas and over 864 contextually grounded attacks.
Is Safety Standard Same for Everyone? User-Specific Safety Evaluation of Large Language Models (2025.findings-emnlp)

Copied to clipboard

Challenge: Extensive benchmarks evaluate LLM safety relying heavily on general standards . no benchmark datasets exist to evaluate the user-specific safety of LLMs .
Approach: a new benchmark is designed to assess user-specific aspect of LLM safety . authors propose a simple remedy based on chain-of-thought to improve user-specified safety.
Outcome: a new benchmark assesses the user-specific aspect of LLM safety . the proposed solution improves user-specified safety by chain-of-thought .
From Recall to Forgetting: Benchmarking Long-Term Memory for Personalized Agents (2026.findings-acl)

Copied to clipboard

Challenge: Existing benchmarks frame long-term memory evaluation as fact retrieval from past conversations, providing limited insight into agents’ ability to consolidate memory over time or handle frequent knowledge updates.
Approach: They propose a long-term memory benchmark that evaluates three memory-grounded tasks: remembering, reasoning, and recommending.
Outcome: The proposed benchmarks evaluate three tasks: remembering, reasoning, and recommending.
ImplicitMemBench: Measuring Unconscious Behavioral Adaptation in Large Language Models (2026.acl-long)

Copied to clipboard

Challenge: Existing memory benchmarks for LLMs evaluate explicit recall of facts, yet overlook implicit memory where experience becomes automated behavior without conscious retrieval.
Approach: They propose a benchmark that evaluates implicit memory using three constructs from non-declarative memory.
Outcome: The new benchmark reframes evaluation from "what agents recall" to "what they automatically enact" no model exceeds 66% overall, with top performers far below human baselines .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations