ConVerse: Benchmarking Contextual Safety in Agent-to-Agent Conversations (2026.findings-eacl)
Copied to clipboard
| Challenge: | Large language models (LLMs) are rapidly transitioning from passive text generators to autonomous agents that act and communicate on behalf of users. |
| Approach: | a new benchmark evaluates privacy and security risks in agent–agent interactions . a converse model enables attackers to embed malicious requests within plausible discourse . the model is based on a three-tier taxonomy assessing abstraction quality . |
| Outcome: | ConVerse tests privacy and security risks in agent–agent interactions with 12 user personas and over 864 contextually grounded attacks. |
Similar Papers
Protecting Users From Themselves: Safeguarding Contextual Privacy in Interactions with Conversational Agents (2025.findings-acl)
Copied to clipboard
Ivoline C. Ngong, Swanand Ravindra Kadhe, Hao Wang, Keerthiram Murugesan, Justin D. Weisz, Amit Dhurandhar, Karthikeyan Natesan Ramamurthy
| Challenge: | Conversational agents are increasingly woven into individuals’ personal lives, yet users underestimate the privacy risks associated with them. |
| Approach: | They propose a framework that allows users to reformulate out-of-context information in user prompts by identifying and reformulating out- of-content information in the context. |
| Outcome: | The proposed framework can achieve strong gains in contextual privacy while preserving the user’s intended interaction goals. |
TAMAS: Benchmarking Adversarial Risks in Multi-Agent LLM Systems (2026.acl-long)
Copied to clipboard
| Challenge: | Existing benchmarks and datasets focus on single-agent settings, failing to capture the unique vulnerabilities of multi-agend LLM dynamics and co-ordination. |
| Approach: | They propose a benchmark to evaluate the robustness and safety of multi-agent LLM systems. |
| Outcome: | The proposed benchmark evaluates the robustness and safety of multi-agent LLM systems. |
Persona-Grounded Safety Evaluation of AI Companions in Multi-Turn Conversations (2026.acl-long)
Copied to clipboard
| Challenge: | Existing safety evaluations rely on self-reported user data or interviews . a recent study evaluated how Replika responds to high-risk user groups . |
| Approach: | They propose a framework for controlled simulation and safety evaluation of multi-turn interactions with AI companion applications. |
| Outcome: | The proposed framework evaluates how Replika responds to high-risk user groups . it incorporates emotion modeling and LLM-assisted utterance-and harm-level classification . |
Privacy in Action: Towards Realistic Privacy Mitigation and Evaluation for LLM-Powered Agents (2025.findings-emnlp)
Copied to clipboard
Shouju Wang, Fenglin Yu, Xirui Liu, Xiaoting Qin, Jue Zhang, Qingwei Lin, Dongmei Zhang, Saravan Rajmohan
| Challenge: | Existing benchmarks for privacy performance of LLM agents are limited to static, simplified scenarios. |
| Approach: | They propose a model-agnostic, contextual integrity based mitigation approach that effectively reduces privacy leakage from 36.08% to 7.30% on DeepSeek-R1 and from 33.06% to 8.32% on GPT-4o. |
| Outcome: | The proposed approach reduces privacy leakage from 36.08% to 7.30% on DeepSeek-R1 and from 33.06% to 8.32% on GPT-4o while preserving task helpfulness. |
SAGE: A Generic Framework for LLM Safety Evaluation (2025.emnlp-industry)
Copied to clipboard
| Challenge: | Current safety evaluation methodologies focus on single-turn interactions with generic policies, failing to capture conversational dynamics of real-world usage and application-specific harms. |
| Approach: | They propose a framework for customized and dynamic harm evaluations that employs prompted adversarial agents with diverse personalities based on the Big Five model. |
| Outcome: | The proposed framework enables system-aware multi-turn conversations that adapt to target applications and harm policies. |
PAC-BENCH: Evaluating Multi-Agent Collaboration under Privacy Constraints (2026.findings-acl)
Copied to clipboard
Minjun Park, Donghyun Kim, Hyeonjong Ju, Seungwon Lim, Dongwook Choi, Taeyoon Kwon, Minju Kim, Jinyoung Yeo
| Challenge: | Recent research explores multi-agent systems where agents collaborate toward shared goals to handle complex tasks. |
| Approach: | They propose a benchmark for systematic evaluation of multi-agent collaboration under privacy constraints. |
| Outcome: | The proposed benchmark shows that privacy constraints degrade collaboration performance and make outcomes depend more on the initiating agent than the partner. |
When Personalization Legitimizes Risks: Uncovering Safety Vulnerabilities in Personalized Dialogue Agents (2026.acl-long)
Copied to clipboard
Jiahe Guo, Xiangran Guo, Yulin Hu, Zimo Long, Xingyu Sui, Xuda Zhi, Yongbo Huang, Hao He, Weixiang Zhao, Yanyan Zhao, Bing Qin
| Challenge: | Existing research on personalized LLM agents focuses on the effectiveness of personalized responses. |
| Approach: | They propose a benchmark to quantify intent legitimation in personalized interactions . they propose 'detection-reflection' method that detects intent legititimation from internal representation space . |
| Outcome: | The proposed method reduces safety degradation by using internal representation space. |
R-Judge: Benchmarking Safety Risk Awareness for LLM Agents (2024.findings-emnlp)
Copied to clipboard
Tongxin Yuan, Zhiwei He, Lingzhong Dong, Yiming Wang, Ruijie Zhao, Tian Xia, Lizhen Xu, Binglin Zhou, Fangqi Li, Zhuosheng Zhang, Rui Wang, Gongshen Liu
| Challenge: | Large language models (LLMs) have shown compelling abilities in reasoning, decision-making, and instruction following. |
| Approach: | They propose a benchmark to evaluate the proficiency of large language models (LLMs) in judging and identifying safety risks given agent interaction records. |
| Outcome: | The proposed model outperforms the best-performing model, GPT-4o, while no other models significantly exceed the random. |
PerMemSafe: Benchmarking Implicit Personalized Safety of Long Horizon Self-Evolving Agents (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing self-evolving agents have a low safety rate in long-horizon interactions . however, this reliance on context-independent safety evaluations is insufficient . |
| Approach: | They propose a framework that explicitly models personalized risk inference and memory evolution. |
| Outcome: | The proposed framework improves implicit personalized safety by 23.8% over prior frameworks while maintaining helpfulness in long-horizon interactions. |
CI-Work: Benchmarking Contextual Integrity in Enterprise LLM Agents (2026.acl-industry)
Copied to clipboard
Wenjie Fu, Xiaoting Qin, Jue Zhang, Qingwei Lin, Lukas Wutschitz, Robert Sim, Saravan Rajmohan, Dongmei Zhang
| Challenge: | Enterprise LLM agents can dramatically improve workplace productivity, but their core capability, retrieving and using internal context to act on a user’s behalf, also creates new risks for sensitive information leakage. |
| Approach: | They propose a Contextual Integrity-grounded benchmark that simulates enterprise workflows across five information-flow directions and evaluates whether agents can convey *essential* content while withholding *sensitive* context in dense retrieval settings. |
| Outcome: | The proposed model demonstrates that privacy failures are prevalent in enterprise workflows and that higher task utility correlates with increased privacy violations. |