Challenge: Large language models (LLMs) are rapidly transitioning from passive text generators to autonomous agents that act and communicate on behalf of users.
Approach: a new benchmark evaluates privacy and security risks in agent–agent interactions . a converse model enables attackers to embed malicious requests within plausible discourse . the model is based on a three-tier taxonomy assessing abstraction quality .
Outcome: ConVerse tests privacy and security risks in agent–agent interactions with 12 user personas and over 864 contextually grounded attacks.

Similar Papers

Protecting Users From Themselves: Safeguarding Contextual Privacy in Interactions with Conversational Agents (2025.findings-acl)

Copied to clipboard

Challenge: Conversational agents are increasingly woven into individuals’ personal lives, yet users underestimate the privacy risks associated with them.
Approach: They propose a framework that allows users to reformulate out-of-context information in user prompts by identifying and reformulating out- of-content information in the context.
Outcome: The proposed framework can achieve strong gains in contextual privacy while preserving the user’s intended interaction goals.
TAMAS: Benchmarking Adversarial Risks in Multi-Agent LLM Systems (2026.acl-long)

Copied to clipboard

Challenge: Existing benchmarks and datasets focus on single-agent settings, failing to capture the unique vulnerabilities of multi-agend LLM dynamics and co-ordination.
Approach: They propose a benchmark to evaluate the robustness and safety of multi-agent LLM systems.
Outcome: The proposed benchmark evaluates the robustness and safety of multi-agent LLM systems.
Persona-Grounded Safety Evaluation of AI Companions in Multi-Turn Conversations (2026.acl-long)

Copied to clipboard

Challenge: Existing safety evaluations rely on self-reported user data or interviews . a recent study evaluated how Replika responds to high-risk user groups .
Approach: They propose a framework for controlled simulation and safety evaluation of multi-turn interactions with AI companion applications.
Outcome: The proposed framework evaluates how Replika responds to high-risk user groups . it incorporates emotion modeling and LLM-assisted utterance-and harm-level classification .
Privacy in Action: Towards Realistic Privacy Mitigation and Evaluation for LLM-Powered Agents (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing benchmarks for privacy performance of LLM agents are limited to static, simplified scenarios.
Approach: They propose a model-agnostic, contextual integrity based mitigation approach that effectively reduces privacy leakage from 36.08% to 7.30% on DeepSeek-R1 and from 33.06% to 8.32% on GPT-4o.
Outcome: The proposed approach reduces privacy leakage from 36.08% to 7.30% on DeepSeek-R1 and from 33.06% to 8.32% on GPT-4o while preserving task helpfulness.
SAGE: A Generic Framework for LLM Safety Evaluation (2025.emnlp-industry)

Copied to clipboard

Challenge: Current safety evaluation methodologies focus on single-turn interactions with generic policies, failing to capture conversational dynamics of real-world usage and application-specific harms.
Approach: They propose a framework for customized and dynamic harm evaluations that employs prompted adversarial agents with diverse personalities based on the Big Five model.
Outcome: The proposed framework enables system-aware multi-turn conversations that adapt to target applications and harm policies.
PAC-BENCH: Evaluating Multi-Agent Collaboration under Privacy Constraints (2026.findings-acl)

Copied to clipboard

Challenge: Recent research explores multi-agent systems where agents collaborate toward shared goals to handle complex tasks.
Approach: They propose a benchmark for systematic evaluation of multi-agent collaboration under privacy constraints.
Outcome: The proposed benchmark shows that privacy constraints degrade collaboration performance and make outcomes depend more on the initiating agent than the partner.
When Personalization Legitimizes Risks: Uncovering Safety Vulnerabilities in Personalized Dialogue Agents (2026.acl-long)

Copied to clipboard

Challenge: Existing research on personalized LLM agents focuses on the effectiveness of personalized responses.
Approach: They propose a benchmark to quantify intent legitimation in personalized interactions . they propose 'detection-reflection' method that detects intent legititimation from internal representation space .
Outcome: The proposed method reduces safety degradation by using internal representation space.
R-Judge: Benchmarking Safety Risk Awareness for LLM Agents (2024.findings-emnlp)

Copied to clipboard

Challenge: Large language models (LLMs) have shown compelling abilities in reasoning, decision-making, and instruction following.
Approach: They propose a benchmark to evaluate the proficiency of large language models (LLMs) in judging and identifying safety risks given agent interaction records.
Outcome: The proposed model outperforms the best-performing model, GPT-4o, while no other models significantly exceed the random.
PerMemSafe: Benchmarking Implicit Personalized Safety of Long Horizon Self-Evolving Agents (2026.findings-acl)

Copied to clipboard

Challenge: Existing self-evolving agents have a low safety rate in long-horizon interactions . however, this reliance on context-independent safety evaluations is insufficient .
Approach: They propose a framework that explicitly models personalized risk inference and memory evolution.
Outcome: The proposed framework improves implicit personalized safety by 23.8% over prior frameworks while maintaining helpfulness in long-horizon interactions.
CI-Work: Benchmarking Contextual Integrity in Enterprise LLM Agents (2026.acl-industry)

Copied to clipboard

Challenge: Enterprise LLM agents can dramatically improve workplace productivity, but their core capability, retrieving and using internal context to act on a user’s behalf, also creates new risks for sensitive information leakage.
Approach: They propose a Contextual Integrity-grounded benchmark that simulates enterprise workflows across five information-flow directions and evaluates whether agents can convey *essential* content while withholding *sensitive* context in dense retrieval settings.
Outcome: The proposed model demonstrates that privacy failures are prevalent in enterprise workflows and that higher task utility correlates with increased privacy violations.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations